Encoding device, decoding device, encoding method, decoding method, and storage medium
By using multiple CTUs to represent sub-picture areas in moving images in video encoding technology, and dynamically determine their shape and position, the problem of improving encoding efficiency and image quality in the prior art is solved, the processing volume and circuit scale are reduced, and the processing speed is improved.
Patent Information
- Application Number
- CN202080062201.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-13
- Filing Date
- 2020-09-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-09-10
AI Technical Summary
When existing video encoding technologies process efficient video data, it is difficult to simultaneously improve encoding efficiency, picture quality, processing volume reduction and circuit scale reduction.
By using multiple encoding tree units (CTUs) in the encoding device to represent the sub-picture areas in the moving image, the area shape, position and size of the sub-picture are dynamically determined, and parameters such as filters, block sizes, motion vectors, etc. are appropriately selected during the encoding and decoding process.
It has achieved improvements in encoding efficiency, image quality, processing volume and circuit scale reduction, while improving processing speed and encoding/decoding flexibility.
Smart Images

Figure CN114521331B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an encoding device, a decoding device, an encoding method, a decoding method, and a storage medium. Background Art
[0002] Video encoding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to process the continuously increasing amount of digital video data in various applications, there is always a need to provide improvements and optimizations to video encoding technology. The present invention relates to further progress, improvements, and optimizations in video encoding.
[0003] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above-described video encoding technology.
[0004] Prior Art Documents
[0005] Non-Patent Documents
[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] Regarding the encoding method as described above, for the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, or the appropriate selection of elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, etc., it is desired to propose a new method.
[0009] The present invention provides a structure or method that can contribute to one or more of, for example, the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, the improvement of processing speed, and the appropriate selection of elements or operations. In addition, the present invention may include a structure or method that can contribute to benefits other than the above.
[0010] Means for Solving the Problems
[0011] For example, an encoding apparatus according to one aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit encodes sub-picture information. The sub-picture information uses a coding tree unit (CTU) to respectively represent the horizontal position and the vertical position of a region of a sub-picture, which is a rectangular region within a picture. The horizontal position is represented by the position of the CTU that constitutes the sub-picture, with reference to the left end of the picture. The vertical position is represented by the position of the CTU that constitutes the sub-picture, with reference to the upper end of the picture.
[0012] In video encoding technology, in order to improve encoding efficiency, improve picture quality, reduce circuit scale, etc., it is desirable to propose new methods.
[0013] Each embodiment or a part of the structure or method in the present invention can respectively achieve at least any one of improvements in encoding efficiency, improvements in picture quality, reduction in the amount of encoding / decoding processing, reduction in circuit scale, or improvement in encoding / decoding processing speed, etc. Or, each embodiment or a part of the structure or method in the present invention can respectively make an appropriate selection of elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, etc. during encoding and decoding. In addition, the present invention also includes the disclosure of structures or methods that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in the amount of processing.
[0014] Based on the specification and the drawings, further advantages and effects in one aspect of the present invention have been clarified. These advantages and / or effects are respectively obtained through several embodiments and the features described in the specification and the drawings, but it is not necessary to provide all of them in order to obtain one or more of the advantages and / or effects.
[0015] In addition, these general or specific aspects can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0016] Advantages of the Invention
[0017] The structure or method according to one aspect of the present invention can, for example, contribute to one or more of improvements in encoding efficiency, improvements in picture quality, reduction in the amount of processing, reduction in circuit scale, improvement in processing speed, and appropriate selection of elements or actions. In addition, the structure or method according to one aspect of the present invention can also contribute to benefits other than the above. Brief Description of the Drawings
[0018] Figure 1It is a schematic diagram showing an example of the structure of the transmission system of the embodiment.
[0019] Figure 2 It is a diagram showing an example of the hierarchical structure of the data in the stream.
[0020] Figure 3 It is a diagram showing an example of the structure of the slice.
[0021] Figure 4 It is a diagram showing an example of the structure of the tile.
[0022] Figure 5 It is a diagram showing an example of the coding structure in scalable coding.
[0023] Figure 6 It is a diagram showing an example of the coding structure in scalable coding.
[0024] Figure 7 It is a block diagram showing an example of the functional structure of the coding device of the embodiment.
[0025] Figure 8 It is a block diagram showing an example of the installation of the coding device.
[0026] Fig. 9 It is a flowchart showing an example of the overall coding process performed by the coding device.
[0027] Fig.10 It is a diagram showing an example of block segmentation.
[0028] Fig.11 It is a diagram showing an example of the functional structure of the segmentation unit.
[0029] Fig.12 It is a diagram showing an example of the segmentation pattern.
[0030] Fig.13A It is a diagram showing an example of the syntax tree of the segmentation pattern.
[0031] Fig. 13B It is a diagram showing another example of the syntax tree of the segmentation pattern.
[0032] Fig.14 It is a table showing the transform basis functions corresponding to each transform type.
[0033] Fig.15 It is a diagram showing an example of SVT.
[0034] Fig.16 It is a flowchart showing an example of the process performed by the transform unit.
[0035] Fig.17 It is a flowchart showing another example of the process performed by the transform unit.
[0036] Fig.18 It is a block diagram showing an example of the functional structure of the quantization unit.
[0037] Fig.19 It is a flowchart showing an example of the quantization performed by the quantization unit.
[0038] Fig. 20 It is a block diagram showing an example of the functional structure of the entropy encoding unit.
[0039] Fig.21 It is a diagram showing the process of CABAC in the entropy encoding unit.
[0040] Fig. 22 It is a block diagram showing an example of the functional structure of the loop filtering unit.
[0041] Fig.23A It is a diagram showing an example of the shape of the filter used in ALF (adaptive loop filter).
[0042] Fig. 23B It is a diagram showing another example of the shape of the filter used in ALF.
[0043] Fig.23C It is a diagram showing another example of the shape of the filter used in ALF.
[0044] Fig.23D It is a diagram showing an example where the Y sample (first component) is used for CCALF of Cb and Cr (multiple components different from the first component).
[0045] Fig.23E It is a diagram showing a diamond-shaped filter.
[0046] Fig.23F It is a diagram showing an example of JC-CCALF.
[0047] Figure 23G It is a diagram showing an example of the weight_index candidate of JC-CCALF.
[0048] Fig.24 It is a block diagram showing an example of the detailed structure of the loop filtering unit that functions as a DBF.
[0049] Fig.25 It is a diagram showing an example of deblocking filtering with filtering characteristics symmetric with respect to the block boundary.
[0050] Fig.26 It is a diagram for explaining an example of the block boundary where deblocking filtering is performed.
[0051] Fig. 27It is a diagram showing an example of the Bs value.
[0052] Fig.28 It is a flowchart showing an example of the processing performed by the prediction unit of the encoding device.
[0053] Fig.29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.
[0054] Fig.30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.
[0055] Fig.31 It is a diagram showing an example of 67 intra prediction modes in intra prediction.
[0056] Fig.32 It is a flowchart showing an example of the processing performed by the intra prediction unit.
[0057] Fig.33 It is a diagram showing an example of each reference picture.
[0058] Fig.34 It is a conceptual diagram showing an example of a reference picture list.
[0059] Fig.35 It is a flowchart showing the flow of the basic processing of inter prediction.
[0060] Fig.36 It is a flowchart showing an example of MV derivation.
[0061] Fig.37 It is a flowchart showing another example of MV derivation.
[0062] Fig.38A It is a diagram showing an example of the classification of each mode of MV derivation.
[0063] Fig.38B It is a diagram showing an example of the classification of each mode of MV derivation.
[0064] Fig.39 It is a flowchart showing an example of inter prediction based on the normal inter mode.
[0065] Fig.40 It is a flowchart showing an example of inter prediction based on the normal merge mode.
[0066] Fig.41 It is a diagram for explaining an example of the MV derivation process based on the normal merge mode.
[0067] Fig.42 It is a diagram for explaining an example of the MV derivation process based on the HMVP mode.
[0068] Fig.43 It is a flowchart showing an example of FRUC (frame rate up conversion).
[0069] Fig.44 It is a diagram showing an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0070] Fig.45 It is a diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture.
[0071] Fig.46A It is a diagram showing an example of the derivation of the MV in sub-block units in the affine mode using two control points.
[0072] Fig.46B It is a diagram showing an example of the derivation of the MV in sub-block units in the affine mode using three control points.
[0073] Fig.47A It is a conceptual diagram showing an example of the derivation of the MV of the control points in the affine mode.
[0074] Fig.47B It is a conceptual diagram showing an example of the derivation of the MV of the control points in the affine mode.
[0075] Fig.47C It is a conceptual diagram showing an example of the derivation of the MV of the control points in the affine mode.
[0076] Fig.48A It is a diagram showing the affine mode with two control points.
[0077] Fig.48B It is a diagram showing the affine mode with three control points.
[0078] Fig.49A It is a conceptual diagram showing an example of the method for deriving the MV of the control points when the number of control points in the encoded block and the current block is different.
[0079] Fig.49B It is a conceptual diagram showing another example of the method for deriving the MV of the control points when the number of control points in the encoded block and the current block is different.
[0080] Fig.50 It is a flowchart showing an example of the process of the affine merge mode.
[0081] Fig.51 It is a flowchart showing an example of the process of the affine inter-frame mode.
[0082] Fig.52A It is a diagram for explaining the generation of predicted images of two triangles.
[0083] Fig.52B It is a conceptual diagram showing an example of the first part of the first partition, the first sample set, and the second sample set.
[0084] Fig.52C It is a conceptual diagram showing the first part of the first partition.
[0085] Fig.53 It is a flowchart showing an example of a triangular pattern.
[0086] Fig.54 It is a diagram showing an example of the ATMVP mode for deriving MV in sub-block units.
[0087] Fig.55 It is a diagram showing the relationship between the merge mode and DMVR (dynamic motion vector refreshing).
[0088] Fig.56 It is a conceptual diagram for explaining an example of DMVR.
[0089] Fig.57 It is a conceptual diagram for explaining another example of DMVR for determining MV.
[0090] Fig.58A It is a diagram showing an example of motion search in DMVR.
[0091] Fig.58B It is a flowchart showing an example of motion search in DMVR.
[0092] Fig.59 It is a flowchart showing an example of the generation of a predicted image.
[0093] Fig.60 It is a flowchart showing another example of the generation of a predicted image.
[0094] Fig.61 It is a flowchart for explaining an example of the predicted image correction process based on OBMC (overlapped block motion compensation).
[0095] Fig.62 It is a conceptual diagram for explaining an example of the predicted image correction process based on OBMC.
[0096] Fig.63 It is a diagram for explaining a model assuming uniform linear motion.
[0097] Fig.64 It is a flowchart showing an example of inter-frame prediction according to BIO.
[0098] Fig.65 It is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.
[0099] Fig.66A It is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC (local illumination compensation).
[0100] Fig.66B It is a flowchart showing an example of a method for generating a predicted image using a luminance correction process based on LIC.
[0101] Fig.67 It is a block diagram showing the functional structure of a decoding device according to an embodiment.
[0102] Fig.68 It is a block diagram showing an example of the installation of a decoding device.
[0103] Fig.69 It is a flowchart showing an example of the overall decoding process performed by a decoding device.
[0104] Fig.70 It is a diagram showing the relationship between a segmentation determination unit and other components.
[0105] Fig.71 It is a block diagram showing an example of the functional structure of an entropy decoding unit.
[0106] Fig.72 It is a diagram showing the process of CABAC in an entropy decoding unit.
[0107] Fig.73 It is a block diagram showing an example of the functional structure of an inverse quantization unit.
[0108] Fig.74 It is a flowchart showing an example of inverse quantization performed by an inverse quantization unit.
[0109] Fig.75 It is a flowchart showing an example of the process performed by an inverse transform unit.
[0110] Fig.76 It is a flowchart showing another example of the process performed by an inverse transform unit.
[0111] Fig.77 It is a block diagram showing an example of the functional structure of a loop filter unit.
[0112] Fig.78It is a flowchart showing an example of the processing performed by the prediction unit of the decoding device.
[0113] Fig.79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device.
[0114] Fig.80A It is a flowchart showing a part of another example of the processing performed by the prediction unit of the decoding device.
[0115] Fig.80B It is a flowchart showing the remaining part of another example of the processing performed by the prediction unit of the decoding device.
[0116] Fig.81 It is a diagram showing an example of the processing performed by the intra prediction unit of the decoding device.
[0117] Fig.82 It is a flowchart showing an example of MV derivation in the decoding device.
[0118] Fig.83 It is a flowchart showing another example of MV derivation in the decoding device.
[0119] Fig.84 It is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding device.
[0120] Fig.85 It is a flowchart showing an example of inter prediction based on the normal merge mode in the decoding device.
[0121] Fig.86 It is a flowchart showing an example of inter prediction based on the FRUC mode in the decoding device.
[0122] Fig.87 It is a flowchart showing an example of inter prediction based on the affine merge mode in the decoding device.
[0123] Fig.88 It is a flowchart showing an example of inter prediction based on the affine inter mode in the decoding device.
[0124] Fig.89 It is a flowchart showing an example of inter prediction based on the triangle mode in the decoding device.
[0125] Fig.90 It is a flowchart showing an example of motion search based on DMVR in the decoding device.
[0126] Fig.91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device.
[0127] Fig.92It is a flowchart showing an example of the generation of a predicted image in a decoding device.
[0128] Fig.93 It is a flowchart showing another example of the generation of a predicted image in a decoding device.
[0129] Fig.94 It is a flowchart showing an example of the correction of a predicted image based on OBMC in a decoding device.
[0130] Fig.95 It is a flowchart showing an example of the correction of a predicted image based on BIO in a decoding device.
[0131] Fig.96 It is a flowchart showing an example of the correction of a predicted image based on LIC in a decoding device.
[0132] Fig.97 It is a conceptual diagram showing the relationship between sub - pictures, slices, and tiles.
[0133] Fig.98 It is a conceptual diagram showing the relationship between tiles and slices.
[0134] Fig.99 It is a syntactic structure diagram related to sub - pictures determined by a grid.
[0135] Fig.100 It is a syntactic structure diagram related to sub - pictures determined by a CTU.
[0136] Fig.101A It is a conceptual diagram showing the grid index of each grid element in a picture.
[0137] Fig.101B It is a conceptual diagram showing the sub - picture index of each grid element in a picture.
[0138] Fig.102 It is a syntactic structure diagram showing another example related to sub - pictures determined by a grid.
[0139] Fig.103 It is a syntactic structure diagram showing another example related to sub - pictures determined by a CTU.
[0140] Fig.104 It is a flowchart showing the operation of an encoding device according to an embodiment.
[0141] Fig.105 It is a flowchart showing the operation of a decoding device according to an embodiment.
[0142] Fig.106 It is an overall structure diagram of a content supply system for implementing a content distribution service.
[0143] Fig.107It is a diagram showing an example of a display screen of a web page.
[0144] Fig.108 It is a diagram showing an example of a display screen of a web page.
[0145] Fig.109 It is a diagram showing an example of a smartphone.
[0146] Fig.110 It is a block diagram showing an example of the structure of a smartphone. Detailed implementation
[0147] [Introduction]
[0148] In the encoding of moving images, multiple pictures constituting the moving image are each divided into various regions. For example, a picture is divided into multiple CTUs (Coding Tree Units), or divided into multiple tiles, or divided into multiple slices.
[0149] For example, a CTU corresponds to a square region of a fixed size. A tile is a rectangular region determined according to one or more rows within a picture, one or more columns within a picture, or both. A slice corresponds to one NAL unit as a data packet and corresponds to one or more tiles or one or more consecutive CTUs within one tile.
[0150] However, sometimes the processing target region does not match one of the CTUs, tiles, and slices, and it is sometimes difficult to determine the processing target region.
[0151] Therefore, an encoding apparatus according to one aspect of the present invention includes a circuit and a memory connected to the circuit. In operation, the circuit encodes sub-picture information. The sub-picture information uses multiple CTUs (Coding Tree Units) to represent the region of a sub-picture commonly determined in multiple pictures constituting a moving image. The multiple CTUs are common in the multiple pictures and are determined in a grid pattern. The sub-picture information includes information on the horizontal position and the vertical position of the region of the sub-picture commonly determined in the multiple pictures using the multiple CTUs.
[0152] Thereby, sometimes the region of the sub-picture can be flexibly determined based on multiple CTUs. In addition, sometimes the horizontal position and the vertical position of the region of the sub-picture can be simply represented based on multiple CTUs. Therefore, sometimes the sub-picture can be appropriately determined as the processing target region.
[0153] For example, the sub-picture information is as follows. This information represents the shape, horizontal position, and vertical position of the area of the sub-picture by using the multiple CTUs, so as to represent the area of the sub-picture by using the multiple CTUs.
[0154] Thus, sometimes the shape, horizontal position, and vertical position of the area of the sub-picture can be flexibly determined based on multiple CTUs.
[0155] In addition, for example, the sub-picture information is as follows. This information represents the width, height, horizontal position, and vertical position of the area of the sub-picture by using the multiple CTUs, so as to represent the area of the sub-picture by using the multiple CTUs.
[0156] Thus, sometimes the width, height, horizontal position, and vertical position of the area of the sub-picture can be flexibly determined based on multiple CTUs.
[0157] In addition, for example, the sub-picture information is as follows. This information represents the horizontal position and vertical position of the upper left corner of the area of the sub-picture and the horizontal position and vertical position of the lower right corner of the area of the sub-picture as the horizontal position and vertical position of the area of the sub-picture, so as to represent the area of the sub-picture by using the multiple CTUs.
[0158] Thus, sometimes the horizontal position and vertical position of the upper left and lower right corners of the area of the sub-picture can be flexibly determined based on multiple CTUs.
[0159] In addition, for example, the sub-picture information is as follows. This information represents (1) the horizontal position of the upper left corner of the area of the sub-picture as the horizontal position of the area of the sub-picture, (2) the vertical position of the upper left corner of the area of the sub-picture as the vertical position of the area of the sub-picture, (3) the difference between the horizontal position of the upper left corner of the area of the sub-picture and the horizontal position of the upper right corner of the area of the sub-picture, and (4) the difference between the vertical position of the upper left corner of the area of the sub-picture and the vertical position of the lower left corner of the area of the sub-picture, so as to represent the area of the sub-picture by using the multiple CTUs.
[0160] Thus, sometimes the horizontal position and vertical position of the upper left corner of the area of the sub-picture can be flexibly determined based on multiple CTUs. In addition, sometimes the width and height of the area of the sub-picture can be flexibly determined based on multiple CTUs. In addition, sometimes the coding amount can be reduced.
[0161] In addition, for example, the circuit further encodes the CTU size information earlier than the sub-picture information, and the CTU size information represents the size commonly determined for the multiple CTUs.
[0162] Thus, sometimes, the area of the sub-picture can be appropriately determined based on a plurality of CTUs whose sizes are determined.
[0163] In addition, for example, the circuit encodes the sub-picture information only when there are two or more sub-pictures as the sub-pictures commonly determined in the plurality of pictures.
[0164] Thus, sometimes, the signaling of unnecessary information can be omitted. Sometimes, the amount of coding can be reduced.
[0165] In addition, for example, the circuit encodes the sub-picture information into a sequence parameter set.
[0166] Thus, sometimes, the area of the sub-picture can be appropriately determined for a sequence including a plurality of pictures.
[0167] In addition, for example, in each of the plurality of pictures, a CTU that is neither at the right end nor at the lower end of the picture among the plurality of CTUs is determined as a square area of a fixed size, and a CTU that is at the right end or the lower end of the picture among the plurality of CTUs is determined as the square area of the fixed size or an area smaller than the square area of the fixed size.
[0168] Thus, sometimes, the sizes of the picture and the CTU can be determined flexibly.
[0169] Furthermore, for example, a decoding device according to one aspect of the present invention includes a circuit and a memory connected to the circuit. In operation, the circuit decodes sub-picture information that uses a plurality of CTUs (Coding Tree Units) to represent an area of a sub-picture commonly determined in a plurality of pictures that constitute a moving image. The plurality of CTUs are common and determined in a grid pattern in the plurality of pictures, and the sub-picture information includes information on the horizontal position and the vertical position of the area of the sub-picture commonly determined in the plurality of pictures using the plurality of CTUs.
[0170] Thus, sometimes, the area of the sub-picture can be determined flexibly based on a plurality of CTUs. In addition, sometimes, the horizontal position and the vertical position of the area of the sub-picture can be simply represented based on a plurality of CTUs. Therefore, sometimes, the sub-picture can be appropriately determined as a processing target area.
[0171] In addition, for example, the sub-picture information is information that represents the area of the sub-picture by using the plurality of CTUs to represent the shape, the horizontal position, and the vertical position of the area of the sub-picture.
[0172] Thus, sometimes, the shape, horizontal position, and vertical position of the region of the sub-picture can be flexibly determined based on multiple CTUs.
[0173] In addition, for example, the sub-picture information is information that represents the width, height, the horizontal position, and the vertical position of the region of the sub-picture by using the multiple CTUs, so as to represent the region of the sub-picture by using the multiple CTUs.
[0174] Thus, sometimes, the width, height, horizontal position, and vertical position of the region of the sub-picture can be flexibly determined based on multiple CTUs.
[0175] In addition, for example, the sub-picture information is information that represents the horizontal position and vertical position of the upper left of the region of the sub-picture and the horizontal position and vertical position of the lower right of the region of the sub-picture as the horizontal position and the vertical position of the region of the sub-picture, so as to represent the region of the sub-picture by using the multiple CTUs.
[0176] Thus, sometimes, the horizontal position and vertical position of the upper left and lower right of the region of the sub-picture can be flexibly determined based on multiple CTUs.
[0177] In addition, for example, the sub-picture information is information that represents (1) the horizontal position of the upper left of the region of the sub-picture as the horizontal position of the region of the sub-picture, (2) the vertical position of the upper left of the region of the sub-picture as the vertical position of the region of the sub-picture, (3) the difference between the horizontal position of the upper left of the region of the sub-picture and the horizontal position of the upper right of the region of the sub-picture, and (4) the difference between the vertical position of the upper left of the region of the sub-picture and the vertical position of the lower left of the region of the sub-picture, so as to represent the region of the sub-picture by using the multiple CTUs.
[0178] Thus, sometimes, the horizontal position and vertical position of the upper left of the region of the sub-picture can be flexibly determined based on multiple CTUs. In addition, sometimes, the width and height of the region of the sub-picture can be flexibly determined based on multiple CTUs. In addition, sometimes, the coding amount can be reduced.
[0179] In addition, for example, the circuit further decodes the CTU size information earlier than the sub-picture information, and the CTU size information represents a size commonly determined for the multiple CTUs.
[0180] Thus, sometimes, the region of the sub-picture can be appropriately determined based on the multiple CTUs whose sizes are determined.
[0181] In addition, for example, the circuit decodes the sub-picture information only when there are two or more sub-pictures as the sub-pictures commonly determined among the plurality of pictures.
[0182] Thereby, signalization of unnecessary information can sometimes be omitted. Sometimes, the coding amount can be reduced.
[0183] In addition, for example, the circuit decodes the sub-picture information from a sequence parameter set.
[0184] Thereby, the area of the sub-picture can sometimes be appropriately determined for a sequence including a plurality of pictures.
[0185] In addition, for example, in each of the plurality of pictures, a CTU that is neither at the right end nor at the lower end of the picture among the plurality of CTUs is determined as a square region of a fixed size, and a CTU that is at the right end or the lower end of the picture among the plurality of CTUs is determined as the square region of the fixed size or a region smaller than the square region of the fixed size.
[0186] Thereby, the size of the picture and the size of the CTU can sometimes be determined flexibly.
[0187] In addition, for example, a coding method according to one aspect of the present invention encodes sub-picture information, where the sub-picture information uses a plurality of CTUs (Coding Tree Unit) to represent a region of a sub-picture commonly determined among a plurality of pictures constituting a moving image, the plurality of CTUs are common among the plurality of pictures and are determined in a grid pattern, and the sub-picture information includes information on the horizontal position and the vertical position of the region of the sub-picture commonly determined among the plurality of pictures using the plurality of CTUs.
[0188] Thereby, the region of the sub-picture can sometimes be determined flexibly based on a plurality of CTUs. In addition, the horizontal position and the vertical position of the region of the sub-picture can sometimes be simply represented based on a plurality of CTUs. Therefore, the sub-picture can sometimes be appropriately determined as a processing target region.
[0189] In addition, for example, a decoding method according to one aspect of the present invention decodes sub-picture information, where the sub-picture information uses a plurality of CTUs (Coding Tree Unit) to represent a region of a sub-picture commonly determined among a plurality of pictures constituting a moving image, the plurality of CTUs are common among the plurality of pictures and are determined in a grid pattern, and the sub-picture information includes information on the horizontal position and the vertical position of the region of the sub-picture commonly determined among the plurality of pictures using the plurality of CTUs.
[0190] Thus, sometimes the region of the sub-picture can be determined flexibly based on multiple CTUs. In addition, sometimes the horizontal position and the vertical position of the region of the sub-picture can be simply represented based on multiple CTUs. Therefore, sometimes the sub-picture can be appropriately determined as the processing target region.
[0191] In addition, for example, an encoding apparatus according to one aspect of the present invention includes an input unit, a segmentation unit, an intra prediction unit, an inter prediction unit, a loop filter unit, a transform unit, a quantization unit, an entropy encoding unit, and an output unit.
[0192] The current picture is input to the input unit. The segmentation unit divides the current picture into a plurality of blocks.
[0193] The intra prediction unit generates a prediction signal of a current block included in the current picture using a reference image included in the current picture. The inter prediction unit generates a prediction signal of the current block included in the current picture using a reference image included in a reference picture different from the current picture. The loop filter unit applies a filter to a reconstructed block of the current block included in the current picture.
[0194] The transform unit transforms a prediction error between an original signal of the current block included in the current picture and a prediction signal generated by the intra prediction unit or the inter prediction unit to generate transform coefficients. The quantization unit quantizes the transform coefficients to generate quantized coefficients. The entropy encoding unit applies variable length coding to the quantized coefficients to generate a coded bitstream. Then, the coded bitstream including the quantized coefficients to which variable length coding has been applied and control information is output from the output unit.
[0195] Further, for example, in operation, the entropy encoding unit encodes sub-picture information that uses a plurality of CTUs (Coding Tree Units) to represent a region of a sub-picture that is commonly determined in a plurality of pictures constituting a moving image. The plurality of CTUs are common in the plurality of pictures and are determined in a grid pattern. The sub-picture information includes information representing the horizontal position and the vertical position of the region of the sub-picture that is commonly determined in the plurality of pictures using the plurality of CTUs.
[0196] In addition, for example, a decoding apparatus according to one aspect of the present invention includes an input unit, an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, a loop filter unit, and an output unit.
[0197] The coded bitstream is input to the input unit. The entropy decoding unit applies variable length decoding to the coded bitstream to derive quantized coefficients. The inverse quantization unit inverse quantizes the quantized coefficients to derive transform coefficients. The inverse transform unit inverse transforms the transform coefficients to derive a prediction error.
[0198] The intra prediction unit generates a prediction signal for a current block included in the current picture using a reference picture included in the current picture. The inter prediction unit generates a prediction signal for the current block included in the current picture using a reference picture included in a reference picture different from the current picture.
[0199] The loop filter unit applies a filter to a reconstructed block of a current block included in the current picture. Then, the current picture is output from the output unit.
[0200] In addition, for example, during operation, the entropy decoding unit decodes sub-picture information that uses a plurality of CTUs (Coding Tree Units) to represent a region of a sub-picture that is commonly determined in a plurality of pictures constituting a moving image. The plurality of CTUs are common and determined in a grid pattern in the plurality of pictures, and the sub-picture information includes information on the horizontal position and the vertical position of the region of the sub-picture that is commonly determined in the plurality of pictures using the plurality of CTUs.
[0201] Moreover, these inclusive or specific forms can also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non-transitory recording medium such as a computer-readable CD-ROM, or can be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0202] [Definition of Terms]
[0203] As an example, each term can be defined as follows.
[0204] (1) Image
[0205] Is a unit of data composed of a set of pixels, composed of a picture or a block smaller than a picture, and includes still images in addition to moving images.
[0206] (2) Picture
[0207] Is a processing unit of an image composed of a set of pixels, and is sometimes referred to as a frame or a field.
[0208] (3) Block
[0209] Is a processing unit including a set of a specific number of pixels, and as listed in the following examples, the name is not limited. In addition, the shape is not limited, for example, of course, includes a rectangle composed of M×N pixels, a square composed of M×M pixels, and also includes triangles, circles, and other shapes.
[0210] (Examples of Blocks)
[0211] · Slice / Tile / Brick
[0212] ·CTU / Superblock / Base segmentation unit
[0213] ·VPDU / Hardware processing segmentation unit
[0214] ·CU / Processing block unit / Prediction block unit (PU) / Orthogonal transform block unit (TU) / Unit
[0215] ·Sub-block
[0216] (4) Pixel / Sample
[0217] Is the point that is the smallest unit constituting an image, including not only the pixels at integer positions but also the pixels at fractional positions generated based on the pixels at integer positions.
[0218] (5) Pixel value / Sample value
[0219] Is the inherent value of a pixel, including of course the luminance value, color difference value, grayscale of RGB, and also including the depth value, or binary values of 0 and 1.
[0220] (6) Flag
[0221] In addition to 1 bit, it also includes cases of multiple bits. For example, it can also be a parameter or index of 2 bits or more. In addition, not only the binary-valued two-state, but also multi-state using other number systems can be used.
[0222] (7) Signal
[0223] Is symbolized and encoded for information transmission, including not only the discretized digital signal but also the analog signal taking continuous values.
[0224] (8) Stream / Bitstream
[0225] Refers to a data string of digital data or a stream of digital data. The stream / bitstream can be divided into multiple layers and composed of multiple streams in addition to one stream. In addition, in addition to the case of being transmitted on a single transmission path through serial communication, it also includes the case of being transmitted through packet communication on multiple transmission paths.
[0226] (9) Difference / Differential
[0227] In the case of a scalar, in addition to the simple difference (x - y), as long as it includes the operation of difference, it includes the absolute value of the difference (|x - y|), the square difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a and b are constants), and the offset difference (x - y + a: a is the offset).
[0228] (10) Sum
[0229] In the case of a scalar, in addition to the simple sum (x + y), any operation involving a sum is acceptable, including the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted sum (ax + by where a and b are constants), and the offset sum (x + y + a where a is the offset).
[0230] (11) Based on
[0231] This also includes cases where elements other than those that are the basis object are added. Additionally, in addition to cases where the direct result is obtained, it also includes cases where the result is obtained via intermediate results.
[0232] (12) Used, using
[0233] This also includes cases where elements other than those that are the used object are added. Additionally, in addition to cases where the direct result is obtained, it also includes cases where the result is obtained via intermediate results.
[0234] (13) Prohibit, forbid
[0235] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation.
[0236] (14) Limit, restriction / restrict / restricted
[0237] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation. And as long as a part is prohibited in terms of quantity or quality, it also includes cases of total prohibition.
[0238] (15) Chroma
[0239] It is an adjective represented by the notations Cb and Cr that designates one of two color difference signals associated with the primary colors for a sample arrangement or a single sample representation. Additionally, the term chrominance can also be used instead of the term chroma.
[0240] (16) Luma
[0241] It is an adjective represented by the notation or subscript Y or L that designates a monochrome signal associated with the primary colors for a sample arrangement or a single sample representation. The term luminance can also be used instead of the term luma.
[0242] [Regarding the explanations in the record]
[0243] In the drawings, the same reference numerals denote the same or similar components. In addition, the dimensions and relative positions of the components in the drawings are not necessarily drawn to a specific scale.
[0244] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and do not limit the meaning of the claims.
[0245] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processes and / or structures described in various aspects of the present invention. The processes and / or structures can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processes and / or structures applied to the embodiments, any of the following can be done, for example.
[0246] (1) Among a plurality of components of the encoding device or decoding device of the embodiment described in each aspect of the present invention, a certain component can be replaced with another component described in a certain aspect of the present invention, or they can be combined;
[0247] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. of functions or processes performed by a part of the components of the encoding device or decoding device can also be made. For example, any function or process can be replaced with another function or process described in a certain aspect of the present invention, or they can be combined;
[0248] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can also be made to a part of the processes included in the method. For example, any process in the method can be replaced with another process described in a certain aspect of the present invention, or they can be combined;
[0249] (4) A part of the components constituting the encoding device or decoding device of the embodiment can be combined with the components described in a certain aspect of the present invention, can also be combined with the components having a part of the functions described in a certain aspect of the present invention, or can be combined with the components performing a part of the processes performed by the components described in the aspects of the present invention;
[0250] (5) A component that is part of the function of the encoding device or decoding device of the embodiment, or a component that performs part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced by a component described in a certain one of the various aspects of the present invention, a component that is part of the function described in a certain one of the various aspects of the present invention, or a component that performs part of the processing described in a certain one of the various aspects of the present invention;
[0251] (6) In the method implemented by the encoding device or decoding device of the embodiment, a certain one of the multiple processes included in the method is replaced by a process described in a certain one of the various aspects of the present invention or a similar process, or they are combined;
[0252] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various aspects of the present invention.
[0253] (8) The manner of implementing the processes and / or structures described in the various aspects of the present invention is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the motion picture encoding or motion picture decoding disclosed in the embodiment.
[0254] [System Structure]
[0255] Figure 1 It is a schematic diagram showing an example of the structure of the transmission system of the present embodiment.
[0256] The transmission system Trs is a system that transmits the stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs, for example, as Figure 1 shown, includes an encoding device 100, a network Nw, and a decoding device 200.
[0257] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream, for example, includes the encoded image and control information for decoding the encoded image. The image is compressed through this encoding.
[0258] In addition, the original image input to the encoding device 100 before being encoded is also referred to as the original image, original signal, or original sample. Additionally, the image can be a moving image or a still image. Further, the image is a superordinate concept of sequences, pictures, blocks, etc., and is not restricted by spatial and temporal regions unless otherwise specified. Moreover, the image is composed of an arrangement of pixels or pixel values, and the signal or pixel value representing the image is also referred to as a sample. Furthermore, the stream can be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal. Additionally, the encoding device can also be called an image encoding device or a moving image encoding device, and the encoding method of the encoding device 100 can also be called an encoding method, image encoding method, or moving image encoding method.
[0259] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Additionally, the network Nw can also be replaced by a storage medium that records the stream such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0260] The decoding device 200 generates a decoded image, such as an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method of the encoding device 100.
[0261] Additionally, the decoding device can also be called an image decoding device or a moving image decoding device, and the decoding method of the decoding device 200 can also be called a decoding method, image decoding method, or moving image decoding method.
[0262] [Data Structure]
[0263] Figure 2 is a diagram showing an example of the hierarchical structure of the data in the stream. The stream includes, for example, a video sequence. This video sequence, for example, as shown in (a) of Figure 2 contains a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and a plurality of pictures.
[0264] In a moving image in which a VPS includes a plurality of layers, the encoding parameters common to the plurality of layers and the plurality of layers included in the moving image, or the encoding parameters associated with each layer.
[0265] An SPS includes parameters used for a sequence, that is, encoding parameters that a decoding device 200 refers to for decoding the sequence. For example, the encoding parameters may also represent the width or height of a picture. In addition, there may be multiple SPSs.
[0266] A PPS includes parameters used for a picture, that is, encoding parameters that a decoding device 200 refers to for decoding each picture in the sequence. For example, the encoding parameters may also include a reference value of a quantization width used in decoding the picture and a flag indicating the application of weighted prediction. In addition, there may be multiple PPSs. In addition, SPSs and PPSs are sometimes simply referred to as parameter sets.
[0267] As Figure 2 shown in (b) of , a picture may include a picture header and one or more slices. The picture header includes encoding parameters that a decoding device 200 refers to for decoding the one or more slices.
[0268] As Figure 2 shown in (c) of , a slice includes a slice header and one or more tiles. The slice header includes encoding parameters that a decoding device 200 refers to for decoding the one or more tiles.
[0269] As Figure 2 shown in (d) of , a tile includes one or more CTUs (Coding Tree Units).
[0270] In addition, a picture may not include a slice and include a tile group instead of the slice. In this case, the tile group includes one or more tiles. In addition, a slice may be included in a tile.
[0271] A CTU is also referred to as a super block or a basic segmentation unit. As Figure 2 shown in (e) of , such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes encoding parameters that a decoding device 200 refers to for decoding the one or more CUs.
[0272] A CU may also be divided into a plurality of small CUs. In addition, as Figure 2As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual described later. In addition, the CU is basically the same as the PU (Prediction Unit) and the TU (Transform Unit), but for example, in the SBT described later, it may also include a plurality of TUs smaller than the CU. In addition, the CU can also process each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. The VPDU is, for example, a fixed unit that can be processed in one stage when performing pipeline processing in hardware.
[0273] In addition, the stream may not have Figure 2 any part of the hierarchies shown. In addition, the order of these hierarchies can be swapped, and any one hierarchy can be replaced with another hierarchy. In addition, the picture that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is called the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded, and if the processing is decoding, the current picture is synonymous with the picture to be decoded. In addition, a block such as a CU or a CU that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded, and if the processing is decoding, the current block is synonymous with the block to be decoded.
[0274] [Slice / Tile of Picture Structure]
[0275] In order to decode pictures in parallel, a picture may sometimes be composed of slice units or tile units.
[0276] A slice is the basic encoding unit that constitutes a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs.
[0277] Figure 3This is a diagram showing an example of the structure of slices. For example, the picture contains 11×8 CTUs and is divided into 4 slices (Slice 1 - 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU within the picture belongs to a certain slice. The shape of the slice is the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can also be somewhere among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. Additionally, the slice contains a slice header and encoded data. In the slice header, features of the slice such as the starting CTU address and slice type of the slice can also be described.
[0278] A tile is the unit of the rectangular area that constitutes the picture. Numbers called TileId can also be assigned to each tile in the raster scan order.
[0279] Figure 4 This is a diagram showing an example of the structure of tiles. For example, the picture contains 11×8 CTUs and is divided into 4 rectangular area tiles (Tile 1 - 4). When using tiles, compared with the case of not using tiles, the processing order of CTUs is changed. When not using tiles, multiple CTUs within the picture are processed in the raster scan order, for example. When using tiles, in each of the multiple tiles, at least 1 CTU is processed in the raster scan order, for example. For example, as Figure 4 shown, the processing order of the multiple CTUs contained in Tile 1 is the order from the left end of the first column of Tile 1 towards the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 towards the right end of the second column of Tile 1.
[0280] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.
[0281] Furthermore, the picture can also be composed of tile set units. A tile set can contain more than one tile group and can also contain more than one tile. The picture can be composed of only one of the tile set, tile group, and tile. For example, the order of scanning multiple tiles in the raster order for each tile set is set as the basic encoding order of the tiles. A set of one or more tiles whose basic encoding order is continuous within each tile set is set as a tile group. Such a picture can also be composed of the later-described division unit 102 (refer to Figure 7 ).
[0282] [Scalable Coding]
[0283] Figure 5 and Figure 6This is a diagram showing an example of the structure of a scalable stream.
[0284] As Figure 5 shown, the encoding device 100 can perform encoding by dividing multiple pictures into a certain layer among multiple layers, thereby generating a temporally / spatially scalable stream. For example, the encoding device 100 realizes the scalability where the enhancement layer exists above the base layer by encoding pictures for each layer. The encoding of each such picture is called scalable encoding. Thus, the decoding device 200 can switch the picture quality of the image displayed by decoding this stream. That is, the decoding device 200 determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth. As a result, the decoding device 200 can freely switch the same content between low-resolution content and high-resolution content for decoding. For example, a user of this stream, while on the move, watches a moving image of this stream halfway through using a smartphone, and after returning home, watches the subsequent part of this moving image using a device such as an Internet TV. Additionally, decoding devices 200 with the same or different performances are respectively assembled in the above-mentioned smartphone and device. In this case, if the device decodes to the upper layer in this stream, the user can watch a high-quality moving image after returning home. Thus, the encoding device 100 does not need to generate multiple streams with the same content but different picture qualities, and can reduce the processing load.
[0285] Furthermore, the enhancement layer can also include meta-information such as based on the statistical information of the image. It can also be that the decoding device 200 generates a high-quality moving image by super-resolving the pictures of the base layer based on the meta-information. Super-resolution can be to improve the SNR at the same resolution or to increase the resolution. The meta-information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least squares operation used in the super-resolution process, etc.
[0286] Alternatively, the picture can also be segmented into tiles, etc. according to the meaning of each object, etc. within the picture. In this case, the decoding device 200 can decode only a partial area within the picture by selecting the tile to be decoded. Moreover, the attributes of the object (person, car, ball, etc.) and the position within the picture (coordinate position in the same picture, etc.) can be saved as meta-information. In this case, the decoding device 200 can determine the position of the desired object based on the meta-information and decide the tile containing the object. For example, as Figure 6 shown, a data storage structure different from the pixel data, such as SEI in HEVC, can also be used to store the meta-information. This meta-information represents, for example, the position, size, or color of the main object.
[0287] In addition, meta-information can also be stored in units composed of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding device 200 can obtain the time when a specific person appears in the moving image, etc. By using this time and the information of the picture unit, the picture in which the target exists and the position of the target in that picture can be determined.
[0288] [Encoding device]
[0289] Next, the encoding device 100 of the embodiment will be described. Figure 7 It is a block diagram showing an example of the functional structure of the encoding device 100 of the embodiment. The encoding device 100 encodes an image in units of blocks.
[0290] As Figure 7 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. In addition, the intra prediction unit 124 and the inter prediction unit 126 are respectively configured as part of the prediction processing unit.
[0291] [Installation example of the encoding device]
[0292] Figure 8 It is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 7 as shown, the multiple components of the encoding device 100 are implemented by Figure 8 the processor a1 and the memory a2 shown.
[0293] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes an image. The processor a1 can also be a processor such as a CPU. In addition, the processor a1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 can also play the role of Figure 7 multiple components of the encoding device 100 shown, except for the components used for storing information.
[0294] Memory A2 is a dedicated or general-purpose memory that stores information for the processor A1 to encode images. Memory A2 can be either an electronic circuit or connected to the processor A1. Additionally, Memory A2 can also be included in the processor A1. Moreover, Memory A2 can also be an aggregate of multiple electronic circuits. Additionally, Memory A2 can be a magnetic disk, optical disk, etc., or can be represented as a storage or a recording medium, etc. Additionally, Memory A2 can be either a non-volatile memory or a volatile memory.
[0295] For example, Memory A2 can store the encoded images or can store the stream corresponding to the encoded images. Additionally, a program for the processor A1 to encode images can also be stored in Memory A2.
[0296] Additionally, for example, Memory A2 can also serve as Figure 7 the component for storing information among the multiple components of the encoding device 100 shown. Specifically, Memory A2 can serve as Figure 7 the block memory 118 and the frame memory 122 shown. More specifically, reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.) can be stored in Memory A2.
[0297] Additionally, in the encoding device 100, not all of the multiple components shown in Figure 7 may be installed, and not all of the above-mentioned multiple processes may be performed. Figure 7 A part of the multiple components shown in
[0298] may be included in other devices, or a part of the above-mentioned multiple processes may be performed by other devices.
[0299] [Overall Process of Encoding Processing]
[0300] Fig. 9 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.
[0301] First, the segmentation unit 102 of the encoding device 100 divides the pictures included in the original image into multiple blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation pattern for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block with a fixed size into multiple blocks that make up the selected segmentation pattern. Then, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 for each of the multiple blocks.
[0302] The prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126, and the prediction control unit 128 generate a prediction image of the current block (step Sa_3). In addition, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.
[0303] Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual is also referred to as a prediction error.
[0304] Next, the transformation unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by transforming and quantizing the prediction image (step Sa_5).
[0305] Next, the entropy encoding unit 110 generates a stream by encoding the plurality of quantization coefficients and prediction parameters related to the generation of the prediction image (specifically, entropy encoding) (step Sa_6).
[0306] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore the prediction residual by inverse quantizing and inverse transforming the plurality of quantization coefficients (step Sa_7).
[0307] Next, the addition unit 116 reconstructs the current block by adding the restored prediction residual to the prediction image (step Sa_8). Thereby, a reconstructed image is generated. In addition, the reconstructed image is also referred to as a reconstructed block. In particular, the reconstructed image generated by the encoding device 100 is also referred to as a local decoded block or a local decoded image.
[0308] When generating the reconstructed image, the loop filtering unit 120 filters the reconstructed image as needed (step Sa_9).
[0309] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.
[0310] In addition, in the above example, the encoding device 100 selects one segmentation style for a block of a fixed size and encodes each block according to the segmentation style, but each block may also be encoded according to each of a plurality of segmentation styles. In this case, the encoding device 100 can evaluate the cost for each of the plurality of segmentation styles, and for example, can select the stream obtained by encoding according to the segmentation style with the minimum cost as the finally output stream.
[0311] In addition, the processing of these steps Sa_1 to Sa_10 can be sequentially performed by the encoding device 100, a part of the plurality of processes among these processes can be performed in parallel, or the order can be changed.
[0312] The encoding process of such an encoding device 100 uses hybrid encoding that combines predictive encoding and transform encoding. In addition, the predictive encoding is performed through an encoding loop, which is composed of a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. That is, the prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126 constitutes a part of the encoding loop.
[0313] [Splitting Unit]
[0314] The splitting unit 102 splits each picture included in the original image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128 pixels). Such blocks of the fixed size are sometimes referred to as coding tree units (CTUs). And, the splitting unit 102 splits each block of the fixed size into blocks of variable sizes (e.g., 64×64 pixels or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects a splitting pattern. Such blocks of variable sizes are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). Additionally, in various implementation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture can also be used as the processing unit for CUs, PUs, or TUs.
[0315] Fig.10 is a diagram showing an example of the block splitting in the embodiment. In Fig.10 it, the solid lines represent block boundaries based on quadtree block splitting, and the dashed lines represent block boundaries based on binary tree block splitting.
[0316] Here, the block 10 is a square block of 128×128 pixels. This block 10 is first split into 4 square blocks of 64×64 pixels (quadtree block splitting).
[0317] The upper-left 64×64 pixel square block is further vertically split into 2 rectangular blocks each composed of 32×64 pixels, and the left 32×64 pixel rectangular block is further vertically split into 2 rectangular blocks each composed of 16×64 pixels (binary tree block splitting). As a result, the upper-left 64×64 pixel square block is split into 2 rectangular blocks 11 and 12 of 16×64 pixels and a 32×64 pixel rectangular block 13.
[0318] The upper-right 64×64 pixel square block is horizontally split into 2 rectangular blocks 14 and 15 each composed of 64×32 pixels (binary tree block splitting).
[0319] The 64×64 pixel square block in the lower left is divided into 4 square blocks (quad-tree block division), each consisting of 32×32 pixels. The upper left and lower right blocks among the 4 square blocks, each consisting of 32×32 pixels, are further divided. The upper left 32×32 pixel square block is vertically divided into 2 rectangular blocks, each consisting of 16×32 pixels, and the right rectangular block consisting of 16×32 pixels is further horizontally divided into 2 square blocks, each consisting of 16×16 pixels (binary tree block division). The lower right 32×32 pixel square block is horizontally divided into 2 rectangular blocks, each consisting of 32×16 pixels (binary tree block division). As a result, the 64×64 pixel square block in the lower left is divided into 16 rectangular blocks 16 of 16×32 pixels, 2 square blocks 17 and 18 of 16×16 pixels each, 2 square blocks 19 and 20 of 32×32 pixels each, and 2 rectangular blocks 21 and 22 of 32×16 pixels each.
[0320] The block 23 consisting of 64×64 pixels in the lower right is not divided.
[0321] As described above, in Fig.10 , the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary tree block division. Such division is sometimes called QTBT (quad-tree plus binary tree) division.
[0322] In addition, in Fig.10 , 1 block is divided into 4 or 2 blocks (quad-tree or binary tree block division), but the division is not limited to these. For example, 1 block may be divided into 3 blocks (ternary tree division). Division including such ternary tree division is sometimes called MBT (multi type tree) division.
[0323] Fig.11 is a diagram showing an example of the functional structure of the division unit 102. As Fig.11 shown, the division unit 102 may also include a block division determination unit 102a. As an example, the block division determination unit 102a may perform the following processing.
[0324] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the above division pattern based on this block information. The division unit 102 divides the original image according to this division pattern and outputs one or more blocks obtained by this division to the subtraction unit 104.
[0325] In addition, the block segmentation determination unit 102a outputs, for example, parameters representing the above-described segmentation pattern to the transformation unit 106, the inverse transformation unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy encoding unit 110. The transformation unit 106 can transform the prediction residual based on this parameter, and the intra prediction unit 124 and the inter prediction unit 126 can generate a prediction image based on this parameter. In addition, the entropy encoding unit 110 can also perform entropy encoding on this parameter.
[0326] As an example, parameters related to the segmentation pattern can also be written into the stream as follows.
[0327] Fig.12 FIG. is an example showing a segmentation pattern. Examples of the segmentation pattern include: four-way split (QT), in which a block is split into two in the horizontal and vertical directions respectively; three-way split (HT or VT), in which a block is split in the same direction at a ratio of 1:2:1; two-way split (HB or VB), in which a block is split in the same direction at a ratio of 1:1; and no split (NS).
[0328] In addition, in the case of four-way split and no split, the segmentation pattern does not have a block segmentation direction, and in the case of two-way split and three-way split, the segmentation pattern has segmentation direction information.
[0329] Fig.13A and Fig. 13B FIG. is an example of a syntax tree showing a segmentation pattern. In the example of Fig.13A , first, there is information indicating whether to perform segmentation (S: Split flag), and then there is information indicating whether to perform four-way split (QT: QT flag). Next, there is information indicating whether to perform three-way split or two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the segmentation direction (Ver: Vertical flag or Hor: Horizontal flag). In addition, for each of one or more blocks obtained by such segmentation based on the segmentation pattern, the same process can be further repeatedly applied for segmentation. That is, as an example, it is also possible to recursively perform determination of whether to perform segmentation, whether to perform four-way split, whether the segmentation method is horizontal or vertical, and whether to perform three-way split or two-way split, and encode the determination results implemented into the stream in the encoding order disclosed in the syntax tree shown in Fig.13A .
[0330] In addition, in the syntax tree shown in Fig.13A , these information are arranged in the order of S, QT, TT, Ver, but they can also be arranged in the order of S, QT, Ver, BT. That is, in Fig. 13BIn the example, first, there is information indicating whether to perform splitting (S: Split flag), then there is information indicating whether to perform four-way splitting (QT: QT flag). Next, there is information indicating the splitting direction (Ver: Vertical flag or Hor: Horizontal flag), and finally there is information indicating whether to perform two-way splitting or three-way splitting (BT: BT flag or TT: TT flag).
[0331] In addition, the splitting pattern described here is an example. It is possible to use a splitting pattern other than the described one, or only a part of the described splitting pattern.
[0332] [Subtraction unit]
[0333] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units that are input and split by the splitting unit 102. That is, the subtraction unit 104 calculates the prediction residual of the current block. And the subtraction unit 104 outputs the calculated prediction residual to the transformation unit 106.
[0334] The original image is an input signal of the encoding device 100, for example, a signal representing an image of each picture constituting a moving image (for example, a luma signal and two chroma signals).
[0335] [Transformation unit]
[0336] The transformation unit 106 transforms the prediction residual in the spatial domain into transformation coefficients in the frequency domain and outputs the transformation coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction residual in the spatial domain.
[0337] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types, use a transformation basis function corresponding to the selected transformation type, and transform the prediction residual into transformation coefficients. Such a transformation is called EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform) in some cases. In addition, the transformation basis function is sometimes simply referred to as the basis.
[0338] Multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. In addition, these transformation types can be described as DCT2, DCT5, DCT8, DST1, and DST7 respectively. Fig.14is a table representing transform basis functions corresponding to respective transform types. In Fig.14 , N represents the number of input pixels. The selection of a transform type from among these multiple transform types can depend, for example, on the type of prediction (intra prediction, inter prediction, etc.) or on the intra prediction mode.
[0339] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. Additionally, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0340] Furthermore, the transform unit 106 may also perform a re - transform on the transform coefficients (i.e., the transform result). Such a re - transform may be a case of what is called AST (adaptive secondary transform) or NSST (non - separable secondary transform). For example, the transform unit 106 performs a re - transform for each sub - block (e.g., a 4×4 pixel sub - block) included in a block of transform coefficients corresponding to the intra prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are generally signaled at the CU level. Additionally, the signaling of this information need not be limited to the CU level and may also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0341] In the transform unit 106, a separable transform and a non - separable transform may also be applied. A separable transform is a method of performing multiple transforms separately for each direction according to the number of input dimensions, and a non - separable transform is a method of treating two or more dimensions as one dimension and performing a transform together when the input is multi - dimensional.
[0342] For example, as an example of a non - separable transform, when the input is a 4×4 pixel block, it can be regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.
[0343] Furthermore, in a further example of a non - separable transform, after regarding a 4×4 pixel input block as a permutation having 16 elements, a transform of performing multiple Givens rotations on this permutation (Hypercube Givens Transform) may also be performed.
[0344] In the transformation in the transformation unit 106, the type of transformation of the transformation basis function to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform).
[0345] Fig.15 It is a diagram showing an example of SVT.
[0346] In SVT, as Fig.15 shown, the CU is bisected in the horizontal or vertical direction, and only one of the arbitrary regions is transformed into the frequency domain. The transformation type can be set for each region. For example, DST7 and DCT8 are used. For example, for the region at position 0 among the two regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used. Or, for the region at position 1 among the two regions, DST7 is used. Similarly, for the region at position 0 among the two regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 are used. Or, for the region at position 1 among the two regions, DST7 is used. In such an Fig.15 example shown, only one of the two regions within the CU is transformed, and the other is not transformed, but it is also possible to transform the two regions separately. In addition, the splitting method is not limited to bisection, and it can also be quartering. Furthermore, it can be more flexible, encoding the information indicating the splitting method and performing signaling in the same way as CU splitting. In addition, SVT is sometimes also referred to as SBT (Sub-block Transform).
[0347] The aforementioned AMT and EMT may also be referred to as MTS (Multiple Transform Selection). In the case of applying MTS, transform types such as DST7 or DCT8 can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. On the other hand, as a process of selecting the transform type used in the orthogonal transform based on the shape of the CU without encoding the index information, there is a process called IMTS (Implicit MTS). In the case of applying IMTS, for example, if the shape of the CU is rectangular, DST7 is used on the short side of the rectangle and DCT2 is used on the long side, and orthogonal transforms are performed respectively. Additionally, for example, in the case where the shape of the CU is square, if MTS is effective in the sequence, DCT2 is used for orthogonal transform, and if MTS is ineffective, DST7 is used for orthogonal transform. DCT2 and DST7 are just examples, and other transform types can be used, or different combinations of the used transform types can be set. IMTS can be used only in blocks for intra prediction, or can be used together in blocks for intra prediction and blocks for inter prediction.
[0348] As described above, as a selection process for selectively switching the transform type used in the orthogonal transform, three processes, namely MTS, SBT, and IMTS, have been described. However, all three selection processes can be effective, or only some of the selection processes can be selectively made effective. Regarding whether each selection process is effective, it can be identified by flag information in headers such as SPS. For example, if all three selection processes are effective, one is selected from the three selection processes for orthogonal transform in units of CU. Additionally, as long as the selection process for selectively switching the transform type can achieve at least one of the following four functions [1] to [4], a selection process different from the above three selection processes can be used, or the above three selection processes can be replaced with other processes respectively. Function [1] is the function of performing orthogonal transform on the entire range within the CU and encoding the information indicating the transform type used in the transform. Function [2] is the function of performing orthogonal transform on the entire range of the CU, determining the transform type based on a predetermined rule without encoding the information indicating the transform type. Function [3] is the function of performing orthogonal transform on a partial region of the CU and encoding the information indicating the transform type used in the transform. Function [4] is the function of performing orthogonal transform on a partial region of the CU, determining the transform type based on a predetermined rule without encoding the information indicating the transform type used in the transform, and so on.
[0349] In addition, the presence or absence of the application of MTS, IMTS, and SBT respectively can also be determined for each processing unit. For example, it can be determined for each sequence unit, picture unit, tile unit, slice unit, CTU unit, or CU unit.
[0350] In addition, the tool for selectively switching the transformation type in the present invention can also be renamed as a method for adaptively selecting the basis used in the transformation process, a selection process, or a process for selecting a basis. In addition, the tool for selectively switching the transformation type can also be renamed as a mode for adaptively selecting the transformation type.
[0351] Fig.16 It is a flowchart showing an example of the process performed by the transformation unit 106.
[0352] For example, the transformation unit 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined to perform an orthogonal transformation (Yes in step St_1), the transformation unit 106 selects the transformation type used for the orthogonal transformation from among multiple transformation types (step St_2). Next, the transformation unit 106 performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Then, the transformation unit 106 outputs the information indicating the selected transformation type to the entropy encoding unit 110 to encode this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transformation (No in step St_1), the transformation unit 106 outputs the information indicating that no orthogonal transformation is performed to the entropy encoding unit 110 to encode this information (step St_5). In addition, the determination of whether to perform an orthogonal transformation in step St_1 can be made, for example, based on the size of the transformation block, the prediction mode applied to the CU, etc. Also, the information indicating the transformation type used for the orthogonal transformation may not be encoded, and an orthogonal transformation may be performed using a pre-specified transformation type.
[0353] Fig.17 It is a flowchart showing another example of the process performed by the transformation unit 106. In addition, Fig.17 The example shown is the same as the example shown in Fig.16 and is an example of an orthogonal transformation in the case of applying a method for selectively switching the transformation type used for the orthogonal transformation.
[0354] As an example, the first transformation type group may include DCT2, DST7, and DCT8. In addition, as an example, the second transformation type group may include DCT2. Also, the transformation types included in the first transformation type group and the second transformation type group may partially overlap or may be all different transformation types.
[0355] Specifically, the transformation unit 106 determines whether the transformation size is equal to or less than a specified value (step Su_1). Here, when it is determined that the size is equal to or less than the specified value (yes in step Su_1), the transformation unit 106 orthogonally transforms the prediction residual of the current block using the transformation types included in the first transformation type group (step Su_2). Then, the transformation unit 106 encodes the information indicating which transformation type among one or more transformation types included in the first transformation type group is used by outputting the information to the entropy encoding unit 110 (step Su_3). On the other hand, when it is determined that the transformation size is not equal to or less than the specified value (no in step Su_1), the transformation unit 106 orthogonally transforms the prediction residual of the current block using the second transformation type group (step Su_4).
[0356] In step Su_3, the information indicating the transformation type used for the orthogonal transformation may be information representing a combination of the transformation type applied to the vertical direction of the current block and the transformation type applied to the horizontal direction. Additionally, the first transformation type group may include only one transformation type, and the information indicating the transformation type used for the orthogonal transformation may not be encoded. The second transformation type group may include multiple transformation types, and the information indicating the transformation type used in the orthogonal transformation among one or more transformation types included in the second transformation type group may also be encoded.
[0357] Alternatively, the transformation type may be determined based only on the transformation size. In addition, if the process of determining the transformation type used for the orthogonal transformation is based on the transformation size, it is not limited to the determination of whether the transformation size is equal to or less than the specified value.
[0358] [Quantization unit]
[0359] The quantization unit 108 quantizes the transformation coefficients output from the transformation unit 106. Specifically, the quantization unit 108 scans the multiple transformation coefficients of the current block in a specified scan order and quantizes the transformation coefficients based on the quantization parameter (QP) corresponding to the scanned transformation coefficients. Then, the quantization unit 108 outputs the quantized multiple transformation coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0360] The specified scan order is the order used for quantization / inverse quantization of the transformation coefficients. For example, the specified scan order is defined by the ascending order of frequencies (from low frequency to high frequency) or the descending order (from high frequency to low frequency).
[0361] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the error of the quantization coefficient (quantization error) increases.
[0362] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transformation sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing values sampled at a predetermined interval by associating them with predetermined levels, and in this technical field, expressions such as rounding, truncation, or scaling are sometimes used.
[0363] As methods of using the quantization matrix, there are a method of using the quantization matrix directly set on the encoding device 100 side and a method of using the default quantization matrix (default matrix). On the encoding device 100 side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix. In addition, instead of directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix used in the quantization of the current block can be generated based on the default quantization matrix or the encoded quantization matrix.
[0364] On the other hand, there is also a method of performing quantization in such a way that the coefficients of the high-frequency components and the coefficients of the low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix (flat matrix) in which all the coefficients are the same value.
[0365] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level.
[0366] When using the quantization matrix, the quantization unit 108, for example, scales the quantization width and the like obtained according to quantization parameters and the like for each transform coefficient using the values of the quantization matrix. The quantization process without using the quantization matrix can also be a process of quantizing the transform coefficients based on the quantization width obtained according to quantization parameters and the like. In addition, in the quantization process without using the quantization matrix, a predetermined value common to all the transform coefficients in the block can also be multiplied by the quantization width.
[0367] Fig.18 It is a block diagram showing an example of the functional structure of the quantization unit 108.
[0368] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108a, a prediction quantization parameter generation unit 108b, a quantization parameter generation unit 108c, a quantization parameter storage unit 108d, and a quantization processing unit 108e.
[0369] Fig.19 It is a flowchart showing an example of the quantization performed by the quantization unit 108.
[0370] As an example, the quantization unit 108 can be based on Fig.19 The flowchart shown performs quantization for each CU. Specifically, the quantization parameter generation unit 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generation unit 108c generates quantization parameters for the current block (step Sv_2) and saves the quantization parameters to the quantization parameter storage unit 108d (step Sv_3).
[0371] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). Then, the predicted quantization parameter generation unit 108b obtains quantization parameters of a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). The predicted quantization parameter generation unit 108b generates predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generation unit 108c and the predicted quantization parameters of the current block generated by the predicted quantization parameter generation unit 108b (step Sv_7). By calculating this difference, differential quantization parameters are generated. The differential quantization parameter generation unit 108a outputs the differential quantization parameters to the entropy encoding unit 110, thereby encoding the differential quantization parameters (step Sv_8).
[0372] In addition, the differential quantization parameters can also be encoded at the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameters can be generated using the initial values of the quantization parameters and the differential quantization parameters.
[0373] In addition, the quantization unit 108 can include multiple quantizers, and dependent quantization that quantizes the transform coefficients using a quantization method selected from multiple quantization methods can also be applied.
[0374] [Entropy Encoding Unit]
[0375] Fig. 20 is a block diagram showing an example of the functional structure of the entropy encoding unit 110.
[0376] The entropy encoding unit 110 performs entropy encoding on the quantized coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130, thereby generating a stream. In this entropy encoding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy encoding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization that transforms multi-value signals such as quantized coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, etc. The context control unit 110b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. In the method for deriving this context value, for example, there are bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic coding unit 110c performs arithmetic coding on the binarized signal using the derived context value.
[0377] Fig.21 It is a diagram showing the process of CABAC in the entropy encoding unit 110.
[0378] First, in the CABAC in the entropy encoding unit 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding unit 110c and setting of the initial context value are performed. Then, the binarization unit 110a and the binary arithmetic coding unit 110c sequentially perform binarization and arithmetic coding on the multiple quantized coefficients of the CTU, respectively. At this time, the context control unit 110b updates the context value each time arithmetic coding is performed. Then, the context control unit 110b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.
[0379] [Inverse Quantization Unit]
[0380] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scan order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0381] [Inverse Transform Unit]
[0382] The inverse transform unit 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.
[0383] In addition, since the restored prediction residual usually loses information through quantization, it is inconsistent with the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual usually contains a quantization error.
[0384] [Addition unit]
[0385] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 to the predicted image input from the prediction control unit 128. As a result, a reconstructed image is generated. Then, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.
[0386] [Block memory]
[0387] The block memory 118 is, for example, a storage unit for storing blocks referred to in intra prediction and that are blocks within the current picture. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.
[0388] [Frame memory]
[0389] The frame memory 122 is, for example, a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.
[0390] [Loop filter unit]
[0391] The loop filter unit 120 performs a loop filter process on the reconstructed image output from the addition unit 116 and outputs the reconstructed image after the filter process to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).
[0392] Fig. 22 It is a block diagram showing an example of the functional structure of the loop filter unit 120.
[0393] For example Fig. 22As shown in the figure, the loop filter unit 120 includes a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 120b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 120c applies the above-described ALF processing to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is a process for improving the image quality by reducing ringing (a phenomenon in which pixel values around the edge are deformed in a fluctuating manner) and correcting the deviation of pixel values. In this SAO processing, for example, there are edge offset processing and band offset processing. In addition, the loop filter unit 120 may not include Fig. 22 all of the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 120 may also be structured to perform the above-described respective processes in an order different from the processing order disclosed in Fig. 22 .
[0394] [Loop Filter Unit > Adaptive Loop Filter]
[0395] In the ALF, a least squares error filter used to remove coding distortion is adopted. For example, for each 2×2 pixel sub-block within the current block, one filter selected from multiple filters based on the direction and activity of the locality-based gradient is adopted.
[0396] Specifically, first, sub-blocks (for example, 2×2 pixel sub-blocks) are classified into multiple classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed, for example, based on the direction and activity of the gradient. In a specific example, using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4), the classification value C (for example, C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into multiple classes.
[0397] The gradient direction value D is derived, for example, by comparing the gradients in multiple directions (for example, horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in multiple directions and quantifying the added result.
[0398] Based on the result of such classification, the filter to be used for the sub-block is determined from among multiple filters.
[0399] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 23A to 23C is a diagram showing multiple examples of the shape of the filter used in the ALF. Fig.23A represents a 5×5 diamond-shaped filter, Fig. 23B represents a 7×7 diamond-shaped filter, Fig.23C represents a 9×9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0400] The on / off of ALF can also be determined, for example, at the picture level or CU level. For example, for luminance, it can be determined at the CU level whether to employ ALF, and for chrominance difference, it can be determined at the picture level whether to employ ALF. Information indicating the on / off of ALF is usually signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of ALF need not be limited to the picture level or CU level and may also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0401] Additionally, as described above, one filter is selected from multiple filters and ALF processing is applied to the sub-block. For each of the multiple filters (e.g., up to 15 or 25 filters), the set of coefficients composed of the multiple coefficients used in the filter is usually signaled at the picture level. Additionally, the signaling of the set of coefficients need not be limited to the picture level and may also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0402] [Loop Filter>Cross Component Adaptive Loop Filter (or Cross Component Adaptive Loop Filter)]
[0403] Fig.23D is a diagram showing an example where the Y sample (first component) is used for CCALF of Cb and CCALF of Cr (multiple components different from the first component). Fig.23E is a diagram showing a diamond-shaped filter.
[0404] One example of CC-ALF is by using a linear diamond-shaped filter ( Fig.23D , Fig.23E) It operates on the luminance channel of each color difference component. For example, filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled to have variable block sizes and is notified by context-encoded flags received for each block of samples. The block size and the CC-ALF enable flag are received at the slice level of each color difference component. The syntax and semantics of CC-ALF are provided in the Appendix. In this document, block sizes of 16x16, 32x32, 64x64, and 128x128 are supported (in color difference samples).
[0405] [Loop Filter> Joint Chroma Cross Component Adaptive Loop Filter]
[0406] Fig.23F It is a diagram showing an example of JC-CCALF. Figure 23G It is a diagram showing an example of the weight_index candidates of JC-CCALF.
[0407] One example of JC-CCALF uses only one CCALF filter, generates one CCALF filter output as the color difference adjustment signal for only one color component, and applies a properly weighted version of the same color difference adjustment signal to other color components. In this way, the complexity of the existing CCALF is approximately halved.
[0408] The weight value is encoded into a sign flag and a weight index. The weight index (denoted as weight_index) is encoded as 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight. It cannot be the same as 0. The magnitude of JcCcWeight is determined as follows.
[0409] · When weight_index is 4 or less, JcCcWeight is equal to weight_index >> 2.
[0410] · In other cases, JcCcWeight is equal to 4 / (weight_index - 4).
[0411] The on / off control at the block level for ALF filtering of Cb and Cr is separate. This is the same as CCALF, and two separate sets of on / off control flags at the block level are encoded. Here, different from CCALF, the on / off control block sizes for Cb and Cr are the same, so only one block size variable is encoded.
[0412] [Loop Filter Section > Deblocking Filter]
[0413] In the deblocking filter process, the loop filter section 120 reduces the distortion generated at the block boundary by filtering the block boundary of the reconstructed image.
[0414] Fig.24 It is a block diagram showing an example of the detailed structure of the deblocking filter processing section 120a.
[0415] The deblocking filter processing section 120a includes, for example, a boundary determination section 1201, a filtering determination section 1203, a filtering processing section 1205, a processing determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.
[0416] The boundary determination section 1201 determines whether there are pixels (i.e., target pixels) for which deblocking filter processing is to be performed near the block boundary. Then, the boundary determination section 1201 outputs its determination result to the switch 1202 and the processing determination section 1208.
[0417] In the case where the boundary determination section 1201 determines that target pixels exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1204. On the contrary, when the boundary determination section 1201 determines that target pixels do not exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1206. In addition, the image before filtering processing is an image composed of the target pixel and at least one surrounding pixel located around the target pixel.
[0418] The filtering determination section 1203 determines whether to perform deblocking filter processing on the target pixel based on the pixel values of at least one surrounding pixel located around the target pixel. Then, the filtering determination section 1203 outputs the determination result to the switch 1204 and the processing determination section 1208.
[0419] In the case where the filtering determination section 1203 determines that deblocking filter processing is to be performed on the target pixel, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the filtering processing section 1205. On the contrary, in the case where the filtering determination section 1203 determines that deblocking filter processing is not to be performed on the target pixel, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the switch 1206.
[0420] In the case where the image before filtering processing is obtained via the switches 1202 and 1204, the filtering processing section 1205 performs deblocking filter processing with the filtering characteristics determined by the filtering characteristic determination section 1207 on the target pixel. Then, the filtering processing section 1205 outputs the pixel after the filtering processing to the switch 1206.
[0421] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs pixels that have not been deblocking-filtered and pixels that have been deblocking-filtered by the filtering processing unit 1205.
[0422] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processing determination unit 1208 outputs the deblocking-filtered pixels from the switch 1206. In addition, in other cases than the above, the processing determination unit 1208 outputs the pixels that have not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206. In addition, Fig.24 The structure shown is an example of the structure in the deblocking filtering unit 120a, and the deblocking filtering unit 120a may have other structures.
[0423] Fig.25 is a diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to the block boundary.
[0424] In deblocking filtering, for example, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Fig.25 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.
[0425] q’0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8
[0426] q’1 = (p0 + q0 + q1 + q2 + 2) / 4
[0427] q’2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8
[0428] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering are filtering coefficients.
[0429] Furthermore, in the deblocking filter processing, the clipping process may also be performed in such a way that the pixel value after the operation does not change when it exceeds the threshold. In this clipping process, the threshold determined according to the quantization parameter is used to clip the pixel value after the operation based on the above formula to "the pixel value before the operation ± 2 × the threshold". Thereby, excessive smoothing can be prevented.
[0430] Fig.26 FIG. is an example of a block boundary for explaining the deblocking filter processing. Fig. 27 FIG. is an example showing the BS value.
[0431] The block boundary for the deblocking filter processing is, for example, Fig.26 the boundary of the CU, PU, or TU of an 8×8 pixel block shown in FIG. The deblocking filter processing is performed, for example, in units of 4 rows or 4 columns. First, for Fig.26 the blocks P and Q shown in FIG., the Bs (Boundary Strength) value is determined as Fig. 27 shown.
[0432] According to Fig. 27 the Bs value, it can be determined whether to perform deblocking filter processing with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filter processing for the chrominance signal is performed. When the Bs value is 1 or more and satisfies a specified condition, deblocking filter processing for the luminance signal is performed. In addition, the determination condition of the Bs value is not limited to Fig. 27 the condition shown in FIG., and it can also be determined based on other parameters.
[0433] [Prediction unit (intra prediction unit / inter prediction unit / prediction control unit)]
[0434] Fig.28 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the encoding apparatus 100. In addition, as an example, the prediction unit is composed of all or a part of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.
[0435] The prediction unit generates a prediction image of the current block (step Sb_1). In addition, in the prediction image, there are, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). Specifically, the prediction unit uses the reconstructed image that has already been obtained by generating a prediction image of other blocks, generating a prediction residual, generating quantization coefficients, restoring the prediction residual, and adding the prediction images, and generates a prediction image of the current block.
[0436] The reconstructed image can be, for example, an image that refers to a reference picture, or an image that includes the current block, i.e., an encoded block within the current picture (i.e., the other blocks described above). The encoded blocks within the current picture are, for example, adjacent blocks of the current block.
[0437] Fig.29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding apparatus 100.
[0438] The prediction unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such prediction methods, the above-described reconstructed image can also be used.
[0439] Next, the prediction unit evaluates the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit evaluates these prediction images by calculating a cost C for the prediction images generated in each of steps Sc_1a, Sc_1b, and Sc_1c, and comparing the costs C of these prediction images. In addition, the cost C is calculated by an equation of an R-D optimization model, such as C = D + λ × R. In this equation, D is the encoding distortion of the prediction image, and is represented, for example, by the sum of the absolute values of the differences between the pixel values of the current block and the pixel values of the prediction image. Further, R is the bit rate of the stream. In addition, λ is, for example, the undetermined multiplier of Lagrange.
[0440] Next, the prediction unit selects one of the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects the method or mode for obtaining the final prediction image. For example, the prediction unit selects the prediction image with the minimum cost C based on the cost C calculated for these prediction images. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 can also be performed based on the parameters used in the encoding process. The encoding apparatus 100 can signal an information signal for determining the selected prediction image, method, or mode as a stream. This information can be, for example, a flag or the like. Thereby, the decoding apparatus 200 can generate a prediction image based on this information in accordance with the method or mode selected in the encoding apparatus 100. In addition, in the Fig.29 example shown, after generating the prediction images by each method, the prediction unit selects any one of the prediction images. However, before generating these prediction images, the prediction unit can select a method or mode based on the parameters used for the above encoding process, and can generate a prediction image according to this method or mode.
[0441] For example, the first mode and the second mode are intra prediction and inter prediction respectively, and the prediction unit can select the final prediction image for the current block from the prediction images generated according to these prediction modes.
[0442] Fig.30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device 100.
[0443] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a), and generates a prediction image by inter prediction (step Sd_1b). In addition, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.
[0444] Next, the prediction unit evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above-mentioned cost C can also be used in this evaluation. Then, the prediction unit can select the prediction image that calculates the minimum cost C from the intra prediction image and the inter prediction image as the final prediction image for the current block (step Sd_3). That is, the prediction mode or pattern used to generate the prediction image for the current block is selected.
[0445] [Intra Prediction Unit]
[0446] The intra prediction unit 124 performs intra prediction (also referred to as intra-picture prediction) of the current block with reference to the block within the current picture stored in the block memory 118, thereby generating a prediction image of the current block (i.e., an intra prediction image). Specifically, the intra prediction unit 124 generates an intra prediction image by performing intra prediction with reference to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control unit 128.
[0447] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0448] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC standard.
[0449] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may also include 32-direction prediction modes (a total of 65 directional prediction modes) in addition to the 33 directions. Fig.31This is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent 33 directions specified by the H.265 / HEVC standard, and the dashed arrows represent 32 additional directions (the 2 non-directional prediction modes are not shown in Fig.31 .
[0450] In various installation examples, in the intra prediction of chrominance blocks, the luminance block can also be referred to. That is, the chrominance component of the current block can also be predicted based on the luminance component of the current block. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block in this way (for example, called the CCLM mode) can also be added as one of the intra prediction modes of the chrominance block.
[0451] The intra prediction unit 124 can also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes called PDPC (position dependent intraprediction combination). The information indicating whether PDPC is used (for example, called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0452] Fig.32 This is a flowchart showing an example of the processing performed by the intra prediction unit 124.
[0453] The intra prediction unit 124 selects one intra prediction mode from multiple intra prediction modes (step Sw_1). Then, the intra prediction unit 124 generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM consists of, for example, 6 intra prediction modes. Two of the 6 intra prediction modes can be the Planar prediction mode and the DC prediction mode, and the remaining 4 modes can be directional prediction modes. Then, the intra prediction unit 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0454] Here, when it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). In addition, the MPM flag set to 1 and the information indicating this intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110.
[0455] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction unit 124 does not set the MPM flag. Then, the intra prediction unit 124 generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). In addition, the MPM flag set to 0 and the information indicating this intra prediction mode are respectively encoded as prediction parameters by the entropy encoding unit 110. The information indicating this intra prediction mode represents, for example, any value from 0 to 60.
[0456] [Inter - prediction unit]
[0457] The inter - prediction unit 126 performs inter - prediction (also called inter - picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a predicted image (inter - prediction image). The inter - prediction is performed in units of the current block or a current sub - block within the current block. A sub - block is included in a block and is a unit smaller than the block. The size of the sub - block can be 4x4 pixels, can be 8x8 pixels, or can be other sizes. The size of the sub - block can also be switched in units of slices, bricks, or pictures, etc.
[0458] For example, the inter - prediction unit 126 performs motion search (motion estimation) within the reference picture for the current block or current sub - block to find the reference block or sub - block that most closely matches the current block or current sub - block. And, the inter - prediction unit 126 obtains motion information (such as a motion vector) that compensates for the motion or change from the reference block or sub - block to the current block or sub - block. The inter - prediction unit 126 performs motion compensation (or motion prediction) based on this motion information, thereby generating an inter - prediction image of the current block or sub - block. And, the inter - prediction unit 126 outputs the generated inter - prediction image to the prediction control unit 128.
[0459] The motion information used in motion compensation is signaled as an inter - prediction image in various forms. For example, a motion vector can also be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.
[0460] [List of reference pictures]
[0461] Fig.33 is a diagram showing an example of each reference picture, Fig.34 is a conceptual diagram showing an example of the list of reference pictures. The list of reference pictures is a list showing one or more reference pictures stored in the frame memory 122. In addition, in Fig.33 , the rectangle represents a picture, the arrow represents the reference relationship between pictures, the horizontal axis represents time, I, P, and B in the rectangle represent intra-frame predicted pictures, single predicted pictures, and bi-predicted pictures respectively, and the numbers in the rectangle represent the decoding order. As Fig.33 shown, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As Fig.34 shown, the list of reference pictures is a list showing candidates for reference pictures. For example, one picture (or slice) can have more than one list of reference pictures. For example, if the current picture is a single predicted picture, one list of reference pictures is used, and if the current picture is a bi-predicted picture, two lists of reference pictures are used. In the Fig.33 and Fig.34 example, picture B3 as the current picture currPic has two lists of reference pictures, namely L0 list and L1 list. When the current picture currPic is picture B3, the candidates for the reference pictures of this current picture currPic are I0, P1, and B2, and each list of reference pictures (i.e., L0 list and L1 list) represents these pictures. The inter-frame prediction unit 126 or the prediction control unit 128 specifies whether to actually refer to which picture in each list of reference pictures through the reference picture index refidxLx. In Fig.34 , the reference pictures P1 and B2 are specified through the reference picture indices refIdxL0 and refIdxL1.
[0462] Such a list of reference pictures can be generated in units of sequence, picture, slice, tile, CTU, or CU. In addition, the reference picture indices of the reference pictures shown in the list of reference pictures that are referred to in inter-frame prediction can be encoded at the sequence level, picture level, slice level, tile level, CTU level, or CU level. In addition, in multiple inter-frame prediction modes, a common list of reference pictures can also be used.
[0463] [Basic process of inter-frame prediction]
[0464] Fig.35 is a flowchart showing the basic process of inter-frame prediction.
[0465] The inter-frame prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).
[0466] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter-frame prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. In addition, in the candidate MV list, the previously derived MV can be added as a candidate MV. Further, in the derivation of the MV, the inter-frame prediction unit 126 may also determine the MV of the current block by further selecting at least one candidate MV from at least one candidate MV and determining the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching for a region of the reference picture indicated by each of the selected at least one candidate MVs. In addition, the action of searching for the region of the reference picture may also be referred to as motion estimation.
[0467] In addition, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126. However, for example, the processing of step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.
[0468] In addition, a candidate MV list may be created for each process in each inter-frame prediction mode, or a common candidate MV list may be used in multiple inter-frame prediction modes. In addition, the processes of steps Se_3 and Se_4 respectively correspond to Fig. 9 the processes of steps Sa_3 and Sa_4 shown. In addition, the process of step Se_3 corresponds to Fig.30 the process of step Sd_1b.
[0469] [Flow of MV Derivation]
[0470] Fig.36 is a flowchart showing an example of MV derivation.
[0471] The inter-frame prediction unit 126 may derive the MV of the current block in a mode of encoding motion information (such as MV). In this case, for example, the motion information may be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.
[0472] Alternatively, the inter prediction unit 126 may derive an MV in a mode where motion information is not encoded. In this case, the motion information is not included in the stream.
[0473] Here, the modes of MV derivation include the ordinary inter mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes in which motion information is encoded include the ordinary inter mode, the ordinary merge mode, and the affine mode (specifically, the affine inter mode and the affine merge mode), etc. In addition, the motion information may include not only the MV, but also the predicted MV selection information described later. In addition, the mode in which motion information is not encoded includes the FRUC mode, etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0474] Fig.37 is a flowchart showing another example of MV derivation.
[0475] The inter prediction unit 126 may derive the MV of the current block in a mode where the differential MV is encoded. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. The differential MV is the difference between the MV of the current block and its predicted MV. In addition, the predicted MV is a predicted motion vector.
[0476] Alternatively, the inter prediction unit 126 may derive an MV in a mode where the differential MV is not encoded. In this case, the encoded differential MV is not included in the stream.
[0477] Here, as described above, the modes of MV derivation include the ordinary inter mode, the ordinary merge mode, the FRUC mode, the affine mode, etc. Among these modes, the modes in which the differential MV is encoded include the ordinary inter mode and the affine mode (specifically, the affine inter mode), etc. In addition, the modes in which the differential MV is not encoded include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode), etc. The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0478] [Modes of MV Derivation]
[0479] Fig.38A and Fig.38B is a diagram showing an example of the classification of each mode of MV derivation. For example, as Fig.38AAs shown, according to whether motion information is encoded and whether differential MVs are encoded, the MV derivation modes are classified into three major modes. The three modes are the inter-frame mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter-frame mode is the mode in which motion search is performed, and it is the mode in which motion information and differential MVs are encoded. For example, as Fig.38B shown, the inter-frame mode includes the affine inter-frame mode and the normal inter-frame mode. The merge mode is the mode in which motion search is not performed, and it is the mode in which an MV is selected from the surrounding encoded blocks and used to derive the MV of the current block. This merge mode is basically the mode in which motion information is encoded and differential MVs are not encoded. For example, as Fig.38B shown, the merge mode includes the normal merge mode (sometimes also referred to as the usual merge mode or the regular merge mode), the MMVD (Merge with Motion Vector Difference) mode, the CIIP (Combined inter merge / intraprediction) mode, the triangle mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the various modes included in the merge mode, differential MVs are encoded exceptionarily. In addition, the above-mentioned affine merge mode and affine inter-frame mode are the modes included in the affine mode. The affine mode is the mode in which an affine transformation is assumed and the MVs of the respective sub-blocks constituting the current block are derived as the MV of the current block. The FRUC mode is the mode in which the MV of the current block is derived by searching between encoded regions, and it is the mode in which neither motion information nor differential MVs are encoded. In addition, the details of these respective modes will be described later.
[0480] In addition, Fig.38A and Fig.38B the classification of the respective modes shown is an example and is not limited thereto. For example, when differential MVs are encoded in the CIIP mode, this CIIP mode is classified as the inter-frame mode.
[0481] [MV Derivation > Normal Inter-Frame Mode]
[0482] The normal inter-frame mode is an inter-frame prediction mode in which the MV of the current block is derived by finding a block similar to the image of the current block from the region of the reference picture represented by the candidate MVs. In addition, in this normal inter-frame mode, differential MVs are encoded.
[0483] Fig.39 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.
[0484] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0485] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction MV candidates respectively in a predetermined priority order (step Sg_2). In addition, this priority order is predetermined for each of the N candidate MVs.
[0486] Next, the inter-frame prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the selected prediction MV into the stream. That is, the inter-frame prediction unit 126 outputs the prediction MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0487] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.
[0488] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are executed for each block. For example, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. In addition, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. Furthermore, it may be that when the processes of steps Sg_1 to Sg_5 are not executed for all the blocks included in a slice but for some blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, it may be that when the processes of steps Sg_1 to Sg_5 are executed for some blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends.
[0489] In addition, the predicted image is the inter-frame prediction signal described above. Further, information indicating an inter-frame prediction mode (in the above example, the normal inter-frame mode) used in the generation of the predicted image included in the encoded signal is encoded as, for example, a prediction parameter.
[0490] In addition, the candidate MV list can also be commonly used with lists used in other modes. Further, processing related to the candidate MV list can be applied to processing related to lists used in other modes. Processing related to this candidate MV list is, for example, extracting or selecting a candidate MV from the candidate MV list, rearranging the candidate MVs, or deleting a candidate MV, etc.
[0491] [MV derivation > Normal merge mode]
[0492] The normal merge mode is an inter-frame prediction mode in which a candidate MV is selected from the candidate MV list as the MV of the current block to derive the MV. In addition, the normal merge mode is a narrow sense of the merge mode and is sometimes simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.
[0493] Fig.40 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode.
[0494] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of encoded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0495] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream. That is, the inter-frame prediction unit 126 outputs the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0496] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processes of steps Sh_1 to Sh_3 are performed on each block, for example. For example, when the processes of steps Sh_1 to Sh_3 are performed separately on all the blocks included in a slice, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_1 to Sh_3 are performed separately on all the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends. In addition, the processes of steps Sh_1 to Sh_3 may be such that when they are not performed on all the blocks included in a slice but on a part of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processes of steps Sh_1 to Sh_3 may be such that when they are performed on a part of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.
[0497] In addition, information indicating the inter-frame prediction mode (in the above example, the normal merge mode) used in the generation of the predicted image included in the stream is encoded as, for example, prediction parameters.
[0498] Fig.41 is a diagram for explaining an example of the MV derivation process of the current picture based on the normal merge mode.
[0499] First, the inter-frame prediction unit 126 generates a candidate MV list in which candidate MVs are registered. As candidate MVs, there are: spatially adjacent candidate MVs, which are MVs possessed by a plurality of encoded blocks located in the spatial vicinity of the current block; temporally adjacent candidate MVs, which are MVs possessed by blocks in the vicinity where the position of the current block in the encoded reference picture is projected; combined candidate MVs, which are MVs generated by combining the MV values of the spatially adjacent candidate MVs and the temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero, etc.
[0500] Next, the inter-frame prediction unit 126 determines one candidate MV as the MV of the current block by selecting one candidate MV from among the plurality of candidate MVs registered in the candidate MV list.
[0501] Moreover, in the entropy encoding unit 110, a signal indicating which candidate MV is selected, i.e., merge_idx, is described in the stream and encoded.
[0502] In addition, Fig.41 the candidate MVs registered in the candidate MV list described in [] are an example, and the number may be different from that in the figure, or the structure may not include some of the candidate MVs in the figure, or the structure may be one to which candidate MVs other than the types of candidate MVs in the figure are added.
[0503] The MV of the current block exported through the normal merge mode can also be used to perform the subsequent DMVR (dynamic motion vector refreshing) to determine the final MV. In addition, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same way as the normal merge mode, but encodes the differential MV. As Fig.38B shown, such an MMVD can also be classified as a merge mode together with the normal merge mode. In addition, the differential MV in the MMVD mode may not be the same as the differential MV used in the inter-frame mode. For example, the derivation of the differential MV in the MMVD mode may also be a process with a smaller processing amount than the derivation of the differential MV in the inter-frame mode.
[0504] In addition, the predicted image generated in the inter-frame prediction can be made to coincide with the predicted image generated in the intra-frame prediction to perform the CIIP (Combined inter merge / intra prediction) mode for generating the predicted image of the current block.
[0505] In addition, the candidate MV list may also be referred to as a candidate list. In addition, merge_idx is MV selection information.
[0506] [MV Derivation > HMVP Mode]
[0507] Fig.42 is a diagram for explaining an example of the MV derivation process of the current picture based on the HMVP mode.
[0508] In the normal merge mode, one candidate MV is selected from the candidate MV list generated by referring to the encoded block (e.g., CU), thereby determining the MV of, for example, the CU of the current block. Here, other candidate MVs can also be registered in the candidate MV list. The mode of registering such other candidate MVs is called the HMVP mode.
[0509] In the HMVP mode, separately from the candidate MV list of the normal merge mode, a FIFO (First-In First-Out) buffer for HMVP is used to manage the candidate MVs.
[0510] In the FIFO buffer, motion information such as the MV of the blocks processed in the past is sequentially stored from the new FIFO buffer. In the management of this FIFO buffer, every time one block is processed, the MV of the latest block (i.e., the immediately preceding processed CU) is stored in the FIFO buffer, and instead, the MV of the earliest CU (i.e., the CU that was processed first) in the FIFO buffer is deleted from the FIFO buffer. In Fig.42In the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the earliest block.
[0511] Then, for example, the inter-frame prediction unit 126 sequentially checks each MV managed in the FIFO buffer starting from HMVP1 to see if the MV is different from all the candidate MVs already registered in the candidate MV list in the ordinary merge mode. Also, when it is determined that the MV is different from all the candidate MVs, the inter-frame prediction unit 126 can add the MV managed in the FIFO buffer as a candidate MV to the candidate MV list in the ordinary merge mode. At this time, the number of candidate MVs registered in the FIFO buffer can be one or more.
[0512] In this way, by using the HMVP mode, not only can the MVs of spatially or temporally adjacent blocks of the current block be added to the candidates, but also the MVs of the blocks processed in the past can be added to the candidates. As a result, by expanding the change of the candidate MVs in the ordinary merge mode, the possibility of improving the coding efficiency becomes higher.
[0513] In addition, the above MV can also be motion information. That is, the information stored in the candidate MV list and the FIFO buffer can include not only the value of the MV, but also information such as the information indicating the reference picture, the reference direction, and the number of pictures. In addition, the above block is, for example, a CU.
[0514] In addition, Fig.42 The candidate MV list and the FIFO buffer are an example, and the candidate MV list and the FIFO buffer can also be lists or buffers of different sizes from Fig.42 or a structure in which candidate MVs are registered in a different order from Fig.42 Here, the processing described is common to both the encoding device 100 and the decoding device 200.
[0515] In addition, the HMVP mode can also be applied to modes other than the ordinary merge mode. For example, motion information such as the MVs of the blocks processed in the affine mode in the past can be sequentially stored in a new FIFO buffer and used as candidate MVs. The mode in which the HMVP mode is applied in the affine mode can also be called the historical affine mode.
[0516] [MV Derivation>FRUC Mode]
[0517] Motion information may not be signaled from the encoding device 100 side, but may be derived on the decoding device 200 side. For example, motion information may also be derived by performing a motion search on the decoding device 200 side. In such a case, the motion search is performed without using the pixel values of the current block on the decoding device 200 side. Such a mode of performing a motion search on the decoding device 200 side includes a FRUC (frame rate up-conversion) mode, a PMMVD (pattern matched motion vector derivation) mode, or the like.
[0518] Fig.43 An example of FRUC processing is shown. First, referring to the MVs of each encoded block adjacent to the current block in space or time, a list representing these MVs as candidate MVs is generated (i.e., it is a candidate MV list and may also be common to the candidate MV list in the normal merge mode) (step Si_1). Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected as the best candidate MV based on this evaluation value. And, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. In addition, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral area of the position in the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and evaluation values in the reference picture may be performed on the peripheral area of the best candidate MV, and if there is an MV with a better evaluation value, the best candidate MV is updated to this MV and used as the final MV for the current block. The update to an MV with a better evaluation value may not be performed.
[0519] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_1 to Si_5 are performed, for example, for each block. For example, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. Also, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. In addition, the processes of steps Si_1 to Si_5 may be such that when they are not performed for all the blocks included in a slice but for some blocks, the inter-frame prediction using the FRUC mode for that slice ends. Similarly, the processes of steps Si_1 to Si_5 may be such that when they are performed for some blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends.
[0520] The same processing as that in the above block unit may also be performed in the case of processing in units of sub-blocks.
[0521] The evaluation value may also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the MV is compared with the reconstructed image of a specified region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). Then, the difference in pixel values of the two reconstructed images may be calculated for use as the evaluation value of the MV. In addition, it may be that other information is used in addition to the difference value to calculate the evaluation value.
[0522] Next, the pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (also referred to as the merge list) is selected as the starting point for the search based on pattern matching. As the pattern matching, the first pattern matching or the second pattern matching may be used. The first pattern matching and the second pattern matching may be cases respectively referred to as bilateral matching and template matching.
[0523] [MV Derivation>FRUC>Bilateral Matching]
[0524] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate MV, a region in another reference picture along the motion trajectory of the current block is used.
[0525] Fig.44This is a diagram illustrating an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along a motion trajectory. As Fig.44 shown, in the first pattern matching, by searching for the best-matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two MVs (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. Among multiple candidate MVs, the candidate MV with the best evaluation value can be selected as the best candidate MV.
[0526] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional MVs are derived.
[0527] [MV Derivation > FRUC > Template Matching]
[0528] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (the block adjacent to the current block in the current picture (e.g., the upper and / or left adjacent block)) and the block in the reference picture. Thus, in the second pattern matching, the block adjacent to the current block in the current picture is used as the specified region for calculating the evaluation value of the candidate MV as described above.
[0529] Fig.45 This is a diagram illustrating an example of the pattern matching (template matching) between the template in the current picture and the block in the reference picture. As Fig.45 shown, in the second pattern matching, the MV of the current block is derived by searching for the block in the reference picture (Ref0) that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the left adjacent and / or upper adjacent encoded region and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. Among multiple candidate MVs, the candidate MV with the best evaluation value can be selected as the best candidate MV.
[0530] Information indicating whether such a representation adopts the FRUC mode (e.g., referred to as the FRUC flag) is signaled at the CU level. In addition, in the case of adopting the FRUC mode (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (the first pattern matching or the second pattern matching) is signaled at the CU level. Additionally, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0531] [MV Derivation > Affine Mode]
[0532] The affine mode is a mode that generates MVs using an affine transformation. For example, MVs can also be derived in sub-block units based on the MVs of multiple adjacent blocks. This mode is sometimes referred to as the affine motion compensation prediction mode.
[0533] Fig.46A is a diagram for illustrating an example of the derivation of MVs in sub-block units based on the MVs of multiple adjacent blocks. In Fig.46A , the current block includes, for example, 16 sub-blocks composed of 4×4 pixels. Here, the motion vector v of the upper left control point of the current block is derived based on the MVs of adjacent blocks 0 , and similarly, the motion vector v of the upper right control point of the current block is derived based on the MVs of adjacent sub-blocks 1 . Then, according to the following equation (1A), the two motion vectors v 0 and v 1 are projected, and the motion vectors (v x , v y ) of each sub-block within the current block are derived.
[0534]
Equation 1
[0535]
[0536] Here, x and y respectively represent the horizontal position and vertical position of the sub-block, and w represents a predetermined weight coefficient.
[0537] Information indicating this affine mode (e.g., referred to as the affine flag) can be signaled at the CU level. Additionally, the signaling of the information indicating this affine mode does not need to be limited to the CU level and can be at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0538] In addition, in such an affine mode, several modes in which the methods for deriving the MVs of the upper-left and upper-right control points are different may also be included. For example, in the affine mode, there are two modes: the affine inter-frame (also referred to as the affine normal inter-frame) mode and the affine merge mode.
[0539] Fig.46B FIG. is a diagram for explaining an example of the derivation of the sub-block unit MVs in the affine mode using three control points. In Fig.46B , the current block includes 16 sub-blocks of 4×4 pixels. Here, the motion vector v 0 of the upper-left control point of the current block is derived based on the MVs of the adjacent blocks. Similarly, the motion vector v 1 of the upper-right control point of the current block is derived based on the MVs of the adjacent blocks, and the motion vector v 2 of the lower-left control point of the current block is derived based on the MVs of the adjacent blocks. Then, according to the following equation (1B), the three motion vectors v 0 , v 1 , and v 2 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.
[0540]
Equation 2
[0541]
[0542] Here, x and y represent the horizontal position and vertical position of the sub-block center, respectively, and w and h represent predetermined weight coefficients. Alternatively, w may represent the width of the current block, and h may represent the height of the current block.
[0543] Affine modes using different numbers of control points (e.g., two and three) can also be switched and signaled at the CU level. In addition, information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0544] In addition, in such an affine mode with three control points, several modes in which the methods for deriving the MVs of the upper-left, upper-right, and lower-left control points are different may also be included. For example, in the affine mode with three control points, similar to the affine mode with two control points, there are two modes: the affine inter-frame mode and the affine merge mode.
[0545] In addition, in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and may be other sizes. For example, the size of each sub-block may also be 8×8 pixels.
[0546] [MV Derivation > Affine Mode > Control Points]
[0547] Fig.47A , Fig.47B and Fig.47C are conceptual diagrams for explaining an example of MV derivation for control points in the affine mode.
[0548] In the affine mode, as Fig.47A shown, for example, based on a plurality of MVs corresponding to blocks encoded in the affine mode in the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) adjacent to the current block, a predicted MV for each of the control points of the current block is calculated. Specifically, these blocks are checked in the order of the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), and the first valid block encoded in the affine mode is determined. The MV of the control point of the current block is calculated based on the plurality of MVs corresponding to the determined block.
[0549] For example, as Fig.47B shown, when encoding the block A adjacent to the left side of the current block in the affine mode with 2 control points, the motion vectors v 3 and v 4 projected onto the positions of the upper left corner and the upper right corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v 3 and v 4 , the motion vector v 0 of the upper left corner control point of the current block and the motion vector v 1 of the upper right corner control point are calculated.
[0550] For example, as Fig.47C shown, when encoding the block A adjacent to the left side of the current block in the affine mode with 3 control points, the motion vectors v 3 , v 4 and v 5 projected onto the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v 3 , v 4 and v 5 , the motion vector v 0 of the upper left corner control point of the current block, the motion vector v 1 of the upper right corner control point, and the motion vector v 2 of the lower left corner control point are calculated.
[0551] In addition, Figures 47A to 47C the method for deriving the MVs shown can be used for deriving the MVs of each control point of the current block in step Sk_1 described later, and can also be used for Fig.50 in step Sk_2 described later. Fig.51Derivation of the predicted MVs of the control points of the current block in step Sj_1 shown above.
[0552] Fig.48A and Fig.48B is a conceptual diagram for explaining another example of the derivation of the control point MVs in the affine mode.
[0553] Fig.48A is a diagram for explaining the affine mode with two control points.
[0554] In this affine mode, as Fig.48A shown, the MV selected from the MVs of the respective encoded blocks A, B, and C adjacent to the current block is used as the motion vector v of the upper left control point of the current block 0 . Similarly, the MV selected from the MVs of the respective encoded blocks D and E adjacent to the current block is used as the motion vector v of the upper right control point of the current block 1 .
[0555] Fig.48B is a diagram for explaining the affine mode with three control points.
[0556] In this affine mode, as Fig.48B shown, the MV selected from the MVs of the respective encoded blocks A, B, and C adjacent to the current block is used as the motion vector v of the upper left control point of the current block 0 . Similarly, the MV selected from the MVs of the respective encoded blocks D and E adjacent to the current block is used as the motion vector v of the upper right control point of the current block 1 . In addition, the MV selected from the MVs of the respective encoded blocks F and G adjacent to the current block is used as the motion vector v of the lower left control point of the current block 2 .
[0557] In addition, Fig.48A and Fig.48B the MV derivation method shown can be used for the derivation of the MVs of the respective control points of the current block in step Sk_1 shown later, and can also be used for the derivation of the predicted MVs of the respective control points of the current block in step Sj_1 of Fig.50 shown later. Fig.51
[0558] Here, for example, in the case where the affine mode with different numbers of control points (e.g., two and three) is signaled at the CU level, etc., the number of control points may be different depending on the encoded blocks and the current block.
[0559] Fig.49A and Fig.49B It is a conceptual diagram showing an example of an MV derivation method for control points in the case where the number of control points in an encoded block and the current block is different.
[0560] For example, as Fig.49A shown, the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points. In this case, the motion vectors v 3 and v 4 projected onto the positions of the upper left corner and upper right corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v 3 and v 4 , the motion vector v 0 of the upper left corner control point of the current block and the motion vector v 1 of the upper right corner control point are calculated. Furthermore, based on the derived motion vectors v 0 and v 1 , the motion vector v 2 of the lower left corner control point is calculated.
[0561] For example, as Fig.49B shown, the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left side of the current block is encoded in an affine mode with three control points. In this case, the motion vectors v 3 , v 4 and v 5 projected onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v 3 , v 4 and v 5 , the motion vector v 0 of the upper left corner control point of the current block and the motion vector v 1 of the upper right corner control point are calculated.
[0562] In addition, Fig.49A and Fig.49B The MV derivation method shown can be used for the derivation of the MV of each control point of the current block in step Sk_1 described later, and can also be used for the derivation of the predicted MV of each control point of the current block in step Sj_1 of Fig.50 described later. Fig.51
[0563] [MV Derivation > Affine Mode > Affine Merge Mode]
[0564] Fig.50 It is a flowchart showing an example of the affine merge mode.
[0565] In the affine merge mode, first, the inter prediction unit 126 derives the MV of each control point of the current block (step Sk_1). The control points are, as shown in Fig.46A , the upper left and upper right points of the current block, or, as shown in Fig.46B , the upper left, upper right, and lower left points of the current block. At this time, the inter prediction unit 126 may also encode the MV selection information for identifying the two or three derived MVs into the stream.
[0566] For example, in the case of using the MV derivation method shown in Figures 47A to 47C , as shown in Fig.47A , the inter prediction unit 126 checks these blocks in the order of the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), and determines the initial valid block encoded in the affine mode.
[0567] The inter prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control point. For example, in the case where block A is determined and block A has two control points, as shown in Fig.47B , the inter prediction unit 126 calculates the motion vector v 3 and v 4 of the upper left control point of the current block and the motion vector v 0 and v 1 of the upper right control point of the current block based on the motion vectors v 3 and v 4 of the upper left and upper right corners of the encoded block containing block A. For example, by projecting the motion vectors v 3 and v 4 of the upper left and upper right corners of the encoded block onto the current block, the inter prediction unit 126 calculates the motion vector v 0 of the upper left control point of the current block and the motion vector v 1 of the upper right control point of the current block.
[0568] Alternatively, in the case where block A is determined and block A has three control points, as shown in Fig.47C , the inter prediction unit 126 calculates the motion vector v 3 , v 4 , and v 5 of the upper left, upper right, and lower left corners of the encoded block containing block A to calculate the motion vector v 0 of the upper left control point of the current block, the motion vector v 1 of the upper right control point of the current block, and the motion vector v 2 of the lower left control point of the current block. For example, by projecting the motion vectors v 3 , v 4 , and v 5 of the upper left, upper right, and lower left corners of the encoded block onto the current block.Projected onto the current block, the inter-frame prediction unit 126 calculates the motion vector v of the upper left control point of the current block 0 , the motion vector v of the upper right control point 1 and the motion vector v of the lower left control point 2 .
[0569] In addition, as shown above Fig.49A , block A can be determined. When block A has two control points, the MVs of three control points are calculated. Also, as shown above Fig.49B , block A can be determined. When block A has three control points, the MVs of two control points are calculated.
[0570] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, for each of the multiple sub-blocks, the inter-frame prediction unit 126 uses two motion vectors v 0 and v 1 and the above formula (1A), or uses three motion vectors v 0 , v 1 and v 2 and the above formula (1B) to calculate the MV of the sub-block as an affine MV (step Sk_2). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-block (step Sk_3). When the processes of steps Sk_2 and Sk_3 are respectively executed for all the sub-blocks included in the current block, the process of generating the predicted image using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and the predicted image of the current block is generated.
[0571] In addition, in step Sk_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can, for example, also be a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be Figures 47A to 47C the MV derivation method shown Fig.48A and Fig.48B the MV derivation method shown Fig.49A and Fig.49B the MV derivation method shown, as well as any combination of other MV derivation methods.
[0572] In addition, the candidate MV list can also include candidate MVs of prediction modes that perform prediction in units of sub-blocks other than the affine mode.
[0573] In addition, as a candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode having two control points and candidate MVs of an affine merge mode having three control points may also be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode having two control points and a candidate MV list including candidate MVs of an affine merge mode having three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes having two control points and the affine merge mode having three control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be the MV of the valid block among these blocks.
[0574] In addition, as MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent.
[0575] [MV Derivation > Affine Mode > Affine Inter Frame Mode]
[0576] Fig.51 is a flowchart showing an example of the affine inter frame mode.
[0577] In the affine inter frame mode, first, the inter frame prediction unit 126 derives the predicted MV (v 0 , v 1 ) or (v 0 , v 1 , v 2 ) of each of the two or three control points of the current block (step Sj_1). As shown in Fig.46A or Fig.46B , the control point is a point at the upper left corner, upper right corner, or lower left corner of the current block.
[0578] For example, in the case of using the MV derivation method shown in Fig.48A and Fig.48B , the inter frame prediction unit 126 derives the predicted MV (v 0 , v 1 ) or (v 0 , v 1 , v 2 ) of the control point of the current block by selecting the MV of a certain block among the encoded blocks near each control point of the current block shown in Fig.48A or Figure 48B . At this time, the inter frame prediction unit 126 encodes the prediction MV selection information for identifying the selected two or three predicted MVs into the stream.
[0579] For example, the inter-frame prediction unit 126 may determine which block's MV among the encoded blocks adjacent to the current block is to be used as the predicted MV of the control point by using cost evaluation or the like, and may describe in the bitstream a flag indicating which predicted MV is selected. That is, the inter-frame prediction unit 126 outputs prediction MV selection information such as a flag as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0580] Next, while respectively updating the predicted MVs selected or derived in step Sj_1 (step Sj_2), the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4). That is, the inter-frame prediction unit 126 sets the MV of each sub-block corresponding to the predicted MV to be updated as the affine MV, and calculates it using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference pictures (step Sj_4). Whenever the predicted MV is updated in step Sj_2, the processes of steps Sj_3 and Sj_4 are executed for all blocks within the current block. As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted MV that can obtain the minimum cost as the MV of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0581] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference pictures (step Sj_6).
[0582] In addition, in step Sj_1, the above-mentioned candidate MV list may also be generated. The candidate MV list may, for example, also be a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods may be Figures 47A to 47C the MV derivation methods shown, Figure 48A and Figure 48B the MV derivation methods shown, Figure 49A and Figure 49B the MV derivation methods shown, and any combination of other MV derivation methods.
[0583] In addition, the candidate MV list may also include candidate MVs of prediction modes that perform prediction in units of sub-blocks other than the affine mode.
[0584] In addition, as a candidate MV list, a candidate MV list including candidate MVs of an affine inter-frame mode having two control points and candidate MVs of an affine inter-frame mode having three control points may also be generated. Alternatively, a candidate MV list including candidate MVs of an affine inter-frame mode having two control points and a candidate MV list including candidate MVs of an affine inter-frame mode having three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of a mode of either an affine inter-frame mode having two control points or an affine inter-frame mode having three control points may be generated. The candidate MV may be, for example, the MVs of the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left), or may be the MVs of valid blocks among these blocks.
[0585] In addition, as prediction MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent out.
[0586] [MV Derivation > Triangular Mode]
[0587] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the current rectangular block. However, the inter-frame prediction unit 126 may generate a plurality of prediction images having shapes different from the rectangle for the current rectangular block, and generate a final rectangular prediction image by combining these plurality of prediction images. The shape different from the rectangle may also be a triangle, for example.
[0588] Figure 52A is a diagram for explaining the generation of two triangular prediction images.
[0589] The inter-frame prediction unit 126 performs motion compensation for the first partition of the triangle within the current block using the first MV of the first partition, thereby generating a triangular prediction image. Similarly, the inter-frame prediction unit 126 performs motion compensation for the second partition of the triangle within the current block using the second MV of the second partition, thereby generating a triangular prediction image. Then, the inter-frame prediction unit 126 combines these prediction images, thereby generating a rectangular prediction image identical to the current block.
[0590] In addition, as the prediction image of the first partition, a first rectangular prediction image corresponding to the current block may also be generated using the first MV. In addition, as the prediction image of the second partition, a second rectangular prediction image corresponding to the current block may also be generated using the second MV. The prediction image of the current block may also be generated by weighted addition of the first prediction image and the second prediction image. In addition, the region where the weighted addition is performed may also be only a partial region sandwiching the boundary between the first partition and the second partition.
[0591] Figure 52BIt is a conceptual diagram showing an example of a first part of a first partition overlapping a second partition, and a first sample set and a second sample set that can be weighted as part of a correction process. The first part can be, for example, one-fourth of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero. For example, N can be the integer 2. Figure 52B A rectangular partition representing a rectangular part having a width one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B An example in the center represents a rectangular partition having a height one-fourth of the height of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B An example on the right represents a triangular partition having a polygonal part with a height corresponding to 2 samples. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.
[0592] The first part can be a part of the first partition that overlaps an adjacent partition. Figure 52C It is a conceptual diagram showing a first part of a first partition that is a part of the first partition overlapping a part of an adjacent partition. For simplicity of explanation, a rectangular partition having a part overlapping a spatially adjacent rectangular partition is shown. Partitions with other shapes such as triangular partitions can be used, and the overlapping part can also overlap a partition that is spatially or temporally adjacent.
[0593] In addition, an example of generating a prediction image for two partitions respectively using inter-frame prediction is shown, but an intra-frame prediction can also be used to generate a prediction image for at least one partition.
[0594] Figure 53 It is a flowchart showing an example of a triangular mode.
[0595] In the triangular mode, first, the inter-frame prediction unit 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame prediction unit 126 can encode partition information, which is information related to the division into each partition, as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 can output the partition information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0596] Next, the inter-frame prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially located around the current block (step Sx_2). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0597] Then, the inter-frame prediction unit 126 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_1 (step Sx_3). At this time, the inter-frame prediction unit 126 may also encode the MV selection information for identifying the selected candidate MV as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 may output the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0598] Next, the inter-frame prediction unit 126 uses the selected first MV and the encoded reference picture to perform motion compensation, thereby generating a first prediction image (step Sx_4). Similarly, the inter-frame prediction unit 126 uses the selected second MV and the encoded reference picture to perform motion compensation, thereby generating a second prediction image (step Sx_5).
[0599] Finally, the inter-frame prediction unit 126 performs weighted addition on the first prediction image and the second prediction image, thereby generating a prediction image of the current block (step Sx_6).
[0600] In addition, in Figure 52A the example shown, the first partition and the second partition are respectively triangles, but they may also be trapezoids, or may be respectively of different shapes. Moreover, in Figure 52A the example shown, the current block is composed of two partitions, but it may also be composed of three or more partitions.
[0601] In addition, the first partition and the second partition may also overlap. That is, the first partition and the second partition may include the same pixel area. In this case, the prediction image in the first partition and the prediction image in the second partition may also be used to generate the prediction image of the current block.
[0602] In addition, in this example, an example in which prediction images are generated by inter-frame prediction in both two partitions is shown, but prediction images may also be generated by intra-frame prediction for at least one partition.
[0603] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different, or may be the same candidate MV list.
[0604] In addition, the partition information may include an index indicating a splitting direction for splitting at least the current block into a plurality of partitions. The MV selection information may also include an index indicating the selected first MV and an index indicating the selected second MV. One index may also represent multiple pieces of information. For example, one index summarizing a part or the whole of the partition information and a part or the whole of the MV selection information may be encoded.
[0605] [MV Derivation>ATMVP Mode]
[0606] Figure 54 FIG. is an example of the ATMVP mode for deriving MVs in units of sub-blocks.
[0607] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, candidate MVs in units of sub-blocks are registered in the candidate MV list for the normal merge mode.
[0608] Specifically, in the ATMVP mode, first, as Figure 54 shown, in the coded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, a temporal MV reference block corresponding to the current block is determined. Then, for each sub-block within the current block, the MV used for coding the region corresponding to the sub-block within the temporal MV reference block is determined. The MVs thus determined are included in the candidate MV list as candidate MVs for the sub-blocks of the current block. When selecting the candidate MVs for each sub-block from the candidate MV list, motion compensation is performed on the sub-block using the candidate MV as the MV of the sub-block. Thereby, a predicted image for each sub-block is generated.
[0609] In addition, in Figure 54 the example shown, as the surrounding MV reference block, the block adjacent to the lower left of the current block is used, but other blocks may also be used. In addition, the size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may also be switched in units of slices, bricks, or pictures, etc.
[0610] [Motion Search>DMVR]
[0611] Figure 55 FIG. is a diagram showing the relationship between the merge mode and DMVR.
[0612] The inter-frame prediction unit 126 derives the MV of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform MV search, that is, motion search (step Sl_2). Here, when it is determined not to perform motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in the merge mode.
[0613] On the other hand, when it is determined in step Sl_1 to perform motion search (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by DMVR.
[0614] Figure 56 It is a conceptual diagram for explaining an example of DMVR for determining MV.
[0615] First, for example, in the merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are determined based on the encoded picture in the L0 list, that is, the first reference picture (L0). Similarly, according to the candidate MV (L1), reference pixels are determined based on the encoded picture in the L1 list, that is, the second reference picture (L1). A template is generated by taking the average of these reference pixels.
[0616] Next, using this template, the peripheral areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV of the current block. In addition, the cost can also be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values.
[0617] Even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV to derive the final MV, any processing can be used.
[0618] Figure 57 It is a conceptual diagram for explaining another example of DMVR for determining MV. Figure 57 The example shown is different from Figure 56 an example of DMVR shown, and does not generate a template but calculates the cost.
[0619] First, the inter-frame prediction unit 126 searches the periphery of the reference block included in each of the reference pictures in the L0 list and the L1 list based on the candidate MV, that is, the initial MV, obtained from the candidate MV list. For example, as Figure 57As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion search, the inter-frame prediction unit 126 first sets the search position for the reference picture in the L0 list. The differential vector representing the set search position, specifically, the differential vector from the position represented by the initial MV (i.e., InitMV_L0) to this search position is MVd_L0. Then, the inter-frame prediction unit 126 determines the search position in the reference picture of the L1 list. This search position is represented by the differential vector from the position indicated by the initial MV (i.e., InitMV_L1) to this search position. Specifically, the inter-frame prediction unit 126 determines this differential vector as MVd_L1 by mirroring MVd_L0. That is, the inter-frame prediction unit 126 sets the positions that are symmetric from the position represented by the initial MV in the reference pictures of the L0 list and the L1 list as the search positions. The inter-frame prediction unit 126 calculates, for each search position, the sum of the absolute differences (SAD) of the pixel values within the block at this search position as a cost, and finds the search position with the minimum cost.
[0620] Figure 58A is a diagram showing an example of motion search in DMVR, Figure 58B is a flowchart showing an example of this motion search.
[0621] First, in Step1, the inter-frame prediction unit 126 calculates the costs of the search position represented by the initial MV (also called the starting point) and the 8 search positions around it. And the inter-frame prediction unit 126 determines whether the cost of the search position other than the starting point is the minimum. Here, when it is determined that the cost of the search position other than the starting point is the minimum, the inter-frame prediction unit 126 moves to the search position with the minimum cost and performs the processing of Step2. On the other hand, if the cost of the starting point is the minimum, the inter-frame prediction unit 126 skips the processing of Step2 and performs the processing of Step3.
[0622] In Step2, the inter-frame prediction unit 126 uses the search position moved according to the processing result of Step1 as the new starting point and performs the same search as the processing of Step1. And the inter-frame prediction unit 126 determines whether the cost of the search position other than this starting point is the minimum. Here, if the cost of the search position other than the starting point is the minimum, the inter-frame prediction unit 126 performs the processing of Step4. On the other hand, if the cost of the starting point is the minimum, the inter-frame prediction unit 126 performs the processing of Step3.
[0623] In Step4, the inter-frame prediction unit 126 processes the search position of this starting point as the final search position, and determines the difference between the position represented by the initial MV and this final search position as the differential vector.
[0624] In Step 3, the inter-frame prediction unit 126 determines the pixel position with sub-pixel precision having the minimum cost based on the costs at four points above, below, to the left, and to the right of the start point in Step 1 or Step 2, and sets this pixel position as the final search position. This pixel position with sub-pixel precision is determined by weighted addition of the vectors at the four points above, below, to the left, and to the right ((0, 1), (0, -1), (-1, 0), (1, 0)) using the costs at the respective search positions of the four points as weights. Then, the inter-frame prediction unit 126 determines the difference between the position represented by the initial MV and this final search position as the difference vector.
[0625] [Motion Compensation > BIO / OBMC / LIC]
[0626] In motion compensation, there are modes for generating a predicted image and correcting the predicted image. Such modes are, for example, BIO, OBMC, and LIC described later.
[0627] Figure 59 It is a flowchart showing an example of the generation of a predicted image.
[0628] The inter-frame prediction unit 126 generates a predicted image (step Sm_1), and corrects the predicted image by any of the above modes (step Sm_2).
[0629] Figure 60 It is a flowchart showing another example of the generation of a predicted image.
[0630] The inter-frame prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image using this MV (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting the predicted image (step Sn_4). Also, in LIC described later, it is also possible to correct the luminance and color difference in step Sn_4. On the other hand, when it is determined not to perform the correction process (No in step Sn_3), the inter-frame prediction unit 126 outputs the predicted image without correcting it as the final predicted image (step Sn_5).
[0631] [Motion Compensation > OBMC]
[0632] Not only the motion information of the current block obtained by motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame predicted image. Specifically, an inter-frame predicted image can also be generated in units of sub-blocks within the current block by weighted addition of a predicted image based on the motion information obtained by motion search (within the reference picture) and a predicted image based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or the OBMC mode.
[0633] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Additionally, the level at which these pieces of information are signaled is not limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0634] A more specific description of the OBMC mode will be given. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for explaining the outline of the predicted image correction process based on OBMC.
[0635] First, as Figure 62 shown, using the MV assigned to the current block, a predicted image (Pred) based on normal motion compensation is obtained. In Figure 62 , the arrow "MV" points to the reference picture and indicates which block in the reference picture the current block in the current picture refers to for obtaining the predicted image.
[0636] Next, the MV (MV_L) that has been derived for the already encoded left adjacent block is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two predicted images Pred and Pred_L, the first correction of the predicted image is performed. This has the effect of blending the boundaries between adjacent blocks.
[0637] Similarly, the Motion Vector (MV_U) that has been derived for the encoded upper adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_U). The MV (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted images that have undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block where the boundaries with adjacent blocks are blended (smoothed).
[0638] In addition, in the above example, a two-path correction method using the left adjacent and upper adjacent blocks is used, but this correction method can also be a three-path or more path correction method that also uses the right adjacent and / or lower adjacent blocks.
[0639] Furthermore, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.
[0640] In addition, the prediction image correction process of OBMC has been described here, where the prediction image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can be applied to each of the multiple reference pictures. In this case, through the OBMC image correction based on multiple reference pictures, after obtaining the corrected predicted images from each reference picture, the final predicted image is obtained by further overlapping the obtained multiple corrected predicted images.
[0641] In addition, in OBMC, the unit of the current block can be the PU unit or a sub-block unit obtained by further dividing the PU.
[0642] As a method for determining whether to apply OBMC, for example, there is a method of using a signal indicating whether to apply OBMC, namely obmc_flag. As a specific example, the encoding device 100 can also determine whether the current block belongs to a region with complex motion. When the encoding device 100 belongs to a region with complex motion, it sets the obmc_flag value to 1 and applies OBMC for encoding. When it does not belong to a region with complex motion, it sets the obmc_flag value to 0 and encodes the block without applying OBMC. On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, it switches whether to apply OBMC for decoding according to this value.
[0643] [Motion Compensation > BIO]
[0644] Next, a method for deriving the MV will be described. First, a mode for deriving the MV based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi - directional optical flow) mode. Additionally, the bi - directional optical flow can also be expressed as BDOF instead of BIO.
[0645] Figure 63 is a diagram for explaining the model assuming uniform linear motion. In Figure 63 , (v x , v y ) represents the velocity vector, and τ 0 , τ 1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) represents the MV corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) represents the MV corresponding to the reference picture Ref 1 .
[0646] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1 , MVy 1 ) are respectively expressed as (vxτ 0 , vyτ 0 ) and (-vxτ 1 , -vyτ 1 ), and the following optical flow equation (2) holds.
[0647]
Equation 3
[0648]
[0649] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vector in block units obtained from a candidate MV list or the like can be corrected in pixel units.
[0650] In addition, the MV can be derived on the decoding device 200 side by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, the motion vector can be derived in sub-block units based on the MVs of multiple adjacent blocks.
[0651] Figure 64 is a flowchart showing an example of inter-frame prediction according to BIO. In addition, Figure 65 is a diagram showing an example of the functional structure of the inter-frame prediction unit 126 that performs the inter-frame prediction according to BIO.
[0652] As Figure 65 shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a predicted image correction unit 126f. In addition, the memory 126a can also be the frame memory 122.
[0653] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 126 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_1). In addition, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0654] Next, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . In addition, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_2). Here, the interpolation image I 0 is the image included in the reference picture Ref0 derived for the current block, and the interpolation image I 1Is the image exported for the current block and referring to the image contained in Picture Ref1. Interpolated image I 0 And interpolated image I 1 Can each be the same size as the current block. Alternatively, in order to appropriately export the gradient image described later, interpolated image I 0 And interpolated image I 1 Can each be an image larger than the current block. In addition, interpolated image I 0 And I 1 Can include the predicted image derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1), and the motion compensation filter.
[0655] In addition, the gradient image derivation unit 126c derives the gradient image (Ix 0 And interpolated image I 1 Of the current block, (Ix 0 , Ix 1 , Iy 0 , Iy 1 )(Step Sy_3). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivation unit 126c can also derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image only needs to represent the spatial change amount of the pixel values along the horizontal direction or the vertical direction.
[0656] Next, the optical flow derivation unit 126d uses the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) to derive the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block (Step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and can also be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can be a 4x4 pixel sub-CU. In addition, the derivation of the optical flow can be performed in other units such as pixel units instead of sub-block units.
[0657] Next, the inter-frame prediction unit 126 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives a correction value for the values of the pixels included in the current block using the optical flow (vx, vy) (step Sy_5). Further, the predicted image correction unit 126f may also correct the predicted image of the current block using the correction value (step Sy_6). Additionally, the correction value may be derived on a per-pixel basis, or on a basis of multiple pixels or sub-blocks.
[0658] In addition, the processing flow of the BIO is not limited to Figure 64 the disclosed processing. It is possible to implement only Figure 64 a part of the disclosed processing, or to add or replace different processing, or to execute in a different processing order.
[0659] [Motion Compensation>LIC]
[0660] Next, an example of a mode for generating a predicted image (prediction) using LIC (local illumination compensation) will be described.
[0661] Figure 66A is a diagram for explaining an example of a method for generating a predicted image using a luminance correction process based on LIC. Further, Figure 66B is a flowchart showing an example of a method for generating a predicted image using this LIC.
[0662] First, the inter-frame prediction unit 126 derives an MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).
[0663] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values change between the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions within the reference picture specified by the derived MV. Then, the inter-frame prediction unit 126 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_3).
[0664] The inter-frame prediction unit 126 performs a brightness correction process on the reference image within the reference picture specified by the MV by applying the brightness correction parameter, and generates a prediction image for the current block (step Sz_4). That is, the prediction image, which is the reference image within the reference picture specified by the MV, is corrected based on the brightness correction parameter. In this correction, the brightness can be corrected, or the color difference can be corrected. That is, information indicating how the color difference changes can also be used to calculate the correction parameter for the color difference, and the color difference correction process can be performed.
[0665] In addition, Figure 66A The shape of the peripheral reference area in
[0666] is an example, and shapes other than this can also be used.
[0667] As a method for determining whether to apply LIC, for example, there is a method of using lic_flag, which is a signal indicating whether to apply LIC. As a specific example, in the encoding device 100, it is determined whether the current block belongs to an area where a brightness change has occurred. If it belongs to an area where a brightness change has occurred, the value 1 is set as lic_flag, and LIC is applied for encoding. If it does not belong to an area where a brightness change has occurred, the value 0 is set as lic_flag, and encoding is performed without applying LIC. On the other hand, in the decoding device 200, it is also possible to decode lic_flag described in the stream and switch whether to apply LIC for decoding according to its value.
[0668] As another method for determining whether to apply LIC, for example, there is also a method of determining according to whether LIC has been applied to the surrounding blocks. As a specific example, when the current block is processed in the merge mode, the inter-frame prediction unit 126 determines whether the surrounding encoded blocks selected when deriving the MV in the merge mode have been encoded by applying LIC. The inter-frame prediction unit 126 switches whether to apply LIC for encoding according to the result. In addition, in this example case, the same process also applies to the decoding device 200 side.
[0669] Using Figure 66A and Figure 66B LIC (brightness correction process) has been described above. Hereinafter, the detailed content thereof will be described.
[0670] First, the inter-frame prediction unit 126 derives the MV for obtaining the reference image corresponding to the current block from the reference picture that is an encoded picture.
[0671] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance pixel values change between the reference picture and the current picture for the current block, using the luminance pixel values of the left and upper adjacent encoded peripheral reference regions and the luminance pixel values at the same positions within the reference picture specified by the MV, and calculates the luminance correction parameter. For example, let the luminance pixel value of a certain pixel in the peripheral reference region within the current picture be p0, and the luminance pixel value of the pixel at the same position in the peripheral reference region within the reference picture be p1. The inter-frame prediction unit 126 calculates the coefficients A and B for optimizing A×p1 + B = p0 for multiple pixels within the peripheral reference region as the luminance correction parameter.
[0672] Next, the inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image within the reference picture specified by the MV using the luminance correction parameter. For example, let the luminance pixel value within the reference image be p2, and the luminance pixel value of the prediction image after the luminance correction process be p3. The inter-frame prediction unit 126 generates the prediction image after the luminance correction process by calculating A×p2 + B = p3 for each pixel within the reference image.
[0673] In addition, a part of the peripheral reference region shown in Figure 66A may also be used. For example, a region including a specified number of pixels separately removed at intervals from the upper adjacent pixel and the left adjacent pixel may be used as the peripheral reference region. In addition, the peripheral reference region is not limited to the region adjacent to the current block, and may also be a region not adjacent to the current block. In addition, in the example shown in Figure 66A , the peripheral reference region within the reference picture is the region specified by the MV of the current picture from the peripheral reference region within the current picture, but it may also be the region specified by other MVs. For example, this other MV may also be the MV of the peripheral reference region within the current picture.
[0674] In addition, here, the operations in the encoding device 100 have been described, but the operations in the decoding device 200 are the same.
[0675] In addition, LIC is not only applied to luminance, but can also be applied to color difference. At this time, the correction parameters can be derived separately for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.
[0676] In addition, LIC can also be applied in units of sub-blocks. For example, the correction parameter can also be derived using the peripheral reference region of the current sub-block and the peripheral reference region of the reference sub-block within the reference picture specified by the MV of the current sub-block.
[0677] [Prediction control unit]
[0678] The prediction control unit 128 selects one of the intra-predicted image (pixels or signals output from the intra-prediction unit 124) and the inter-predicted image (pixels or signals output from the inter-prediction unit 126), and outputs the selected predicted image to the subtraction unit 104 and the addition unit 116.
[0679] [Prediction parameter generation unit]
[0680] The prediction parameter generation unit 130 can output information related to the selection of the predicted image in intra-prediction, inter-prediction, and the prediction control unit 128, etc. as prediction parameters to the entropy encoding unit 110. The entropy encoding unit 110 can generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream, and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selected prediction signal (e.g., MV, prediction type, or the prediction mode used by the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.
[0681] [Decoding device]
[0682] Next, a decoding device 200 that can decode the stream output from the above encoding device 100 will be described. Figure 67 It is a block diagram showing an example of the functional structure of the decoding device 200 of the embodiment. The decoding device 200 is a device that decodes the encoded image, i.e., the stream, in block units.
[0683] As Figure 67 shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filtering unit 212, a frame memory 214, an intra-prediction unit 216, an inter-prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. In addition, the intra-prediction unit 216 and the inter-prediction unit 218 are each configured as part of the prediction processing unit.
[0684] [Installation example of the decoding device]
[0685] Figure 68 It is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 67 as shown in the decoding device 200, the multiple components Figure 68The illustrated processor b1 and memory b2 are installed and implemented.
[0686] The processor b1 is a circuit for information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit for decoding a stream. The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor b1 can also function as Figure 67 among the multiple components of the decoding device 200 shown in etc., except for the components for storing information.
[0687] The memory b2 is a dedicated or general-purpose memory for storing information for the processor b1 to decode a stream. The memory b2 can be either an electronic circuit or connected to the processor b1. In addition, the memory b2 can also be included in the processor b1. In addition, the memory b2 can also be an aggregate of multiple electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disc, etc., or can be represented as a storage or a recording medium, etc. In addition, the memory b2 can be either a non-volatile memory or a volatile memory.
[0688] For example, the memory b2 can store an image or a stream. In addition, a program for the processor b1 to decode a stream can also be stored in the memory b2.
[0689] In addition, for example, the memory b2 can also function as Figure 67 the component for storing information among the multiple components of the decoding device 200 shown in etc. Specifically, the memory b2 can function as Figure 67 the block memory 210 and the frame memory 214 shown. More specifically, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.) can be stored in the memory b2.
[0690] In addition, in the decoding device 200, not all of the multiple components shown in Figure 67 etc. need to be installed, and not all of the above multiple processes need to be performed. Figure 67 A part of the multiple components shown in etc. can be included in other devices, or a part of the above multiple processes can be performed by other devices.
[0691] Hereinafter, after explaining the overall processing flow of the decoding device 200, each component included in the decoding device 200 will be described. In addition, for the components that perform the same processing as the components included in the encoding device 100 among the components included in the decoding device 200, detailed descriptions will be omitted. For example, the inverse quantization unit 204, inverse transformation unit 206, addition unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 included in the decoding device 200 perform the same processing as the inverse quantization unit 112, inverse transformation unit 114, addition unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120 included in the encoding device 100, respectively.
[0692] [Overall Flow of Decoding Processing]
[0693] Figure 69 It is a flowchart showing an example of the overall decoding processing performed by the decoding device 200.
[0694] First, the segmentation determination unit 224 of the decoding device 200 determines the segmentation style of each of the multiple fixed-size blocks (128×128 pixels) included in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This segmentation style is the segmentation style selected by the encoding device 100. Then, the decoding device 200 performs the processing of steps Sp_2 to Sp_6 on each of the multiple blocks constituting this segmentation style.
[0695] The entropy decoding unit 202 decodes the encoded quantization coefficients and prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).
[0696] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore the prediction residual of the current block by performing inverse quantization and inverse transformation on the multiple quantization coefficients (step Sp_3).
[0697] Next, the prediction processing unit composed of the intra prediction unit 216, inter prediction unit 218, and prediction control unit 220 generates the prediction image of the current block (step Sp_4).
[0698] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction residual to the prediction image (step Sp_5).
[0699] Moreover, when generating this reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).
[0700] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeatedly executed.
[0701] In addition, the processing of steps Sp_1 to Sp_7 can be sequentially performed by the decoding device 200. Multiple processes among these processes can be performed in parallel, or the order can be changed.
[0702] [Segmentation determination unit]
[0703] Figure 70 FIG. is a diagram showing the relationship between the segmentation determination unit 224 and other components. As an example, the segmentation determination unit 224 can also perform the following processing.
[0704] The segmentation determination unit 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoding unit 202. Then, the segmentation determination unit 224 can determine the segmentation pattern of the fixed-size block based on the block information and the parameters. In addition, the segmentation determination unit 224 can also output the information indicating the determined segmentation pattern to the inverse transformation unit 206, the intra-frame prediction unit 216, and the inter-frame prediction unit 218. The inverse transformation unit 206 can perform inverse transformation on the transform coefficients based on the segmentation pattern indicated by the information from the segmentation determination unit 224. The intra-frame prediction unit 216 and the inter-frame prediction unit 218 can generate a predicted image based on the segmentation pattern indicated by the information from the segmentation determination unit 224.
[0705] [Entropy decoding unit]
[0706] Figure 71 FIG. is a block diagram showing an example of the functional structure of the entropy decoding unit 202.
[0707] The entropy decoding unit 202 generates quantization coefficients, prediction parameters, parameters related to the segmentation pattern, etc. by performing entropy decoding on the stream. For example, CABAC is used in this entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a multi-valued conversion unit 202c. The binary arithmetic decoding unit 202a performs arithmetic decoding on the stream as a binary signal using the context value derived by the context control unit 202b. Similar to the context control unit 110b of the encoding device 100, the context control unit 202b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. The multi-valued conversion unit 202c performs multi-valued conversion (debinarize) to convert the binary signal output from the binary arithmetic decoding unit 202a into a multi-valued signal representing the above quantization coefficients, etc. This multi-valued conversion is performed in the above-described binary conversion manner.
[0708] The entropy decoding unit 202 outputs the quantization coefficients to the inverse quantization unit 204 in block units. The entropy decoding unit 202 may also output the prediction parameters included in the stream (refer to Figure 1 ) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device 100 side.
[0709] [Entropy decoding unit]
[0710] Figure 72 is a diagram showing the process of CABAC in the entropy decoding unit 202.
[0711] First, in the CABAC in the entropy decoding unit 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding unit 202c and setting of the initial context value are performed. Then, the binary arithmetic decoding unit 202c and the multi-valued conversion unit 202c perform arithmetic decoding and multi-valued conversion on the encoded data of the CTU, for example. At this time, the context control unit 202b updates the context value each time arithmetic decoding is performed. Then, the context control unit 202b saves the context value as post-processing. The saved context value is used as the initial value of the context value for the next CTU, for example.
[0712] [Inverse quantization unit]
[0713] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the current block as the input from the entropy decoding unit 202. Specifically, the inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the current block based on the quantization parameters corresponding to the quantization coefficients, respectively. And the inverse quantization unit 204 outputs the inverse-quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transformation unit 206.
[0714] Figure 73 is a block diagram showing an example of the functional structure of the inverse quantization unit 204.
[0715] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204a, a predicted quantization parameter generation unit 204b, a quantization parameter storage unit 204d, and an inverse quantization processing unit 204e.
[0716] Figure 74 is a flowchart showing an example of the inverse quantization performed by the inverse quantization unit 204.
[0717] As an example, the inverse quantization unit 204 may be based on Figure 74In the process shown, inverse quantization processing is performed for each CU. Specifically, the quantization parameter generation unit 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation unit 204a obtains the differential quantization parameter of the current block from the entropy decoding unit 202 (step Sv_12).
[0718] Next, the predicted quantization parameter generation unit 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). The predicted quantization parameter generation unit 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).
[0719] Then, the quantization parameter generation unit 204a adds the differential quantization parameter of the current block obtained from the entropy decoding unit 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generation unit 204b (step Sv_15). By this addition, the quantization parameter of the current block is generated. In addition, the quantization parameter generation unit 204a stores the quantization parameter of the current block in the quantization parameter storage unit 204d (step Sv_16).
[0720] Next, the inverse quantization processing unit 204e uses the quantization parameter generated in step Sv_15 to inverse quantize the quantization coefficient of the current block into a transform coefficient (step Sv_17).
[0721] In addition, the differential quantization parameter can also be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0722] In addition, the inverse quantization unit 204 can include multiple inverse quantizers, or can inverse quantize the quantization coefficient using an inverse quantization method selected from multiple inverse quantization methods.
[0723] [Inverse Transform Unit]
[0724] The inverse transform unit 206 restores the prediction residual by performing an inverse transform on the transform coefficient that is the input from the inverse quantization unit 204.
[0725] For example, when the information read from the stream indicates the application of EMT or AMT (for example, the AMT flag is true), the inverse transform unit 206 performs an inverse transform on the transform coefficient of the current block based on the information indicating the transform type read.
[0726] In addition, for example, when the information read from the stream indicates the application of NSST, the inverse transform unit 206 applies an inverse re - transform to the transform coefficient.
[0727] Figure 75 It is a flowchart showing an example of the process performed by the inverse transform unit 206.
[0728] For example, the inverse transform unit 206 determines whether information indicating that orthogonal transformation is not to be performed exists in the stream (step St_11). Here, when it is determined that such information does not exist (No in step St_11), the inverse transform unit 206 acquires the information indicating the transform type that has been decoded by the entropy decoding unit 202 (step St_12). Next, the inverse transform unit 206 determines the transform type to be used in the orthogonal transformation in the encoding apparatus 100 based on this information (step St_13). And the inverse transform unit 206 performs inverse orthogonal transformation using the determined transform type (step St_14).
[0729] Figure 76 It is a flowchart showing another example of the process performed by the inverse transform unit 206.
[0730] For example, the inverse transform unit 206 determines whether the transform size is equal to or less than a specified value (step Su_11). Here, when it is determined that it is equal to or less than the specified value (Yes in step Su_11), the inverse transform unit 206 acquires from the entropy decoding unit 202 the information indicating which one of one or more transform types included in the first transform type group is used by the encoding apparatus 100 (step Su_12). In addition, such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.
[0731] The inverse transform unit 206 determines the transform type to be used in the orthogonal transformation in the encoding apparatus 100 based on this information (step Su_13). Then, the inverse transform unit 206 performs inverse orthogonal transformation on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, when it is determined in step Su_11 that the transform size is not equal to or less than the specified value (No in step Su_11), the inverse transform unit 206 performs inverse orthogonal transformation on the transform coefficients of the current block using the second transform type group (step Su_15).
[0732] In addition, as an example, the inverse orthogonal transformation performed by the inverse transform unit 206 can be implemented for each TU according to the Figure 75 or Figure 76 shown process. In addition, instead of decoding the information indicating the transform type used in the orthogonal transformation, inverse orthogonal transformation can be performed using a pre-specified transform type. Specifically, the transform type is, for example, DST7 or DCT8, etc., and in the inverse orthogonal transformation, the inverse transform basis function corresponding to this transform type is used.
[0733] [Addition unit]
[0734] The adder unit 208 reconstructs the current block by adding the prediction residual, which is the input from the inverse transform unit 206, to the predicted image, which is the input from the predictive control unit 220. That is, a reconstructed image of the current block is generated. Further, the adder unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.
[0735] [Block Memory]
[0736] The block memory 210 is a storage unit that stores blocks within the current picture, which are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the adder unit 208.
[0737] [Loop Filter Unit]
[0738] The loop filter unit 212 applies loop filtering to the reconstructed image generated by the adder unit 208, and outputs the filtered reconstructed image to the frame memory 214, a display device, and the like.
[0739] When the information indicating the ON / OFF of ALF read from the bitstream indicates that ALF is ON, one filter is selected from a plurality of filters based on the direction and activity of the gradient of locality, and the selected filter is applied to the reconstructed image.
[0740] Figure 77 is a block diagram showing an example of the functional configuration of the loop filter unit 212. Further, the loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding device 100.
[0741] The loop filter unit 212, for example, as Figure 77 shown, includes a deblocking filter processing unit 212a, an SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a applies the above-described deblocking filter processing to the reconstructed image. The SAO processing unit 212b applies the above-described SAO processing to the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. Further, the loop filter unit 212 may not include Figure 77 all of the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 212 may also be configured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 77 .
[0742] [Frame Memory]
[0743] The frame memory 214 is a storage unit that stores reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter unit 212.
[0744] [Prediction unit (intra prediction unit / inter prediction unit / prediction control unit)]
[0745] Figure 78 It is a flowchart showing an example of the processing performed by the prediction unit of the decoding device 200. In addition, as an example, the prediction unit is composed of all or some of the constituent elements of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The prediction processing unit includes, for example, the intra prediction unit 216 and the inter prediction unit 218.
[0746] The prediction unit generates a predicted image of the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. In addition, in the prediction signal, there are, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction unit uses the reconstructed image that has already been obtained by generating a predicted image for other blocks, restoring the prediction residual, and adding the predicted images, to generate the predicted image of the current block. The prediction unit of the decoding device 200 generates the same predicted image as the predicted image generated by the prediction unit of the encoding device 100. That is, the methods for generating the predicted images used in these prediction units are mutually common or corresponding.
[0747] The reconstructed image can be, for example, an image of a reference picture, or an image of the decoded blocks (i.e., the above-mentioned other blocks) within the current picture including the current block. The decoded blocks within the current picture are, for example, adjacent blocks of the current block.
[0748] Figure 79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device 200.
[0749] The prediction unit determines the method or mode for generating the predicted image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.
[0750] When it is determined that the first method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the first method (step Sr_2a). In addition, when it is determined that the second method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the second method (step Sr_2b). In addition, when it is determined that the third method is the mode for generating the predicted image, the prediction unit generates the predicted image according to the third method (step Sr_2c).
[0751] The first method, the second method, and the third method are different methods for generating the predicted image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.
[0752] Figure 80A and Figure 80B is a flowchart showing another example of the processing performed in the prediction unit of the decoding device 200.
[0753] As an example, the prediction unit may also perform prediction processing according to the Figure 80A and Figure 80B shown process. Additionally, Figure 80A and Figure 80B The intra-block copy shown is a mode belonging to inter-frame prediction, and it is a mode in which a block included in the current picture is referred to as a reference picture or a reference block. That is, in intra-block copy, a picture different from the current picture is not referred to. Additionally, Figure 80A The PCM mode shown is a mode belonging to intra-frame prediction, and it is a mode in which transformation and quantization are not performed.
[0754] [Intra-frame prediction unit]
[0755] The intra-frame prediction unit 216 generates a predicted image (i.e., an intra-frame predicted image) of the current block by performing intra-frame prediction based on the intra-frame prediction mode read from the stream and referring to the blocks within the current picture stored in the block memory 210. Specifically, the intra-frame prediction unit 216 generates an intra-frame predicted image by referring to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra-frame predicted image to the prediction control unit 220.
[0756] Additionally, when the intra-frame prediction mode of referring to the luminance block is selected in the intra-frame prediction of the chrominance difference block, the intra-frame prediction unit 216 may also predict the chrominance difference component of the current block based on the luminance component of the current block.
[0757] Furthermore, when the information read from the stream indicates the application of PDPC, the intra-frame prediction unit 216 corrects the pixel values after intra-frame prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0758] Figure 81 is a diagram showing an example of the processing performed by the intra-frame prediction unit 216 of the decoding device 200.
[0759] The intra prediction unit 216 first determines whether the MPM flag representing 1 exists in the stream (step Sw_11). Here, when it is determined that the MPM flag representing 1 exists (Yes in step Sw_11), the intra prediction unit 216 obtains information on the intra prediction mode selected in the encoding device 100 from the entropy decoding unit 202 (step Sw_12). Further, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Next, the intra prediction unit 216 determines the MPM (step Sw_13). The MPM is composed of, for example, six intra prediction modes. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information obtained in step Sw_12 from among the plurality of intra prediction modes included in this MPM (step Sw_14).
[0760] On the other hand, when it is determined in step Sw_11 that the MPM flag representing 1 does not exist in the stream (No in step Sw_11), the intra prediction unit 216 obtains information on the intra prediction mode selected in the encoding device 100 (step Sw_15). That is, the intra prediction unit 216 obtains information on the intra prediction mode selected in the encoding device 100 from among one or more intra prediction modes not included in the MPM from the entropy decoding unit 202. Further, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information obtained in step Sw_15 from among one or more intra prediction modes not included in this MPM (step Sw_17).
[0761] The intra prediction unit 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0762] [Inter prediction unit]
[0763] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks within the current block. In addition, a sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can also be switched in units of slices, bricks, or pictures, etc.
[0764] For example, the inter prediction unit 218 performs motion compensation using motion information (e.g., MV) read from the stream (e.g., prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction image of the current block or sub-block, and outputting the inter prediction image to the prediction control unit 220.
[0765] When the information decoded from the stream indicates the application of the OBMC mode, the inter-frame prediction unit 218 generates an inter-frame prediction image using not only the motion information of the current block obtained by motion search but also the motion information of adjacent blocks.
[0766] In addition, when the information decoded from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) decoded from the stream, thereby deriving motion information. And the inter-frame prediction unit 218 uses the derived motion information for motion compensation (prediction).
[0767] In addition, when applying the BIO mode, the inter-frame prediction unit 218 derives the MV based on a model assuming uniform linear motion. In addition, when the information decoded from the stream indicates the application of the affine mode, the inter-frame prediction unit 218 derives the MV in sub-block units based on the MVs of multiple adjacent blocks.
[0768] [Flow of MV derivation]
[0769] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device 200.
[0770] The inter-frame prediction unit 218 determines, for example, whether to decode motion information (e.g., MV). For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream or based on other information included in the stream. Here, when it is determined to decode the motion information, the inter-frame prediction unit 218 derives the MV of the current block in the mode of decoding the motion information. On the other hand, when it is determined not to decode the motion information, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the motion information.
[0771] Here, the modes of MV derivation include the ordinary inter-frame mode, the ordinary merge mode, the FRUC mode, and the affine mode, etc., which will be described later. Among these modes, the modes of decoding motion information include the ordinary inter-frame mode, the ordinary merge mode, and the affine mode (specifically, the affine inter-frame mode and the affine merge mode), etc. In addition, the motion information can include not only the MV but also the predicted MV selection information described later. In addition, the mode of not decoding the motion information includes the FRUC mode, etc. The inter-frame prediction unit 218 selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.
[0772] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device 200.
[0773] The inter-frame prediction unit 218 determines, for example, whether to decode the differential MV. For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the differential MV, the inter-frame prediction unit 218 can derive the MV of the current block in the mode of decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.
[0774] On the other hand, when it is determined not to decode the differential MV, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the differential MV. In this case, the encoded differential MV is not included in the stream.
[0775] Here, as described above, the derivation modes of the MV include the ordinary inter-frame mode, the ordinary merge mode, the FRUC mode, and the affine mode, etc., which will be described later. Among these modes, the modes for encoding the differential MV include the ordinary inter-frame mode and the affine mode (specifically, the affine inter-frame mode), etc. In addition, the modes for not encoding the differential MV include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode), etc. The inter-frame prediction unit 218 selects the mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0776] [MV Derivation > Ordinary Inter-Frame Mode]
[0777] For example, when the information read from the stream indicates that the ordinary inter-frame mode is applied, the inter-frame prediction unit 218 derives the MV in the ordinary merge mode based on the information read from the stream, and performs motion compensation (prediction) using this MV.
[0778] Figure 84 It is a flowchart showing an example of inter-frame prediction by the ordinary inter-frame mode in the decoding device 200.
[0779] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sg_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0780] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_11 as prediction motion vector candidates (also referred to as prediction MV candidates) respectively, in a predetermined priority order (step Sg_12). In addition, this priority order can also be determined in advance for each of the N prediction MV candidates.
[0781] Next, the inter-frame prediction unit 218 decodes the prediction MV selection information from the input stream, and uses the decoded prediction MV selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_13).
[0782] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream, and derives the MV of the current block by adding the difference value, which is the decoded differential MV, to the selected prediction MV (step Sg_14).
[0783] Finally, the inter-frame prediction unit 218 generates the predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are executed for each block. For example, when the processes of steps Sg_11 to Sg_15 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. In addition, when the processes of steps Sg_11 to Sg_15 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. Further, the processes of steps Sg_11 to Sg_15 may also be such that when they are not executed for all the blocks included in a slice but for a part of the blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, when the processes of steps Sg_11 to Sg_15 are executed for a part of the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture may also end.
[0784] [MV Derivation>Normal Merge Mode]
[0785] For example, in the case where the information read from the stream indicates the application of the normal merge mode, the inter-frame prediction unit 218 derives the MV in the normal merge mode and performs motion compensation (prediction) using the MV.
[0786] Figure 85 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode in the decoding apparatus 200.
[0787] The inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially located around the current block (step Sh_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0788] Next, the inter-frame prediction unit 218 derives the MV of the current block by selecting one candidate MV from among the multiple candidate MVs obtained in step Sh_11 (step Sh_12). Specifically, for example, the inter-frame prediction unit 218 obtains MV selection information included as a prediction parameter in the stream, and selects the candidate MV identified by the MV selection information as the MV of the current block.
[0789] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processing of steps Sh_11 to Sh_13 is performed on each block, for example. For example, when the processing of steps Sh_11 to Sh_13 is performed separately on all the blocks included in a slice, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processing of steps Sh_11 to Sh_13 is performed separately on all the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends. In addition, the processing of steps Sh_11 to Sh_13 may be such that when it is not performed on all the blocks included in a slice but on some of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processing of steps Sh_11 to Sh_13 may be such that when it is performed on some of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.
[0790] [MV Derivation > FRUC Mode]
[0791] For example, in the case where the information read from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 derives the MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding device 100 side but is derived on the decoding device 200 side. For example, the decoding device 200 may also derive the motion information by performing a motion search. In this case, the decoding device 200 does not use the pixel values of the current block for the motion search.
[0792] Figure 86 It is a flowchart showing an example of inter-frame prediction based on the FRUC mode in the decoding device 200.
[0793] First, the inter-frame prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block, and generates a list representing these MVs as candidate MVs (i.e., it is a candidate MV list and can also be common with the candidate MV list in the normal merge mode) (step Si_11). Next, the inter-frame prediction unit 218 selects the best candidate MV from among the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter-frame prediction unit 218 calculates the evaluation value of each candidate MV included in the candidate MV list, and selects one candidate MV as the best candidate MV based on this evaluation value. Then, the inter-frame prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Additionally, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the selected best candidate MV. That is, for the region around the best candidate MV, a search using pattern matching and evaluation values in the reference picture is performed. Furthermore, in the case where there is an MV with a good evaluation value, the best candidate MV may also be updated to this MV and used as the final MV for the current block. It is also possible not to perform the update to an MV with a better evaluation value.
[0794] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processes of steps Si_11 to Si_15 are performed on each block, for example. For example, when the processes of steps Si_11 to Si_15 are respectively performed on all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. Additionally, when the processes of steps Si_11 to Si_15 are respectively performed on all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. It is also possible to perform the processing in the same manner as the above block unit in sub-block units.
[0795] [MV Derivation > Affine Merge Mode]
[0796] For example, in the case where the information read from the stream indicates the application of the affine merge mode, the inter-frame prediction unit 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using this MV.
[0797] Figure 87 It is a flowchart showing an example of inter-frame prediction based on the affine merge mode in the decoding device 200.
[0798] In the affine merge mode, the inter-frame prediction unit 218 first derives the MV of each control point of the current block (step Sk_11). As Figure 46AAs shown, the control points are the points at the upper left and upper right corners of the current block, or as Figure 46B shown, are the points at the upper left, upper right, and lower left corners of the current block.
[0799] For example, in the case of using the Figures 47A to 47C MV export method shown, as Figure 47A shown, the inter-frame prediction unit 218 checks these blocks in the order of the decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and determines the first valid block decoded in the affine mode.
[0800] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to export the MV of the control points. For example, in the case where block A is determined and block A has two control points, as Figure 47B shown, the inter-frame prediction unit 218 projects the motion vectors v 3 and v 4 of the upper left and upper right corners of the decoded block containing block A onto the current block to calculate the motion vector v 0 of the upper left control point of the current block and the motion vector v 1 of the upper right control point. Thus, the MV of each control point is exported.
[0801] In addition, as Figure 49A shown, in the case where block A is determined and block A has two control points, it is also possible to calculate the MV of three control points, or it is also possible to determine block A as Figure 49B shown. In the case where block A has three control points, calculate the MV of two control points.
[0802] In addition, in the case where the stream includes MV selection information as a prediction parameter, the inter-frame prediction unit 218 can also use this MV selection information to export the MV of each control point of the current block.
[0803] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 218 uses two motion vectors v 0 and v 1 and the above formula (1A) for each of the multiple sub-blocks, or uses three motion vectors v 0 、v 1 and v 2And for the above formula (1B), the MV of the sub-block is calculated as an affine MV (step Sk_12). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine MVs and the decoded reference pictures (step Sk_13). When the processes of steps Sk_12 and Sk_13 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.
[0804] In addition, in step Sk_11, the above candidate MV list may also be generated. The candidate MV list may also be, for example, a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods may be Figures 47A to 47C the MV derivation methods shown, Figure 48A and Figure 48B the MV derivation methods shown, Figure 49A and Figure 49B the MV derivation methods shown, and any combination of other MV derivation methods.
[0805] In addition, the candidate MV list may also include candidate MVs of prediction modes performed in units of sub-blocks other than the affine mode.
[0806] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of the affine merge mode with two control points and candidate MVs of the affine merge mode with three control points may also be generated. Or, a candidate MV list including candidate MVs of the affine merge mode with two control points may be generated separately, and a candidate MV list including candidate MVs of the affine merge mode with three control points may be generated separately. Or, a candidate MV list including candidate MVs of one of the affine merge modes with two control points and the affine merge mode with three control points may be generated.
[0807] [MV Derivation > Affine Inter-frame Mode]
[0808] For example, when the information read from the stream indicates the application of the affine inter-frame mode, the inter-frame prediction unit 218 derives the MV in the affine inter-frame mode and performs motion compensation (prediction) using the MV.
[0809] Figure 88 is a flowchart showing an example of inter-frame prediction based on the affine inter-frame mode in the decoding device 200.
[0810] In the affine inter-frame mode, first, the inter-frame prediction unit 218 derives the predicted MVs (v 0 , v 1 ) or (v 0 , v1 , v 2 )(Step Sj_11). The control point is, for example, as shown in Figure 46A or Figure 46B , the point at the upper left corner, upper right corner, or lower left corner of the current block.
[0811] The inter-frame prediction unit 218 obtains the prediction MV selection information included as a prediction parameter in the stream, and uses the MV identified by the prediction MV selection information to derive the prediction MV of each control point of the current block. For example, in the case of using the MV derivation method shown in Figure 48A and Figure 48B , the inter-frame prediction unit 218 selects the MV of the block identified by the prediction MV selection information from the decoded blocks near each control point of the current block shown in Figure 48A or Figure 48B to derive the prediction MV of the control point of the current block (v 0 , v 1 ) or (v 0 , v 1 , v 2 ).
[0812] Next, the inter-frame prediction unit 218, for example, obtains each differential MV included as a prediction parameter in the stream, and adds the prediction MV of each control point of the current block and the differential MV corresponding to the prediction MV (Step Sj_12). Thereby, the MV of each control point of the current block is derived.
[0813] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 218 uses two motion vectors v 0 and v 1 and the above formula (1A) for each of the multiple sub-blocks, or uses three motion vectors v 0 , v 1 and v 2 and the above formula (1B) to calculate the MV of the sub-block as an affine MV (Step Sj_13). Then, the inter-frame prediction unit 218 uses these affine MVs and the decoded reference picture to perform motion compensation on the sub-block (Step Sj_14). When the processes of Steps Sj_13 and Sj_14 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.
[0814] In addition, in Step Sj_11, the above-mentioned candidate MV list may be generated in the same manner as in Step Sk_11.
[0815] [MV Derivation > Triangle Mode]
[0816] For example, in the case where the information read from the stream represents the application of the triangular mode, the inter-frame prediction unit 218 derives the MV in the triangular mode and performs motion compensation (prediction) using this MV.
[0817] Figure 89 FIG. is a flowchart showing an example of inter-frame prediction based on the triangular mode in the decoding apparatus 200.
[0818] In the triangular mode, first, the inter-frame prediction unit 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter-frame prediction unit 218 can obtain, from the stream, information related to the division into each partition, that is, partition information, as a prediction parameter. Further, the inter-frame prediction unit 218 can divide the current block into the first partition and the second partition according to the partition information.
[0819] Next, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sx_12). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0820] Then, the inter-frame prediction unit 218 selects a candidate MV for the first partition and a candidate MV for the second partition as a first MV and a second MV, respectively, from among the plurality of candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter-frame prediction unit 218 can also obtain, from the stream, MV selection information for identifying the selected candidate MVs as a prediction parameter. Then, the inter-frame prediction unit 218 can select the first MV and the second MV according to the MV selection information.
[0821] Next, the inter-frame prediction unit 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame prediction unit 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).
[0822] Finally, the inter-frame prediction unit 218 generates a prediction image of the current block by performing weighted addition on the first prediction image and the second prediction image (step Sx_16).
[0823] [Motion Search > DMVR]
[0824] For example, in the case where the information read from the stream represents the application of DMVR, the inter-frame prediction unit 218 performs motion search by DMVR.
[0825] Figure 90 FIG. is a flowchart showing an example of motion search based on DMVR in the decoding apparatus 200.
[0826] The inter-frame prediction unit 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter-frame prediction unit 218 derives the final MV for the current block by searching the surrounding area of the reference picture represented by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is determined by DMVR.
[0827] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device 200.
[0828] First, in Step1 shown in Figure 58A , the inter-frame prediction unit 218 calculates the search positions (also referred to as start points) represented by the initial MV and the costs of 8 search positions around it. And the inter-frame prediction unit 218 determines whether the cost of the search position other than the start point is the minimum. Here, when it is determined that the cost of the search position other than the start point is the minimum, the inter-frame prediction unit 218 moves to the search position with the minimum cost and performs the Figure 58A processing of Step2 shown in. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 skips the Figure 58A processing of Step2 shown in and performs the processing of Step3.
[0829] In Figure 58A Step2 shown in, the inter-frame prediction unit 218 uses the search position moved according to the processing result of Step1 as a new start point and performs the same search as the processing of Step1. And the inter-frame prediction unit 218 determines whether the cost of the search position other than this start point is the minimum. Here, if the cost of the search position other than the start point is the minimum, the inter-frame prediction unit 218 performs the processing of Step4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 performs the processing of Step3.
[0830] In Step4, the inter-frame prediction unit 218 processes the search position of this start point as the final search position, and determines the difference between the position indicated by the initial MV and this final search position as the difference vector.
[0831] In Figure 58A Step3 shown in, the inter-frame prediction unit 218 determines the pixel position with the minimum cost and the fractional precision based on the costs of the 4 points above, below, left, and right of the start point in Step1 or Step2, and uses this pixel position as the final search position. This pixel position with the fractional precision is determined by weighted addition of the vectors of the 4 points above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)) with the costs of the search positions of these 4 points as weights. Then, the inter-frame prediction unit 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.
[0832] [Motion compensation > BIO / OBMC / LIC]
[0833] For example, in the case where the information read from the stream indicates the application of the correction of the predicted image, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to the correction mode. This mode is, for example, the above-mentioned BIO, OBMC, and LIC, etc.
[0834] Figure 92 It is a flowchart showing an example of the generation of the predicted image in the decoding device 200.
[0835] The inter-frame prediction unit 218 generates a predicted image (step Sm_11), and corrects the predicted image by any of the above modes (step Sm_12).
[0836] Figure 93 It is a flowchart showing another example of the generation of the predicted image in the decoding device 200.
[0837] The inter-frame prediction unit 218 derives the MV of the current block (step Sn_11). Then, the inter-frame prediction unit 218 generates a predicted image using this MV (step Sn_12), and determines whether to perform the correction process (step Sn_13). For example, the inter-frame prediction unit 218 obtains the prediction parameters included in the stream, and determines whether to perform the correction process based on the prediction parameters. This prediction parameter is, for example, a flag indicating whether to apply the above-mentioned various modes. Here, when it is determined to perform the correction process (yes in step Sn_13), the inter-frame prediction unit 218 generates the final predicted image by correcting the predicted image (step Sn_14). In addition, in LIC, the luminance and chrominance differences of the predicted image can be corrected in step Sn_14. On the other hand, when it is determined not to perform the correction process (no in step Sn_13), the inter-frame prediction unit 218 outputs the predicted image without correcting it as the final predicted image (step Sn_15).
[0838] [Motion compensation > OBMC]
[0839] For example, in the case where the information read from the stream indicates the application of OBMC, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to OBMC.
[0840] Figure 94 It is a flowchart showing an example of the correction of the predicted image based on OBMC in the decoding device 200. In addition, Figure 94 The flowchart of Figure 62 shows the process of correcting the predicted image using the current picture and the reference picture shown in
[0841] First, asFigure 62 As shown, the inter-frame prediction unit 218 uses the MV assigned to the current block to obtain a prediction image (Pred) based on normal motion compensation.
[0842] Next, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_L) that has been derived for the decoded left adjacent block to the current block to obtain a prediction image (Pred_L). Then, the inter-frame prediction unit 218 performs the first correction of the prediction image by overlapping the two prediction images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0843] Similarly, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_U) that has been derived for the decoded upper adjacent block to the current block to obtain a prediction image (Pred_U). Then, the inter-frame prediction unit 218 performs the second correction of the prediction image by overlapping the prediction image Pred_U with the prediction image that has been subjected to the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The prediction image obtained through the second correction is the final prediction image of the current block that is blended (smoothed) with the boundaries of adjacent blocks.
[0844] [Motion Compensation > BIO]
[0845] For example, in the case where the information read from the stream indicates the application of BIO, when generating a prediction image, the inter-frame prediction unit 218 corrects the prediction image according to BIO.
[0846] Figure 95 It is a flowchart showing an example of the correction of the prediction image based on BIO in the decoding device 200.
[0847] As Figure 63 shown, the inter-frame prediction unit 218 uses two reference pictures (Ref0, Ref1) different from the picture (CurPic) containing the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 218 uses these two motion vectors (M0, M1) to derive the prediction image of the current block (step Sy_11). In addition, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0848] Next, the inter-frame prediction unit 218 uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . In addition, the inter-frame prediction unit 218 uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_12). Here, the interpolation image I0 is an interpolated image I exported for the current block with reference to the image included in picture Ref0 1 is an image exported for the current block with reference to the image included in picture Ref1. Interpolated image I 0 and interpolated image I 1 can each be the same size as the current block. Alternatively, in order to appropriately export the gradient image described later, interpolated image I 0 and interpolated image I 1 can each also be an image larger than the current block. In addition, interpolated image I 0 and I 1 can include a predicted image derived by applying motion vectors (M0, M1) and reference pictures (L0, L1), and a motion compensation filter.
[0849] In addition, the inter-frame prediction unit 218 derives the gradient image (Ix 0 and interpolated image I 1 for the current block from interpolated image I 0 , Ix 1 , Iy 0 , Iy 1 )(step Sy_13). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Ix 0 , Ix 1 ). The inter-frame prediction unit 218 can also derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image may be an image representing the spatial change amount of pixel values along the horizontal direction or the vertical direction.
[0850] Next, the inter-frame prediction unit 218 uses the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) to derive the optical flow (vx, vy) as the velocity vector described above in units of a plurality of sub-blocks constituting the current block (step Sy_14). As an example, the sub-block may be a 4x4 pixel sub-CU.
[0851] Next, the inter-frame prediction unit 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame prediction unit 218 derives a correction value for the values of the pixels included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter-frame prediction unit 218 can correct the predicted image of the current block using the correction value (step Sy_16). Additionally, the correction value can be derived on a per-pixel basis, or on a basis of multiple pixels or sub-blocks.
[0852] In addition, the processing flow of the BIO is not limited to Figure 95 the disclosed processing. It is possible to implement only Figure 95 a part of the disclosed processing, or to add or replace different processing, or to execute in a different processing order.
[0853] [Motion Compensation > LIC]
[0854] For example, in the case where the information read from the stream indicates the application of LIC, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to LIC.
[0855] Figure 96 is a flowchart showing an example of the correction of the predicted image based on LIC in the decoding device 200.
[0856] First, the inter-frame prediction unit 218 obtains a reference image corresponding to the current block from the decoded reference picture using the MV (step Sz_11).
[0857] Next, the inter-frame prediction unit 218 extracts information indicating how the luminance values change between the reference picture and the current picture for the current block (step Sz_12). As Figure 66A shown, this extraction is based on the luminance pixel values of the decoded left adjacent reference region (peripheral reference region) and the decoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions within the reference picture specified by the derived MV. Then, the inter-frame prediction unit 218 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_13).
[0858] The inter-frame prediction unit 218 generates a predicted image for the current block by performing a luminance correction process applying its luminance correction parameter to the reference image within the reference picture specified by the MV. That is, the predicted image, which is the reference image within the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, it is possible to correct the luminance or the chromatic aberration.
[0859] [Prediction Control Unit]
[0860] The prediction control unit 220 selects one of the intra-predicted image and the inter-predicted image, and outputs the selected predicted image to the addition unit 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-prediction unit 216, and the inter-prediction unit 218 on the decoder device 200 side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-prediction unit 124, and the inter-prediction unit 126 on the encoder device 100 side.
[0861] [Regarding the method of dividing pictures]
[0862] Figure 97 It is a conceptual diagram showing the relationship between sub-pictures, slices, and tiles. Here, an example of a picture division method using sub-pictures, slices, and tiles is shown. Specifically, in Figure 97 the example, the picture is divided into multiple tiles, multiple slices, and multiple sub-pictures.
[0863] Tiles are obtained by dividing the picture horizontally and vertically. In Figure 97 the example, the picture is divided into two parts horizontally and two parts vertically, and thus is divided into a total of 4 tiles.
[0864] Moreover, in Figure 97 the example, each tile is divided into one or more slices. Specifically, the upper-left tile is divided into 1 slice, the upper-right tile is divided into 4 slices, the lower-left tile is divided into 2 slices, and the lower-right tile is divided into 3 slices. In this case, each slice is rectangular, so it is called a rectangular slice. A rectangular slice can be composed of multiple tiles that make up a rectangular area, or can be composed of one or more CTU rows within one tile.
[0865] Moreover, in Figure 97 the example, the picture is divided into multiple sub-pictures, and the multiple sub-pictures are respectively units that summarize one or more slices or one or more tiles.
[0866] Sub-pictures are used, for example, for the following purposes. That is, sub-pictures can be extracted from the stream in units of sub-pictures, and only the extracted sub-pictures are transmitted as a stream and can also be decoded. In addition, the extracted sub-pictures can also be combined with sub-pictures of other streams and reconstructed as one stream.
[0867] Figure 98 It is a conceptual diagram showing the relationship between tiles and slices. In Figure 98 the example, the picture is divided into multiple tiles and multiple slices. Specifically, the picture is divided into four parts horizontally and four parts vertically, and thus is divided into a total of 16 tiles.
[0868] Furthermore, inFigure 98 In the example, multiple tiles form one slice. Slice (1) consists of 5 tiles, slice (2) consists of 5 tiles, and slice (3) consists of 6 tiles. At this time, each slice aggregates multiple tiles that are consecutive in raster scan order and is determined as one slice, so it is called a raster - scan slice.
[0869] [Morphology of syntax related to sub - pictures]
[0870] Figure 99 It is a syntax structure diagram related to sub - pictures determined by using a grid in a sequence parameter set. In this example, sub - pictures composed of multiple unconnected regions are suppressed. In addition, sub - pictures composed of multiple regions forming a non - rectangular region are suppressed. Furthermore, compared with the case where sub - picture index signals are generated for each grid, the number of signaled sub - picture indexes is reduced. Specifically, the grid indexes of the upper - left and lower - right of the sub - picture are signaled.
[0871] Here, top_left_grid_idx[i] is the grid index of the upper - left grid element in the sub - picture with sub - picture index i. bottom_right_grid_idx[i] is the grid index of the lower - right grid element in the sub - picture with sub - picture index i.
[0872] subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 represent the width and height of each grid element. Each grid element can be 4×4 pixels. In this case, subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 can be 3 respectively.
[0873] In addition, in the sequence parameter set, either the size of the CTU can be signaled or the number of sub - pictures can be signaled.
[0874] log2_ctu_size_minus5 represents the size of the CTU. Specifically, log2_ctu_size_minus5 is a value obtained by subtracting 5 from the value representing the number of pixels on one side of the CTU in base - 2 logarithm. For example, when log2_ctu_size_minus5 is 0, the size of the CTU is 32×32 pixels. When log2_ctu_size_minus5 is 1, the size of the CTU is 64×64 pixels.
[0875] The subpics_present_flag indicates whether sub-picture information exists in the sequence parameter set. For example, when the subpics_present_flag is 1 (true), sub-picture information exists in the sequence parameter set. On the other hand, when the subpics_present_flag is 0 (false), no sub-picture information exists.
[0876] num_subpics_minus1 represents the number of sub-pictures. Specifically, num_subpics_minus1 corresponds to the value obtained by subtracting 1 from the number of sub-pictures. num_subpics_minus1 can be replaced by num_subpics_minus2, which corresponds to the value obtained by subtracting 2 from the number of sub-pictures.
[0877] That is to say, it can be that when the subpics_present_flag is true, the number of sub-pictures is 2 or more, or it can be that when the number of sub-pictures is 2 or more, the subpics_present_flag is true. In other words, it can be that when the number of sub-pictures is 2 or more, the sub-picture information is signaled.
[0878] max_subpics_minus1 represents the maximum number of sub-pictures. max_subpics_minus1 may also not exist. That is, the coding and decoding of max_subpics_minus1 can also be omitted. For example, it can also be that only num_subpics_minus1 exists among max_subpics_minus1 and num_subpics_minus1.
[0879] Each grid element can be a CTB (Coding Tree Block) or a CTU (Coding Tree Unit). When each grid element has the same size as the CTU, subpic_grid_col_width_minus1 and subpic_grid_row_height_minus1 may not be signaled.
[0880] Figure 100 It is a syntax structure diagram related to sub-pictures determined using CTUs in the sequence parameter set. In addition, the representation of CTUs is used here, but the representation of CTBs, which basically corresponds to the same area as the CTU, can also be used. Strictly speaking, a CTB corresponds to a block of luminance or a block of chrominance, and the luminance and chrominance in a CTB are processed separately, while a CTU corresponds to a common area for luminance and chrominance.
[0881] Compared with the example of Figure 99 , in the example of Figure 100 , top_left_grid_idx[i] is replaced by top_left_ctu_idx[i]. Additionally, bottom_right_grid_idx[i] is replaced by bottom_right_ctu_idx[i]. Moreover, max_subpics_minus1, subpic_grid_col_width_minus1, and subpic_grid_row_height_minus1 are not signaled.
[0882] In addition to the above, Figure 100 the example of Figure 99 is the same as the example of
[0883] Here, top_left_ctu_idx[i] is the CTU index of the top-left CTU in the subpicture with subpicture index i. bottom_right_ctu_idx[i] is the CTU index of the bottom-right CTU in the subpicture with subpicture index i. That is, the grid and grid elements can be appropriately read as CTUs. Additionally, as described above, the grid and grid elements can also be read as CTBs. Figure 100 Moreover, since top_left_ctu_idx[i] and bottom_right_ctu_idx[i] are used to define the subpicture, information related to the size of the CTU is used. Therefore, as Figure 100 shown, log2_ctu_size_minus5 can be encoded earlier than the signal used to define the subpicture. Additionally, the information related to the size of the CTU is not limited to the
[0884] Figure 101A is a conceptual diagram showing the grid indices of each region in the picture. Specifically, Figure 101A shows the picture and multiple regions within the picture. Figure 101A Each of the multiple regions shown in Figure 101A is a grid element. In
[0885] Figure 101B is a conceptual diagram showing the subpicture indices of each region in the picture. Figure 101B also shows the picture and multiple regions within the picture. Similar to Figure 101A , Figure 101B each of the multiple regions shown in Figure 101BAmong them, a sub-picture index represented by a numerical value is assigned to each grid element.
[0886] For example, 0 is assigned as the sub-picture index to the grid element with grid index 0 and the grid element with grid index 6. That is, the grid element with grid index 0 and the grid element with grid index 6 constitute 1 sub-picture with sub-picture index 0. In this example, top_left_grid_idx[i] and bottom_right_grid_idx[i] are determined as follows.
[0887] top_left_grid_idx[0] = 0
[0888] bottom_right_grid_idx[0] = 6
[0889] top_left_grid_idx[1] = 1
[0890] bottom_right_grid_idx[1] = 7
[0891] top_left_grid_idx[2] = 2
[0892] bottom_right_grid_idx[2] = 8
[0893] …
[0894] As described above, each grid element can be a CTB or a CTU.
[0895] [Variation examples of syntax related to sub-pictures]
[0896] In Figure 99 the signaling of the grid address in the example, the horizontal position and the vertical position of the grid can be signaled separately.
[0897] Figure 102 is a syntax structure diagram showing another example related to the sub-picture determined by using the grid in the sequence parameter set.
[0898] Compared with the example of Figure 99 in Figure 102 the example, top_left_grid_idx[i] is replaced by subpic_grid_addr_TL_x[i] and subpic_grid_addr_TL_y[i]. In addition, bottom_right_grid_idx[i] is replaced by subpic_grid_addr_BR_x[i] and subpic_grid_addr_BR_y[i]. Regarding the others, Figure 102 the example ofFigure 99 is the same as the example of
[0899] Here, subpic_grid_addr_TL_x[i] is the horizontal grid address of the upper-left grid element in the sub-picture with sub-picture index i. That is, subpic_grid_addr_TL_x[i] indicates which grid element the upper-left grid element in the sub-picture with sub-picture index i is from the leftmost grid element in the picture counting from the left. In addition, the leftmost grid element in the picture is the 0th grid element counting from the leftmost grid element to the right.
[0900] In addition, subpic_grid_addr_TL_y[i] is the vertical grid address of the upper-left grid element with sub-picture index = i. That is, subpic_grid_addr_TL_y[i] indicates which grid element the upper-left grid element in the sub-picture with sub-picture index i is from the uppermost grid element in the picture counting from the top. In addition, the uppermost grid element in the picture is the 0th grid element counting from the uppermost grid element to the bottom.
[0901] In addition, subpic_grid_addr_BR_x[i] is the horizontal grid address of the lower-right grid element in the sub-picture with sub-picture index i. That is, subpic_grid_addr_BR_x[i] indicates which grid element the lower-right grid element in the sub-picture with sub-picture index i is from the leftmost grid element in the picture counting from the left.
[0902] In addition, subpic_grid_addr_BR_y[i] is the vertical grid address of the lower-right grid element with sub-picture index = i. That is, subpic_grid_addr_BR_y[i] indicates which grid element the lower-right grid element in the sub-picture with sub-picture index i is from the uppermost grid element in the picture counting from the top.
[0903] In Figure 101A and Figure 101B 's example, these values are determined in the following way.
[0904] subpic_grid_addr_TL_x[0] = 0
[0905] subpic_grid_addr_TL_y[0] = 0
[0906] subpic_grid_addr_BR_x[0] = 0
[0907] subpic_grid_addr_BR_y[0] = 1
[0908] …
[0909] That is, the upper-left grid element in the sub-picture with sub-picture index 0 is the 0th grid element counted from the grid element at the left end of the picture, and is the grid element at the left end of the picture. In addition, the upper-left grid element in the sub-picture with sub-picture index 0 is the 0th grid element counted from the grid element at the upper end of the picture, and is the grid element at the upper end of the picture.
[0910] In addition, the lower-right grid element in the sub-picture with sub-picture index 0 is the 0th grid element counted from the grid element at the left end of the picture, and is the grid element at the left end of the picture. In addition, the lower-right grid element in the sub-picture with sub-picture index 0 is the 1st grid element counted from the grid element at the upper end of the picture, and is the grid element below the grid element at the upper end of the picture.
[0911] Similarly, in Figure 100 the signaling of the CTU address in the example of
[0912] Figure 103 is a syntax structure diagram showing another example related to the sub-picture determined by using CTU in the sequence parameter set.
[0913] Compared with the example of Figure 100 in the example of Figure 103 top_left_ctu_idx[i] is replaced by subpic_ctu_addr_TL_x[i] and subpic_ctu_addr_TL_y[i]. In addition, bottom_right_ctu_idx[i] is replaced by subpic_ctu_addr_BR_x[i] and subpic_ctu_addr_BR_y[i].
[0914] Except for the above, Figure 103 the example of Figure 100 is basically the same as the example of Figure 103 However, in the example of Figure 100 the syntax elements omitted in
[0915] are specifically shown. These syntax elements may not exist.
[0916] In addition, subpic_ctu_addr_TL_y[i] is the vertical CTU address of the top-left CTU in the sub-picture with sub-picture index i. That is, subpic_ctu_addr_TL_y[i] indicates which CTU the top-left CTU in the sub-picture with sub-picture index i is, counting down from the CTU at the top of the picture. Moreover, the CTU at the top of the picture is the 0th CTU counting down from the CTU at the top.
[0917] In addition, subpic_ctu_addr_BR_x[i] is the horizontal CTU address of the bottom-right CTU in the sub-picture with sub-picture index i. That is, subpic_ctu_addr_BR_x[i] indicates which CTU the bottom-right CTU in the su...
Claims
1. An encoding device, wherein, it comprises: a circuit; and a memory connected to the circuit, during operation, the circuit encodes CTU size information representing the size of a CTU, where the CTU is a coding tree unit, into a sequence parameter set, encodes sub - picture information into the sequence parameter set, where the sub - picture information uses the CTU to respectively represent the horizontal position and the vertical position of the area of a sub - picture as a rectangular area within the picture, the horizontal position is represented by the position of the CTU that constitutes the sub - picture with reference to the left end of the picture, the vertical position is represented by the position of the CTU that constitutes the sub - picture with reference to the upper end of the picture, the CTU size information is located before the sub - picture information in the sequence parameter set.
2. The encoding device according to claim 1, wherein, the circuit encodes the sub - picture information only when there are two or more sub - pictures in the picture.
3. The encoding device according to claim 1 or 2, wherein, when the CTU is neither located at the right end nor at the lower end of the picture, the CTU is determined as a square area with a fixed size; when the CTU is located at the right end or the lower end of the picture, the CTU is determined as the square area with the fixed size or an area smaller than the square area with the fixed size.
4. A decoding device, wherein, it comprises: a circuit; and a memory connected to the circuit, during operation, the circuit decodes CTU size information representing the size of a CTU, where the CTU is a coding tree unit, from a sequence parameter set, decodes sub - picture information from the sequence parameter set, where the sub - picture information uses the CTU to respectively represent the horizontal position and the vertical position of the area of a sub - picture as a rectangular area within the picture, the horizontal position is represented by the position of the CTU that constitutes the sub - picture with reference to the left end of the picture, the vertical position is represented by the position of the CTU that constitutes the sub - picture with reference to the upper end of the picture, the CTU size information is located before the sub - picture information in the sequence parameter set.
5. The decoding device according to claim 4, wherein, the circuit decodes the sub - picture information only when there are two or more sub - pictures in the picture.
6. The decoding device according to claim 4 or 5, wherein, when the CTU is neither located at the right end nor at the lower end of the picture, the CTU is determined as a square area with a fixed size; when the CTU is located at the right end or the lower end of the picture, the CTU is determined as the square area with the fixed size or an area smaller than the square area with the fixed size.
7. An encoding method, wherein, CTU size information representing the size of a CTU, where the CTU is a coding tree unit, is encoded into a sequence parameter set, Encode sub - picture information into the sequence parameter set, where the sub - picture information uses the CTU to respectively represent the horizontal position and vertical position of the area of the sub - picture as a rectangular area within the picture. The horizontal position is represented by the position of the CTU that forms the sub - picture with respect to the left end of the picture. The vertical position is represented by the position of the CTU that forms the sub - picture with respect to the upper end of the picture. The CTU size information is located in the sequence parameter set before the sub - picture information.
8. A decoding method wherein, Decode CTU size information representing the size of the CTU from the sequence parameter set, where the CTU is a coding tree unit. Decode sub - picture information from the sequence parameter set, where the sub - picture information uses the CTU to respectively represent the horizontal position and vertical position of the area of the sub - picture as a rectangular area within the picture. The horizontal position is represented by the position of the CTU that forms the sub - picture with respect to the left end of the picture. The vertical position is represented by the position of the CTU that forms the sub - picture with respect to the upper end of the picture. The CTU size information is located in the sequence parameter set before the sub - picture information.
9. A non - transitory computer - readable storage medium storing a bitstream wherein, The bitstream contains a sequence parameter set for a decoding device to perform decoding processing. The sequence parameter set contains: CTU size information representing the size of the CTU, where the CTU is a coding tree unit; and Sub - picture information that uses the CTU to respectively represent the horizontal position and vertical position of the area of the sub - picture as a rectangular area within the picture. In the decoding process, decode the CTU size information and the sub - picture information from the sequence parameter set. The horizontal position is represented by the position of the CTU that forms the sub - picture with respect to the left end of the picture. The vertical position is represented by the position of the CTU that forms the sub - picture with respect to the upper end of the picture. The CTU size information is located in the sequence parameter set before the sub - picture information.