Encoding Device, Decoding Device, Encoding Method, Decoding Method, and Storage Medium
By using constraints in video encoding to determine the relationship between sub-pictures, tiles and slices of the picture, the problem of difficulty in determining the object area in the prior art is solved, and the encoding efficiency and picture quality are improved, the processing volume and circuit scale are reduced, and the processing speed is improved.
Patent Information
- Application Number
- CN202080059890.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-27
- Filing Date
- 2020-09-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-09-25
AI Technical Summary
The existing video encoding technology has the need to improve coding efficiency, picture quality, processing volume, circuit scale and processing speed, especially in determining the processing object area.
By using constraints during the encoding and decoding process to determine the relationship between sub-pictures, tiles and slices of the picture, ensuring that the sub-picture does not lose the containing slices and tiles, and efficient processing object area determination is achieved.
It improves encoding efficiency, improves picture quality, reduces processing volume and circuit scale, and improves processing speed. The elements and actions of encoding and decoding are appropriately selected.
Smart Images

Figure CN114375580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an encoding device, a decoding device, an encoding method, a decoding method, and a storage medium. Background Art
[0002] Video encoding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this progress, in order to process the continuously increasing amount of digital video data in various applications, there is always a need to provide improvements and optimizations to video encoding technology. The present invention relates to further progress, improvements, and optimizations in video encoding.
[0003] In addition, Non-Patent Document 1 relates to an example of an existing standard related to the above video encoding technology.
[0004] Prior Art Documents
[0005] Non-Patent Documents
[0006] Non-Patent Document 1: H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) Summary of the Invention
[0007] Problems to be Solved by the Invention
[0008] Regarding the encoding method as described above, for the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, or the appropriate selection of elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks, it is desired to propose a new method.
[0009] The present invention provides a structure or method that can contribute to one or more of, for example, the improvement of encoding efficiency, the improvement of image quality, the reduction of processing amount, the reduction of circuit scale, the improvement of processing speed, and the appropriate selection of elements or operations. In addition, the present invention may include a structure or method that can contribute to benefits other than the above.
[0010] Means for Solving the Problems
[0011] For example, an encoding device according to one embodiment of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit determines one or more tiles constituting an image and one or more sub-images constituting the image according to a constraint condition, where the constraint condition means that one of the one or more tiles contains the one or more sub-images without loss.
[0012] In video encoding technology, in order to improve encoding efficiency, improve image quality, reduce circuit scale, etc., it is desirable to propose new methods.
[0013] Each embodiment or a part of the structure or method in the present invention can respectively achieve at least any one of, for example, improvement of encoding efficiency, improvement of image quality, reduction of encoding / decoding processing amount, reduction of circuit scale, or improvement of encoding / decoding processing speed. Or, each embodiment or a part of the structure or method in the present invention can respectively make an appropriate selection of elements / actions such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, etc. during encoding and decoding. In addition, the present invention also includes disclosure of structures or methods that can provide benefits other than the above. For example, it is a structure or method that improves encoding efficiency while suppressing an increase in processing amount.
[0014] Based on the specification and the drawings, further advantages and effects in one embodiment of the present invention are clarified. These advantages and / or effects are respectively obtained through several embodiments and the features described in the specification and the drawings, but it is not necessary to provide all of them in order to obtain one or more advantages and / or effects.
[0015] In addition, these general or specific forms can be implemented by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or can also be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0016] Advantages of the Invention
[0017] The structure or method according to one embodiment of the present invention can contribute to one or more of, for example, improvement of encoding efficiency, improvement of image quality, reduction of processing amount, reduction of circuit scale, improvement of processing speed, and appropriate selection of elements or actions. In addition, the structure or method according to one embodiment of the present invention can also contribute to benefits other than the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram showing an example of the structure of a transmission system according to an embodiment.
[0019] Figure 2 It is a diagram showing an example of the hierarchical structure of data in a stream.
[0020] Figure 3 This is a diagram showing an example of the structure of a slice.
[0021] Figure 4 This is a diagram showing an example of the structure of a tile.
[0022] Figure 5 This is a diagram showing an example of the coding structure in scalable coding.
[0023] Figure 6 This is a diagram showing an example of the coding structure in scalable coding.
[0024] Figure 7 This is a block diagram showing an example of the functional structure of the coding device in the embodiment.
[0025] Figure 8 This is a block diagram showing an example of the installation of the coding device.
[0026] Figure 9 This is a flowchart showing an example of the overall coding process performed by the coding device.
[0027] Figure 10 This is a diagram showing an example of block division.
[0028] Figure 11 This is a diagram showing an example of the functional structure of the division unit.
[0029] Figure 12 This is a diagram showing an example of the division pattern.
[0030] Figure 13A This is a diagram showing an example of the syntax tree of the division pattern.
[0031] Figure 13B This is a diagram showing another example of the syntax tree of the division pattern.
[0032] Figure 14 This is a table showing the transform basis functions corresponding to each transform type.
[0033] Figure 15 This is a diagram showing an example of SVT.
[0034] Figure 16 This is a flowchart showing an example of the process performed by the transform unit.
[0035] Figure 17 This is a flowchart showing another example of the process performed by the transform unit.
[0036] Figure 18 This is a block diagram showing an example of the functional structure of the quantization unit.
[0037] Figure 19It is a flowchart showing an example of quantization performed by a quantization unit.
[0038] Figure 20 It is a block diagram showing an example of the functional structure of an entropy encoding unit.
[0039] Figure 21 It is a diagram showing the process of CABAC in an entropy encoding unit.
[0040] Figure 22 It is a block diagram showing an example of the functional structure of a loop filtering unit.
[0041] Figure 23A It is a diagram showing an example of the shape of a filter used in an ALF (adaptive loop filter).
[0042] Figure 23B It is a diagram showing another example of the shape of a filter used in an ALF.
[0043] Figure 23C It is a diagram showing another example of the shape of a filter used in an ALF.
[0044] Figure 23D It is a diagram showing an example where a Y sample (first component) is used in CCALF for Cb and Cr (multiple components different from the first component).
[0045] Figure 23E It is a diagram showing a diamond-shaped filter.
[0046] Figure 23F It is a diagram showing an example of JC-CCALF.
[0047] Figure 23G It is a diagram showing an example of weight_index candidates for JC-CCALF.
[0048] Figure 24 It is a block diagram showing an example of the detailed structure of a loop filtering unit that functions as a DBF.
[0049] Figure 25 It is a diagram showing an example of deblocking filtering with filtering characteristics symmetric with respect to a block boundary.
[0050] Figure 26 It is a diagram for explaining an example of a block boundary where deblocking filtering processing is performed.
[0051] Figure 27 It is a diagram showing an example of a Bs value.
[0052] Figure 28 It is a flowchart showing an example of the processing performed by the prediction unit of an encoding device.
[0053] Figure 29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.
[0054] Figure 30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device.
[0055] Figure 31 It is a diagram showing an example of 67 intra prediction modes in intra prediction.
[0056] Figure 32 It is a flowchart showing an example of the processing performed by the intra prediction unit.
[0057] Figure 33 It is a diagram showing an example of each reference picture.
[0058] Figure 34 It is a conceptual diagram showing an example of a reference picture list.
[0059] Figure 35 It is a flowchart showing the flow of the basic processing of inter prediction.
[0060] Figure 36 It is a flowchart showing an example of MV derivation.
[0061] Figure 37 It is a flowchart showing another example of MV derivation.
[0062] Figure 38A It is a diagram showing an example of the classification of each mode of MV derivation.
[0063] Figure 38B It is a diagram showing an example of the classification of each mode of MV derivation.
[0064] Figure 39 It is a flowchart showing an example of inter prediction based on the normal inter mode.
[0065] Figure 40 It is a flowchart showing an example of inter prediction based on the normal merge mode.
[0066] Figure 41 It is a diagram for explaining an example of the MV derivation process based on the normal merge mode.
[0067] Figure 42 It is a diagram for explaining an example of the MV derivation process based on the HMVP mode.
[0068] Figure 43 It is a flowchart showing an example of FRUC (frame rate up conversion).
[0069] Figure 44 This is a diagram for explaining an example of pattern matching (bidirectional matching) between two blocks along a motion trajectory.
[0070] Figure 45 This is a diagram for explaining an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture.
[0071] Figure 46A This is a diagram for explaining an example of the derivation of the MV of a sub-block unit in an affine mode using two control points.
[0072] Figure 46B This is a diagram for explaining an example of the derivation of the MV of a sub-block unit in an affine mode using three control points.
[0073] Figure 47A This is a conceptual diagram for explaining an example of the derivation of the MV of control points in an affine mode.
[0074] Figure 47B This is a conceptual diagram for explaining an example of the derivation of the MV of control points in an affine mode.
[0075] Figure 47C This is a conceptual diagram for explaining an example of the derivation of the MV of control points in an affine mode.
[0076] Figure 48A This is a diagram for explaining an affine mode with two control points.
[0077] Figure 48B This is a diagram for explaining an affine mode with three control points.
[0078] Figure 49A This is a conceptual diagram for explaining an example of the method for deriving the MV of control points when the number of control points in an encoded block and the current block is different.
[0079] Figure 49B This is a conceptual diagram for explaining another example of the method for deriving the MV of control points when the number of control points in an encoded block and the current block is different.
[0080] Figure 50 This is a flowchart showing an example of the processing of an affine merge mode.
[0081] Figure 51 This is a flowchart showing an example of the processing of an affine inter-frame mode.
[0082] Figure 52A This is a diagram for explaining the generation of predicted images of two triangles.
[0083] Figure 52BIt is a conceptual diagram showing an example of the first part of the first partition, the first sample set, and the second sample set.
[0084] Figure 52C It is a conceptual diagram showing the first part of the first partition.
[0085] Figure 53 It is a flowchart showing an example of a triangular pattern.
[0086] Figure 54 It is a diagram showing an example of the ATMVP mode for deriving MV in sub-block units.
[0087] Figure 55 It is a diagram showing the relationship between the merge mode and DMVR (dynamic motion vector refreshing).
[0088] Figure 56 It is a conceptual diagram for explaining an example of DMVR.
[0089] Figure 57 It is a conceptual diagram for explaining another example of DMVR for determining MV.
[0090] Figure 58A It is a diagram showing an example of motion search in DMVR.
[0091] Figure 58B It is a flowchart showing an example of motion search in DMVR.
[0092] Figure 59 It is a flowchart showing an example of the generation of a predicted image.
[0093] Figure 60 It is a flowchart showing another example of the generation of a predicted image.
[0094] Figure 61 It is a flowchart for explaining an example of the predicted image correction process based on OBMC (overlapped block motion compensation).
[0095] Figure 62 It is a conceptual diagram for explaining an example of the predicted image correction process based on OBMC.
[0096] Figure 63 It is a diagram for explaining a model assuming uniform linear motion.
[0097] Figure 64 It is a flowchart showing an example of inter-frame prediction according to BIO.
[0098] Figure 65It is a diagram showing an example of the functional structure of an inter-frame prediction unit that performs inter-frame prediction according to BIO.
[0099] Figure 66A It is a diagram for explaining an example of a method for generating a prediction image using a luminance correction process based on LIC (local illumination compensation).
[0100] Figure 66B It is a flowchart showing an example of a method for generating a prediction image using a luminance correction process based on LIC.
[0101] Figure 67 It is a block diagram showing the functional structure of a decoding device according to an embodiment.
[0102] Figure 68 It is a block diagram showing an example of the installation of a decoding device.
[0103] Figure 69 It is a flowchart showing an example of the overall decoding process performed by a decoding device.
[0104] Figure 70 It is a diagram showing the relationship between a segmentation determination unit and other components.
[0105] Figure 71 It is a block diagram showing an example of the functional structure of an entropy decoding unit.
[0106] Figure 72 It is a diagram showing the process of CABAC in an entropy decoding unit.
[0107] Figure 73 It is a block diagram showing an example of the functional structure of an inverse quantization unit.
[0108] Figure 74 It is a flowchart showing an example of inverse quantization performed by an inverse quantization unit.
[0109] Figure 75 It is a flowchart showing an example of the process performed by an inverse transform unit.
[0110] Figure 76 It is a flowchart showing another example of the process performed by an inverse transform unit.
[0111] Figure 77 It is a block diagram showing an example of the functional structure of a loop filter unit.
[0112] Figure 78 It is a flowchart showing an example of the process performed by the prediction unit of a decoding device.
[0113] Figure 79It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device.
[0114] Figure 80A It is a flowchart showing a part of another example of the processing performed by the prediction unit of the decoding device.
[0115] Figure 80B It is a flowchart showing the remaining part of another example of the processing performed by the prediction unit of the decoding device.
[0116] Figure 81 It is a diagram showing an example of the processing performed by the intra prediction unit of the decoding device.
[0117] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device.
[0118] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device.
[0119] Figure 84 It is a flowchart showing an example of inter prediction based on the normal inter mode in the decoding device.
[0120] Figure 85 It is a flowchart showing an example of inter prediction based on the normal merge mode in the decoding device.
[0121] Figure 86 It is a flowchart showing an example of inter prediction based on the FRUC mode in the decoding device.
[0122] Figure 87 It is a flowchart showing an example of inter prediction based on the affine merge mode in the decoding device.
[0123] Figure 88 It is a flowchart showing an example of inter prediction based on the affine inter mode in the decoding device.
[0124] Figure 89 It is a flowchart showing an example of inter prediction based on the triangle mode in the decoding device.
[0125] Figure 90 It is a flowchart showing an example of motion search based on DMVR in the decoding device.
[0126] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device.
[0127] Figure 92 It is a flowchart showing an example of the generation of a predicted image in the decoding device.
[0128] Figure 93It is a flowchart showing another example of the generation of a predicted image in a decoding device.
[0129] Figure 94 It is a flowchart showing an example of the correction of a predicted image based on OBMC in a decoding device.
[0130] Figure 95 It is a flowchart showing an example of the correction of a predicted image based on BIO in a decoding device.
[0131] Figure 96 It is a flowchart showing an example of the correction of a predicted image based on LIC in a decoding device.
[0132] Figure 97 It is a conceptual diagram showing the relationship between sub - pictures, slices, and tiles.
[0133] Figure 98 It is a conceptual diagram showing the relationship between tiles and slices.
[0134] Figure 99 It is a syntactic structure diagram related to sub - pictures in the first form.
[0135] Figure 100A It is a conceptual diagram showing the tile index of each region in a picture.
[0136] Figure 100B It is a conceptual diagram showing the sub - picture index of each region in a picture.
[0137] Figure 101 It is a syntactic structure diagram related to sub - pictures in the second form.
[0138] Figure 102 It is a syntactic structure diagram related to sub - pictures in the third form.
[0139] Figure 103A It is a conceptual diagram showing the brick index of each region in a picture.
[0140] Figure 103B It is a conceptual diagram showing the tile index of each region in a picture.
[0141] Figure 103C It is a conceptual diagram showing the sub - picture index of each region in a picture.
[0142] Figure 104 It is a syntactic structure diagram related to sub - pictures in the fourth form.
[0143] Figure 105 It is a syntactic structure diagram related to sub - pictures in the fifth form.
[0144] Figure 106A It is a conceptual diagram showing the slice index of each region in a picture.
[0145] Figure 106B It is a conceptual diagram showing tile indices of respective regions in a picture.
[0146] Figure 106C It is a conceptual diagram showing sub - picture indices of respective regions in a picture.
[0147] Figure 107 It is a syntactic structure diagram related to sub - pictures in the 6th form.
[0148] Figure 108 It is a syntactic structure diagram related to sub - pictures determined by a grid.
[0149] Figure 109 It is a syntactic structure diagram related to sub - pictures determined by a CTU.
[0150] Figure 110A It is a conceptual diagram showing grid indices of respective grid elements in a picture.
[0151] Figure 110B It is a conceptual diagram showing sub - picture indices of respective grid elements in a picture.
[0152] Figure 111 It is a syntactic structure diagram showing another example related to sub - pictures determined by a grid.
[0153] Figure 112 It is a syntactic structure diagram showing another example related to sub - pictures determined by a CTU.
[0154] Figure 113 It is a conceptual diagram showing an example of dividing one picture into multiple sub - pictures, multiple tiles, and multiple slices according to two conditions.
[0155] Figure 114 It is a conceptual diagram showing an example of dividing one picture into multiple sub - pictures, multiple tiles, and multiple slices only according to the first condition.
[0156] Figure 115 It is a flowchart showing the operation of an encoding device in an embodiment.
[0157] Figure 116 It is a flowchart showing the operation of a decoding device in an embodiment.
[0158] Figure 117 It is an overall structure diagram of a content supply system for implementing a content distribution service.
[0159] Figure 118 It is a diagram showing an example of a display screen of a web page.
[0160] Figure 119This is a diagram showing an example of a display screen of a web page.
[0161] Figure 120 This is a diagram showing an example of a smartphone.
[0162] Figure 121 This is a block diagram showing an example of the structure of a smartphone. Detailed Implementation Manner
[0163] [Introduction]
[0164] In the encoding of moving images, multiple pictures constituting the moving image are each divided into various regions. For example, a picture is divided into multiple CTUs (Coding Tree Units), or divided into multiple tiles, or divided into multiple slices.
[0165] For example, a CTU corresponds to a square region of a fixed size. A tile is a rectangular region determined according to one or more rows within the picture, one or more columns within the picture, or both. A slice corresponds to one NAL unit as a data packet and corresponds to one or more tiles or one or more consecutive CTUs within one tile.
[0166] However, sometimes the processing target region does not match one of the CTU, tile, and slice, and it is sometimes difficult to determine the processing target region.
[0167] Therefore, an encoding device according to one aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit determines one or more sub-pictures constituting a picture, one or more tiles constituting the picture, and one or more slices constituting the picture according to both the first constraint condition and the second constraint condition, partitions and encodes the picture according to the one or more sub-pictures, the one or more tiles, and the one or more slices. The first constraint condition is that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture contains without loss one or more slices among the one or more slices constituting the picture. The second constraint condition is that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture contains without loss one or more tiles among the one or more tiles constituting the picture, or one tile among the one or more tiles constituting the picture contains without loss one or more sub-pictures including the sub-picture among the one or more sub-pictures constituting the picture.
[0168] Thereby, sometimes it is possible to suppress the complication of the relationship between the sub-picture and the slice and the relationship between the sub-picture and the tile. Therefore, sometimes it is possible to efficiently determine the sub-picture as the processing target region according to the slice and the tile.
[0169] For example, the one or more sub-images that make up the image are multiple sub-images, the one or more tiles that make up the image are multiple tiles, and the one or more slices that make up the image are multiple slices. The circuit determines the multiple sub-images, the multiple tiles, and the multiple slices that make up the image according to both the first constraint condition and the second constraint condition, partitions and encodes the image according to the multiple sub-images, the multiple tiles, and the multiple slices. The first constraint condition is that for each of the multiple sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the multiple slices that make up the image. The second constraint condition is that for each of the multiple sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the multiple tiles that make up the image, or one of the multiple tiles that make up the image contains, without loss, one or more of the multiple sub-images including the sub-image.
[0170] Accordingly, when multiple sub-images, multiple tiles, and multiple slices are determined within an image, it is sometimes possible to suppress the complication of the relationship between the sub-images and the slices and the relationship between the sub-images and the tiles. Therefore, when multiple sub-images, multiple tiles, and multiple slices are determined within an image, it is sometimes possible to efficiently determine the sub-images as processing target areas according to the slices and the tiles.
[0171] In addition, for example, a decoding device according to an aspect of the present invention includes a circuit and a memory connected to the circuit. During operation, the circuit determines one or more sub-images, one or more tiles, and one or more slices that make up an image according to both the first constraint condition and the second constraint condition, partitions and decodes the image according to the one or more sub-images, the one or more tiles, and the one or more slices. The first constraint condition is that for each of the one or more sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the one or more slices that make up the image. The second constraint condition is that for each of the one or more sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the one or more tiles that make up the image, or one of the one or more tiles that make up the image contains, without loss, one or more of the one or more sub-images including the sub-image.
[0172] Accordingly, sometimes it is possible to suppress the complication of the relationship between the sub-images and the slices and the relationship between the sub-images and the tiles. Therefore, sometimes it is possible to efficiently identify the sub-images as the processing target areas in accordance with the slices and the tiles.
[0173] In addition, for example, the one or more sub-images constituting the picture are a plurality of sub-images, the one or more tiles constituting the picture are a plurality of tiles, the one or more slices constituting the picture are a plurality of slices, and the circuit determines the plurality of sub-images constituting the picture, the plurality of tiles constituting the picture, and the plurality of slices constituting the picture in accordance with both the first constraint condition and the second constraint condition, partitions and decodes the picture in accordance with the plurality of sub-images, the plurality of tiles, and the plurality of slices, the first constraint condition being that, for each of the plurality of sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the plurality of slices constituting the picture, and the second constraint condition being that, for each of the plurality of sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the plurality of tiles constituting the picture, or one of the plurality of tiles constituting the picture contains, without loss, one or more of the plurality of sub-images including the sub-image.
[0174] Accordingly, when a plurality of sub-images, a plurality of tiles, and a plurality of slices are identified within the picture, sometimes it is possible to suppress the complication of the relationship between the sub-images and the slices and the relationship between the sub-images and the tiles. Therefore, when a plurality of sub-images, a plurality of tiles, and a plurality of slices are identified within the picture, sometimes it is possible to efficiently identify the sub-images as the processing target areas in accordance with the slices and the tiles.
[0175] In addition, for example, an encoding method according to one aspect of the present invention determines one or more sub-images constituting the picture, one or more tiles constituting the picture, and one or more slices constituting the picture in accordance with both the first constraint condition and the second constraint condition, partitions and encodes the picture in accordance with the one or more sub-images, the one or more tiles, and the one or more slices, the first constraint condition being that, for each of the one or more sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the one or more slices constituting the picture, and the second constraint condition being that, for each of the one or more sub-images, one sub-image equal to the sub-image contains, without loss, one or more of the one or more tiles constituting the picture, or one of the one or more tiles constituting the picture contains, without loss, one or more of the one or more sub-images including the sub-image.
[0176] Accordingly, sometimes it is possible to suppress the complication of the relationship between sub-pictures and slices and the relationship between sub-pictures and tiles. Therefore, sometimes it is possible to efficiently determine sub-pictures as processing target regions according to slices and tiles.
[0177] In addition, for example, a decoding method according to one aspect of the present invention determines one or more sub-pictures constituting a picture, one or more tiles constituting the picture, and one or more slices constituting the picture according to both the first constraint condition and the second constraint condition, partitions and decodes the picture according to the one or more sub-pictures, the one or more tiles, and the one or more slices, the first constraint condition being that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture contains without loss one or more of the one or more slices constituting the picture, and the second constraint condition being that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture contains without loss one or more of the one or more tiles constituting the picture, or one tile of the one or more tiles constituting the picture contains without loss one or more of the one or more sub-pictures including the sub-picture.
[0178] Accordingly, sometimes it is possible to suppress the complication of the relationship between sub-pictures and slices and the relationship between sub-pictures and tiles. Therefore, sometimes it is possible to efficiently determine sub-pictures as processing target regions according to slices and tiles.
[0179] Furthermore, for example, an encoding apparatus according to one aspect of the present invention includes an input unit, a division unit, an intra prediction unit, an inter prediction unit, a loop filter unit, a transform unit, a quantization unit, an entropy encoding unit, and an output unit.
[0180] A current picture is input to the input unit. The division unit divides the current picture into a plurality of blocks.
[0181] The intra prediction unit generates a prediction signal for a current block included in the current picture using a reference picture included in the current picture. The inter prediction unit generates a prediction signal for the current block included in the current picture using a reference picture included in a reference picture different from the current picture. The loop filter unit applies a filter to a reconstructed block of the current block included in the current picture.
[0182] The transform unit transforms the prediction error between the original signal of the current block included in the current picture and the prediction signal generated by the intra prediction unit or the inter prediction unit, and generates transform coefficients. The quantization unit quantizes the transform coefficients to generate quantized coefficients. The entropy coding unit applies variable length coding to the quantized coefficients to generate a coded bit stream. Then, the coded bit stream including the quantized coefficients to which variable length coding has been applied and control information is output from the output unit.
[0183] In addition, for example, during operation, the entropy coding unit determines one or more sub-pictures constituting a picture, one or more tiles constituting the picture, and one or more slices constituting the picture in accordance with both a first constraint condition and a second constraint condition, partitions and codes the picture in accordance with the one or more sub-pictures, the one or more tiles, and the one or more slices. The first constraint condition is that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture includes, without loss, one or more slices among the one or more slices constituting the picture. The second constraint condition is that for each of the one or more sub-pictures, one sub-picture equal to the sub-picture includes, without loss, one or more tiles among the one or more tiles constituting the picture, or one tile among the one or more tiles constituting the picture includes, without loss, one or more sub-pictures including the sub-picture among the one or more sub-pictures constituting the picture.
[0184] Furthermore, for example, a decoding apparatus according to one aspect of the present invention includes an input unit, an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, a loop filter unit, and an output unit.
[0185] The coded bit stream is input to the input unit. The entropy decoding unit applies variable length decoding to the coded bit stream to derive quantized coefficients. The inverse quantization unit inverse quantizes the quantized coefficients to derive transform coefficients. The inverse transform unit inverse transforms the transform coefficients to derive a prediction error.
[0186] The intra prediction unit generates a prediction signal for a current block included in the current picture using a reference picture included in the current picture. The inter prediction unit generates a prediction signal for the current block included in the current picture using a reference picture included in a reference picture different from the current picture.
[0187] The loop filter unit applies a filter to a reconstructed block of a current block included in the current picture. Then, the current picture is output from the output unit.
[0188] In addition, for example, during operation, the entropy decoding unit determines one or more sub - pictures constituting a picture, one or more tiles constituting the picture, and one or more slices constituting the picture according to both the first constraint condition and the second constraint condition, partitions and decodes the picture according to the one or more sub - pictures, the one or more tiles, and the one or more slices. The first constraint condition is that for each of the one or more sub - pictures, one sub - picture equal to the sub - picture contains, without loss, one or more slices among the one or more slices constituting the picture. The second constraint condition is that for each of the one or more sub - pictures, one sub - picture equal to the sub - picture contains, without loss, one or more tiles among the one or more tiles constituting the picture, or one tile among the one or more tiles constituting the picture contains, without loss, one or more sub - pictures including the sub - picture among the one or more sub - pictures constituting the picture.
[0189] Moreover, these inclusive or specific forms can also be implemented by a system, a device, a method, an integrated circuit, a computer program, or a non - transitory recording medium such as a computer - readable CD - ROM, or can be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0190] [Definition of Terms]
[0191] As an example, each term can be defined as follows.
[0192] (1) Image
[0193] It is a unit of data composed of a set of pixels, composed of a picture or a block smaller than a picture, and includes still images in addition to moving images.
[0194] (2) Picture
[0195] It is a processing unit of an image composed of a set of pixels, and is sometimes called a frame or a field.
[0196] (3) Block
[0197] It is a processing unit containing a set of a specific number of pixels. As listed in the following examples, the name is not limited. In addition, the shape is not restricted. For example, of course, it includes a rectangle composed of M×N pixels, a square composed of M×M pixels, and also includes triangles, circles, and other shapes.
[0198] (Examples of Blocks)
[0199] ·Slice / Tile / Brick
[0200] ·CTU / Super - block / Elementary Segmentation Unit
[0201] ·Processing division unit of VPDU / Hardware
[0202] ·CU / Processing block unit / Prediction block unit (PU) / Orthogonal transform block unit (TU) / Unit
[0203] ·Sub-block
[0204] (4) Pixel / Sample
[0205] It is the point that is the smallest unit constituting an image, including not only the pixels at integer positions but also the pixels at fractional positions generated based on the pixels at integer positions.
[0206] (5) Pixel value / Sample value
[0207] It is the inherent value that a pixel has, of course including luminance value, chrominance difference value, grayscale of RGB, and also including depth value, or binary values of 0 and 1.
[0208] (6) Flag
[0209] In addition to 1 bit, it also includes cases of multiple bits. For example, it can also be a parameter or index of 2 bits or more. In addition, not only the binary-valued two values are used, but also multi-values using other number systems can be used.
[0210] (7) Signal
[0211] It is symbolized and encoded to transmit information, and in addition to the discretized digital signal, it also includes the analog signal taking continuous values.
[0212] (8) Stream / Bitstream
[0213] It refers to the data string of digital data or the stream of digital data. The stream / bitstream can be divided into multiple layers and composed of multiple streams in addition to 1 stream. In addition, in addition to the case of being transmitted on a single transmission path through serial communication, it also includes the case of being transmitted through packet communication in multiple transmission paths.
[0214] (9) Difference / Differential
[0215] In the case of a scalar, in addition to the simple difference (x - y), as long as the operation of difference is included, it includes the absolute value of difference (|x - y|), the square difference (x^2 - y^2), the square root of difference (√(x - y)), the weighted difference (ax - by: a and b are constants), and the offset difference (x - y + a: a is the offset).
[0216] (10) Sum
[0217] In the case of a scalar, in addition to the simple sum (x + y), any operation involving a sum is acceptable, including the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted sum (ax + by: a and b are constants), and the offset sum (x + y + a: a is the offset).
[0218] (11) Based on
[0219] This also includes cases where elements other than those that form the basis object are added. Additionally, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via intermediate results.
[0220] (12) Used, Using
[0221] This also includes cases where elements other than those that form the used object are added. Additionally, in addition to cases where a direct result is obtained, it also includes cases where a result is obtained via intermediate results.
[0222] (13) Prohibit, Forbid
[0223] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation.
[0224] (14) Limit, Restriction / Restrict / Restricted
[0225] This can also be referred to as not allowing. Additionally, not prohibiting or allowing does not necessarily imply an obligation. And as long as a part is prohibited in terms of quantity or quality, it also includes cases where it is completely prohibited.
[0226] (15) Chroma
[0227] It is an adjective represented by the notations Cb and Cr that designates one of two color difference signals associated with the primary colors for a specified sample arrangement or a single sample representation. Additionally, the term chrominance can also be used instead of the term chroma.
[0228] (16) Luma
[0229] It is an adjective represented by the notation or subscript Y or L that designates a monochrome signal associated with the primary colors for a specified sample arrangement or a single sample representation. The term luminance can also be used instead of the term luma.
[0230] [Regarding the explanations in the description]
[0231] In the drawings, the same reference numerals denote the same or similar components. In addition, the dimensions and relative positions of the components in the drawings are not necessarily drawn to a certain scale.
[0232] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments described below all represent inclusive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, relationships and orders of the steps, etc. shown in the following embodiments are examples and do not limit the meaning of the claims.
[0233] Hereinafter, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of an encoding device and a decoding device that can apply the processes and / or structures described in various aspects of the present invention. The processes and / or structures can also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processes and / or structures applied to the embodiments, any of the following can be performed, for example.
[0234] (1) Among a plurality of components of the encoding device or decoding device of the embodiment described in each aspect of the present invention, a certain component can be replaced with another component described in a certain aspect of the present invention, or they can be combined;
[0235] (2) In the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. of functions or processes performed by a part of the components of the encoding device or decoding device can also be made. For example, any function or process can be replaced with another function or process described in a certain aspect of the present invention, or they can be combined;
[0236] (3) In the method implemented by the encoding device or decoding device of the embodiment, arbitrary changes such as addition, replacement, deletion, etc. can also be made to a part of the processes included in the method. For example, any process in the method can be replaced with another process described in a certain aspect of the present invention, or they can be combined;
[0237] (4) A part of the components constituting the encoding device or decoding device of the embodiment can be combined with the components described in a certain aspect of the present invention, or can be combined with the components having a part of the functions described in a certain aspect of the present invention, or can be combined with the components performing a part of the processes implemented by the components described in the various aspects of the present invention;
[0238] (5) A component that is part of the function of the encoding device or decoding device of the embodiment, or a component that processes part of the processing of the encoding device or decoding device of the embodiment, is combined with or replaced by a component described in a certain one of the various aspects of the present invention, a component that is part of the function described in a certain one of the various aspects of the present invention, or a component that processes part of the processing described in a certain one of the various aspects of the present invention;
[0239] (6) In the method implemented by the encoding device or decoding device of the embodiment, a certain one of the multiple processes included in the method is replaced by a process described in a certain one of the various aspects of the present invention or a similar certain process, or they are combined;
[0240] (7) Part of the processes included in the method implemented by the encoding device or decoding device of the embodiment can also be combined with the processes described in any one of the various aspects of the present invention.
[0241] (8) The manner of implementing the processes and / or structures described in the various aspects of the present invention is not limited to the encoding device or decoding device of the embodiment. For example, the processes and / or structures can also be implemented in a device used for a purpose different from the motion image encoding or motion image decoding disclosed in the embodiment.
[0242] [System Structure]
[0243] Figure 1 It is a schematic diagram showing an example of the structure of the transmission system of the present embodiment.
[0244] The transmission system Trs is a system that transmits the stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, as Figure 1 shown, an encoding device 100, a network Nw, and a decoding device 200.
[0245] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding.
[0246] In addition, the original image input to the encoding device 100 before being encoded is also referred to as the original image, original signal, or original sample. Additionally, the image can be a moving image or a still image. Further, the image is a superordinate concept of sequences, pictures, blocks, etc., and is not restricted by spatial and temporal regions unless otherwise specified. Also, the image is composed of an arrangement of pixels or pixel values, and the signal or pixel value representing the image is also called a sample. Moreover, the stream can be referred to as a bitstream, encoded bitstream, compressed bitstream, or encoded signal. Furthermore, the encoding device can also be called an image encoding device or a moving image encoding device, and the encoding method of the encoding device 100 can also be called an encoding method, image encoding method, or moving image encoding method.
[0247] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw can be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Additionally, the network Nw can also be replaced by a storage medium that records the stream such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0248] The decoding device 200 generates a decoded image, such as an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method of the encoding device 100.
[0249] Additionally, the decoding device can also be called an image decoding device or a moving image decoding device, and the decoding method of the decoding device 200 can also be called a decoding method, image decoding method, or moving image decoding method.
[0250] [Data Structure]
[0251] Figure 2 is a diagram showing an example of the hierarchical structure of the data in the stream. The stream includes, for example, a video sequence. This video sequence, as shown in (a) of Figure 2 , contains a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), SEI (Supplemental Enhancement Information), and multiple pictures.
[0252] In a moving picture in which a VPS includes multiple layers, the coding parameters common to the multiple layers and the multiple layers included in the moving picture, or the coding parameters associated with each layer.
[0253] An SPS includes parameters used for a sequence, that is, coding parameters that a decoding device 200 refers to for decoding the sequence. For example, such coding parameters may also represent the width or height of a picture. Additionally, there may be multiple SPSs.
[0254] A PPS includes parameters used for a picture, that is, coding parameters that a decoding device 200 refers to for decoding each picture in the sequence. For example, such coding parameters may also include a reference value of a quantization width used in decoding the picture and a flag indicating the application of weighted prediction. In addition, there may be multiple PPSs. Furthermore, SPS and PPS are sometimes simply referred to as parameter sets.
[0255] As Figure 2 shown in (b) of, a picture may include a picture header and one or more slices. The picture header includes coding parameters that a decoding device 200 refers to for decoding the one or more slices.
[0256] As Figure 2 shown in (c) of, a slice includes a slice header and one or more tiles. The slice header includes coding parameters that a decoding device 200 refers to for decoding the one or more tiles.
[0257] As Figure 2 shown in (d) of, a tile includes one or more CTUs (Coding Tree Units).
[0258] In addition, a picture may not include slices and include a tile group instead of the slices. In this case, the tile group includes one or more tiles. Additionally, slices may be included in a tile.
[0259] A CTU is also referred to as a superblock or a basic segmentation unit. As Figure 2 shown in (e) of, such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes coding parameters that a decoding device 200 refers to for decoding the one or more CUs.
[0260] A CU may also be divided into multiple small CUs. In addition, as Figure 2As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual described later. In addition, the CU is basically the same as the PU (Prediction Unit) and the TU (Transform Unit). However, for example, in the SBT described later, it may also include a plurality of TUs smaller than the CU. In addition, the CU can also process each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. The VPDU is, for example, a fixed unit that can be processed in one stage during pipeline processing in hardware.
[0261] In addition, the stream may not have Figure 2 any part of the hierarchies shown. In addition, the order of these hierarchies can be swapped, and any one hierarchy can be replaced with another hierarchy. In addition, the picture that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is referred to as the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded. If the processing is decoding, the current picture is synonymous with the picture to be decoded. In addition, a block such as a CU or a CU that is the object of the processing performed by a device such as the encoding device 100 or the decoding device 200 at the current time point is referred to as the current block. If the processing is encoding, the current block is synonymous with the block to be encoded. If the processing is decoding, the current block is synonymous with the block to be decoded.
[0262] [Structural Slices / Tiles of Picture]
[0263] For parallel decoding of a picture, the picture may sometimes be composed of slices or tiles.
[0264] A slice is the basic encoding unit that constitutes a picture. A picture is composed of, for example, one or more slices. In addition, a slice is composed of one or more consecutive CTUs.
[0265] Figure 3This is a diagram showing an example of the structure of a slice. For example, the picture contains 11×8 CTUs and is divided into 4 slices (Slice 1 - 4). Slice 1 consists of 16 CTUs, Slice 2 consists of 21 CTUs, Slice 3 consists of 29 CTUs, and Slice 4 consists of 22 CTUs. Here, each CTU within the picture belongs to a certain slice. The shape of the slice is the shape obtained by dividing the picture horizontally. The boundary of the slice does not need to be the edge of the screen and can be anywhere among the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in the slice is, for example, the raster scan order. Additionally, the slice contains a slice header and encoded data. In the slice header, features of the slice such as the starting CTU address and slice type of the slice can also be described.
[0266] A tile is a unit of a rectangular area that makes up a picture. Numbers called TileIds can also be assigned to each tile in the raster scan order.
[0267] Figure 4 This is a diagram showing an example of the structure of a tile. For example, the picture contains 11×8 CTUs and is divided into 4 rectangular - area tiles (Tile 1 - 4). When tiles are used, the processing order of the CTUs is changed compared to the case where tiles are not used. When tiles are not used, multiple CTUs within the picture are processed, for example, in the raster scan order. When tiles are used, in each of the multiple tiles, at least 1 CTU is processed, for example, in the raster scan order. For example, as Figure 4 shown, the processing order of the multiple CTUs contained in Tile 1 is the order from the left end of the first column of Tile 1 towards the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 towards the right end of the second column of Tile 1.
[0268] In addition, sometimes one tile contains more than one slice, and sometimes one slice contains more than one tile.
[0269] Furthermore, a picture can also be composed of tile - set units. A tile - set can contain more than one tile - group and can also contain more than one tile. A picture can be composed of only one of a tile - set, a tile - group, and a tile. For example, the order of scanning multiple tiles in the raster order for each tile - set is set as the basic encoding order of the tiles. A set of one or more tiles that are consecutive in the basic encoding order within each tile - set is set as a tile - group. Such a picture can also be composed of the later - described division unit 102 (refer to Figure 7 ).
[0270] [Scalable Coding]
[0271] Figure 5 and Figure 6This is a diagram showing an example of the structure of a scalable stream.
[0272] As Figure 5 shown, the encoding device 100 can perform encoding by separately dividing multiple pictures into a certain layer among multiple layers, thereby generating a temporally / spatially scalable stream. For example, the encoding device 100 realizes the scalability where the enhancement layer exists above the base layer by encoding pictures for each layer. The encoding of such each picture is called scalable encoding. Thus, the decoding device 200 can switch the image quality of the image displayed by decoding this stream. That is, the decoding device 200 determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication band. As a result, the decoding device 200 can freely switch the same content to a low-resolution content and a high-resolution content for decoding. For example, a user of this stream is on the move and watches the moving image of this stream halfway through using a smartphone, and after returning home, watches the subsequent part of the moving image using a device such as an Internet TV. In addition, decoding devices 200 with the same or different performances are respectively assembled in the above-mentioned smartphone and device. In this case, if the device decodes to the upper layer in this stream, the user can watch a high-quality moving image after returning home. Thus, the encoding device 100 does not need to generate multiple streams with the same content but different image qualities, and can reduce the processing load.
[0273] Furthermore, the enhancement layer can also include meta information such as based on the statistical information of the image. It can also be that the decoding device 200 generates a high-quality moving image by super-resolving the pictures of the base layer based on the meta information. Super-resolution can be to improve the SNR in the same resolution or to expand the resolution. The meta information includes information for determining linear or non-linear filter coefficients used in the super-resolution process, or information for determining parameter values in the filter process, machine learning, or least squares operation used in the super-resolution process, etc.
[0274] Alternatively, the picture can also be divided into tiles, etc. according to the meaning of each object, etc. within the picture. In this case, the decoding device 200 can decode only a part of the area in the picture by selecting the tile to be decoded. Moreover, the attributes of the object (person, car, ball, etc.) and the position within the picture (coordinate position in the same picture, etc.) can be saved as meta information. In this case, the decoding device 200 can determine the position of the desired object based on the meta information and decide the tile containing the object. For example, as Figure 6 shown, a data storage structure different from the pixel data, such as SEI in HEVC, can also be used to store the meta information. This meta information represents, for example, the position, size, or color of the main object.
[0275] In addition, meta-information can also be saved in units composed of multiple pictures, such as streams, sequences, or random access units. Thus, the decoding device 200 can obtain the time when a specific person appears in the moving image, etc. By using this time and the information of the picture unit, the picture in which the target exists and the position of the target in the picture can be determined.
[0276] [Encoding device]
[0277] Next, the encoding device 100 of the embodiment will be described. Figure 7 It is a block diagram showing an example of the functional structure of the encoding device 100 of the embodiment. The encoding device 100 encodes an image in units of blocks.
[0278] As Figure 7 shown, the encoding device 100 is a device that encodes an image in units of blocks, and includes a segmentation unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filtering unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. In addition, the intra prediction unit 124 and the inter prediction unit 126 are respectively configured as part of the prediction processing unit.
[0279] [Installation example of the encoding device]
[0280] Figure 8 It is a block diagram showing an installation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, Figure 7 as shown, the multiple components of the encoding device 100 are implemented by Figure 8 the shown processor a1 and memory a2.
[0281] The processor a1 is a circuit for information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding an image. The processor a1 can also be a processor such as a CPU. In addition, the processor a1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor a1 can also play the role of Figure 7 multiple components of the encoding device 100 shown, excluding the components for storing information.
[0282] Memory a2 is a dedicated or general-purpose memory that stores information for the processor a1 to encode an image. Memory a2 can be either an electronic circuit or connected to the processor a1. Additionally, memory a2 can also be included in the processor a1. Moreover, memory a2 can also be an aggregate of multiple electronic circuits. Additionally, memory a2 can be a magnetic disk or an optical disc, etc., or can be represented as a storage or a recording medium, etc. Additionally, memory a2 can be either a non-volatile memory or a volatile memory.
[0283] For example, memory a2 can store the encoded image or can store the stream corresponding to the encoded image. Additionally, a program for the processor a1 to encode an image can also be stored in memory a2.
[0284] Additionally, for example, memory a2 can also serve as Figure 7 the component for storing information among the multiple components of the encoding device 100 as shown. Specifically, memory a2 can serve as Figure 7 the block memory 118 and the frame memory 122 as shown. More specifically, reconstructed images (specifically, reconstructed blocks and reconstructed pictures, etc.) can be stored in memory a2.
[0285] Additionally, in the encoding device 100, not all of the multiple components as shown in Figure 7 may be installed, and not all of the above-mentioned multiple processes may be performed. Figure 7 A part of the multiple components as shown in
[0286] may be included in other devices, or a part of the above-mentioned multiple processes may be performed by other devices.
[0287] [Overall Process of Encoding Processing]
[0288] Figure 9 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.
[0289] First, the segmentation unit 102 of the encoding device 100 divides the pictures included in the original image into multiple blocks of a fixed size (128×128 pixels) (step Sa_1). Then, the segmentation unit 102 selects a segmentation style for the block of the fixed size (step Sa_2). That is, the segmentation unit 102 further divides the block with a fixed size into multiple blocks constituting the selected segmentation style. Then, the encoding device 100 performs the processing of steps Sa_3 to Sa_9 for each of the multiple blocks.
[0290] A prediction processing unit composed of an intra prediction unit 124 and an inter prediction unit 126, and a prediction control unit 128 generate a prediction image of the current block (step Sa_3). In addition, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.
[0291] Next, the subtraction unit 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual is also referred to as a prediction error.
[0292] Next, the transform unit 106 and the quantization unit 108 generate a plurality of quantization coefficients by performing a transform and quantization on the prediction image (step Sa_5).
[0293] Next, the entropy encoding unit 110 generates a stream by encoding the plurality of quantization coefficients and prediction parameters related to the generation of the prediction image (specifically, entropy encoding) (step Sa_6).
[0294] Next, the inverse quantization unit 112 and the inverse transform unit 114 restore the prediction residual by performing inverse quantization and inverse transform on the plurality of quantization coefficients (step Sa_7).
[0295] Next, the addition unit 116 reconstructs the current block by adding the restored prediction residual to the prediction image (step Sa_8). Thereby, a reconstructed image is generated. In addition, the reconstructed image is also referred to as a reconstructed block, and in particular, the reconstructed image generated by the encoding device 100 is also referred to as a local decoded block or a local decoded image.
[0296] When generating the reconstructed image, the loop filter unit 120 filters the reconstructed image as needed (step Sa_9).
[0297] Then, the encoding device 100 determines whether the encoding of the entire picture has been completed (step Sa_10), and in the case where it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated.
[0298] In addition, in the above example, the encoding device 100 selects one segmentation style for a block of a fixed size and encodes each block according to the segmentation style, but each block can also be encoded according to each of a plurality of segmentation styles. In this case, the encoding device 100 can evaluate the cost for each of the plurality of segmentation styles, and for example, can select the stream obtained by encoding according to the segmentation style with the minimum cost as the finally output stream.
[0299] In addition, the processing of these steps Sa_1 to Sa_10 can be sequentially performed by the encoding device 100, a plurality of parts of these processes can be performed in parallel, or the order can be changed.
[0300] The encoding process of such an encoding device 100 uses hybrid encoding that combines predictive encoding and transform encoding. In addition, the predictive encoding is performed through an encoding loop, which is composed of a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. That is, the prediction processing unit composed of the intra prediction unit 124 and the inter prediction unit 126 forms a part of the encoding loop.
[0301] [Splitting unit]
[0302] The splitting unit 102 splits each picture included in the original image into a plurality of blocks and outputs each block to the subtraction unit 104. For example, the splitting unit 102 first splits the picture into blocks of a fixed size (e.g., 128×128 pixels). Such blocks of the fixed size are sometimes referred to as coding tree units (CTUs). And, the splitting unit 102 splits each block of the fixed size into blocks of variable size (e.g., 64×64 pixels or less) based on, for example, recursive quadtree and / or binary tree block splitting. That is, the splitting unit 102 selects the splitting pattern. Such blocks of the variable size are sometimes referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). In addition, in various installation examples, it is not necessary to distinguish between CUs, PUs, and TUs, and a part or all of the blocks in the picture can also be used as the processing unit for CUs, PUs, or TUs.
[0303] Figure 10 is a diagram showing an example of block splitting in the embodiment. In Figure 10 the solid line represents the block boundary based on quadtree block splitting, and the dashed line represents the block boundary based on binary tree block splitting.
[0304] Here, the block 10 is a square block of 128×128 pixels. This block 10 is first split into 4 square blocks of 64×64 pixels (quadtree block splitting).
[0305] The upper left 64×64 pixel square block is further vertically split into 2 rectangular blocks each composed of 32×64 pixels, and the left 32×64 pixel rectangular block is further vertically split into 2 rectangular blocks each composed of 16×64 pixels (binary tree block splitting). As a result, the upper left 64×64 pixel square block is split into 2 rectangular blocks 11 and 12 of 16×64 pixels and a rectangular block 13 of 32×64 pixels.
[0306] The upper right 64×64 pixel square block is horizontally split into 2 rectangular blocks 14 and 15 each composed of 64×32 pixels (binary tree block splitting).
[0307] The 64×64 pixel square block in the lower left is divided into 4 square blocks (quad-tree block division), each consisting of 32×32 pixels. The upper left and lower right blocks among the 4 square blocks, each consisting of 32×32 pixels, are further divided. The 32×32 pixel square block in the upper left is vertically divided into 2 rectangular blocks, each consisting of 16×32 pixels, and the right rectangular block consisting of 16×32 pixels is further horizontally divided into 2 square blocks, each consisting of 16×16 pixels (binary-tree block division). The 32×32 pixel square block in the lower right is horizontally divided into 2 rectangular blocks, each consisting of 32×16 pixels (binary-tree block division). As a result, the 64×64 pixel square block in the lower left is divided into 16 rectangular blocks 16 of 16×32 pixels, 2 square blocks 17 and 18 of 16×16 pixels each, 2 square blocks 19 and 20 of 32×32 pixels each, and 2 rectangular blocks 21 and 22 of 32×16 pixels each.
[0308] The block 23 consisting of 64×64 pixels in the lower right is not divided.
[0309] As described above, in Figure 10 , the block 10 is divided into 13 variable-size blocks 11 to 23 based on recursive quad-tree and binary-tree block division. Such a division is sometimes called QTBT (quad-tree plus binary tree) division.
[0310] In addition, in Figure 10 , 1 block is divided into 4 or 2 blocks (quad-tree or binary-tree block division), but the division is not limited to these. For example, 1 block may be divided into 3 blocks (ternary-tree division). A division including such ternary-tree division is sometimes called MBT (multi type tree) division.
[0311] Figure 11 is a diagram showing an example of the functional structure of the division unit 102. As Figure 11 shown, the division unit 102 may include a block division determination unit 102a. As an example, the block division determination unit 102a may perform the following processing.
[0312] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the above-described division pattern based on the block information. The division unit 102 divides the original image according to the division pattern and outputs one or more blocks obtained by the division to the subtraction unit 104.
[0313] In addition, the block segmentation determination unit 102a outputs, for example, parameters representing the above-described segmentation patterns to the transformation unit 106, the inverse transformation unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy encoding unit 110. The transformation unit 106 can transform the prediction residual based on this parameter, and the intra prediction unit 124 and the inter prediction unit 126 can generate a prediction image based on this parameter. In addition, the entropy encoding unit 110 can also perform entropy encoding on this parameter.
[0314] As an example, parameters related to the segmentation pattern can also be written into the stream as follows.
[0315] Figure 12 It is a diagram showing an example of the segmentation pattern. Examples of the segmentation pattern include: four-way split (QT), where the block is split into two in the horizontal and vertical directions respectively; three-way split (HT or VT), where the block is split in the same direction at a ratio of 1:2:1; two-way split (HB or VB), where the block is split in the same direction at a ratio of 1:1; and no split (NS).
[0316] In addition, in the case of four-way split and no split, the segmentation pattern does not have a block segmentation direction, and in the case of two-way split and three-way split, the segmentation pattern has segmentation direction information.
[0317] Figure 13A and Figure 13B It is a diagram showing an example of the syntax tree of the segmentation pattern. In the Figure 13A example, first, there is information indicating whether to perform segmentation (S: Split flag), and then there is information indicating whether to perform four-way split (QT: QT flag). Next, there is information indicating whether to perform three-way split or two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the segmentation direction (Ver: Vertical flag or Hor: Horizontal flag). In addition, for each of one or more blocks obtained by such segmentation based on the segmentation pattern, the same processing can be further repeatedly applied for segmentation. That is, as an example, it is also possible to recursively perform the determination of whether to perform segmentation, whether to perform four-way split, whether the segmentation method is horizontal or vertical, and whether to perform three-way split or two-way split, and encode the determination results implemented according to the Figure 13A coding order disclosed in the syntax tree shown into the stream.
[0318] In addition, in the Figure 13A syntax tree shown, these information are arranged in the order of S, QT, TT, Ver, but they can also be arranged in the order of S, QT, Ver, BT. That is, in the Figure 13BIn the example, first, there is information indicating whether to perform splitting (S: Split flag), then, there is information indicating whether to perform four-way splitting (QT: QT flag). Next, there is information indicating the splitting direction (Ver: Vertical flag or Hor: Horizontal flag), and finally, there is information indicating whether to perform two-way splitting or three-way splitting (BT: BT flag or TT: TT flag).
[0319] In addition, the splitting patterns described here are just examples. Splitting patterns other than the described ones can be used, or only a part of the described splitting patterns can be used.
[0320] [Subtraction unit]
[0321] The subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units that are input and split by the splitting unit 102. That is, the subtraction unit 104 calculates the prediction residual of the current block. And the subtraction unit 104 outputs the calculated prediction residual to the transformation unit 106.
[0322] The original image is an input signal of the encoding device 100, for example, a signal representing images of each picture constituting a moving image (such as a luminance (luma) signal and two chrominance (chroma) signals).
[0323] [Transformation unit]
[0324] The transformation unit 106 transforms the prediction residual in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the quantization unit 108. Specifically, the transformation unit 106, for example, performs a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the prediction residual in the spatial domain.
[0325] In addition, the transformation unit 106 can also adaptively select a transformation type from multiple transformation types, use a transform basis function corresponding to the selected transformation type, and transform the prediction residual into transform coefficients. Such a transformation is sometimes called EMT (explicit multiple core transform, multi-core transform) or AMT (adaptive multiple transform, adaptive multi-transform). In addition, the transform basis function is sometimes simply referred to as the basis.
[0326] Multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. In addition, these transformation types can be described as DCT2, DCT5, DCT8, DST1, and DST7 respectively. Figure 14is a table representing transform basis functions corresponding to respective transform types. In Figure 14 , N represents the number of input pixels. The selection of a transform type from among these multiple transform types can depend on, for example, the type of prediction (intra prediction, inter prediction, etc.) or the intra prediction mode.
[0327] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0328] Furthermore, the transform unit 106 can also perform a re - transform on the transform coefficients (i.e., the transform result). Such a re - transform may be a case of what is called AST (adaptive secondary transform) or NSST (non - separable secondary transform). For example, the transform unit 106 performs a re - transform on each sub - block (e.g., a 4×4 pixel sub - block) included in a block of transform coefficients corresponding to the intra - prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are generally signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0329] In the transform unit 106, a separable transform and a non - separable transform can also be applied. A separable transform is a method in which multiple transforms are performed separately in each direction according to the number of input dimensions, and a non - separable transform is a method in which when the input is multi - dimensional, two or more dimensions are regarded as one dimension and transformed together.
[0330] For example, as an example of a non - separable transform, when the input is a 4×4 pixel block, it can be regarded as a permutation having 16 elements, and a transform process is performed on this permutation with a 16×16 transform matrix.
[0331] Furthermore, in a further example of a non - separable transform, after regarding a 4×4 pixel input block as a permutation having 16 elements, a transform (Hypercube Givens Transform) in which Givens rotations are performed on this permutation multiple times can also be performed.
[0332] In the transformation in the transformation unit 106, the type of transform of the transform basis function to be transformed into the frequency domain can be switched according to the region within the CU. As an example, there is SVT (Spatially Varying Transform).
[0333] Figure 15 It is a diagram showing an example of SVT.
[0334] In SVT, as Figure 15 shown, the CU is bisected in the horizontal or vertical direction, and only one of the regions is transformed into the frequency domain. The transform type can be set for each region. For example, DST7 and DCT8 are used. For example, for the region at position 0 among the two regions obtained by bisecting the CU in the vertical direction, DST7 and DCT8 can be used. Or, for the region at position 1 among the two regions, DST7 is used. Similarly, for the region at position 0 among the two regions obtained by bisecting the CU in the horizontal direction, DST7 and DCT8 are used. Or, for the region at position 1 among the two regions, DST7 is used. In such an Figure 15 example shown, only one of the two regions within the CU is transformed, and the other is not transformed. However, it is also possible to transform the two regions separately. In addition, the splitting method is not limited to bisection, and can also be quartering. Furthermore, it can be more flexible, encoding the information indicating the splitting method and performing signaling in the same way as CU splitting. In addition, SVT is sometimes also referred to as SBT (Sub-block Transform).
[0335] The aforementioned AMT and EMT can also be referred to as MTS (Multiple Transform Selection). In the case of applying MTS, transform types such as DST7 or DCT8 can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. On the other hand, there is a process called IMTS (Implicit MTS) which is a process of selecting the transform type used in the orthogonal transform based on the shape of the CU without encoding the index information. In the case of applying IMTS, for example, if the shape of the CU is rectangular, DST7 is used on the short side of the rectangle and DCT2 is used on the long side for orthogonal transform respectively. Additionally, for example, in the case where the shape of the CU is square, if MTS is effective in the sequence, DCT2 is used for orthogonal transform, and if MTS is ineffective, DST7 is used for orthogonal transform. DCT2 and DST7 are just examples, other transform types can be used, and different combinations of the used transform types can also be set. IMTS can be used only in the blocks of intra prediction, or can be used together in the blocks of intra prediction and inter prediction.
[0336] As described above, as the selection process of selectively switching the transform type used in the orthogonal transform, three processes, namely MTS, SBT, and IMTS, have been described. However, all three selection processes can be effective, or only some of the selection processes can be selectively made effective. Whether each selection process is effective can be identified by flag information in headers such as SPS. For example, if all three selection processes are effective, one is selected from the three selection processes for orthogonal transform in CU units. Additionally, as long as the selection process of selectively switching the transform type can achieve at least one of the following four functions [1] to [4], a selection process different from the above three selection processes can be used, or the above three selection processes can be replaced with other processes respectively. Function [1] is the function of performing orthogonal transform on the entire range within the CU and encoding the information indicating the transform type used in the transform. Function [2] is the function of performing orthogonal transform on the entire range of the CU, determining the transform type based on a specified rule without encoding the information indicating the transform type. Function [3] is the function of performing orthogonal transform on a partial region of the CU and encoding the information indicating the transform type used in the transform. Function [4] is the function of performing orthogonal transform on a partial region of the CU, not encoding the information indicating the transform type used in the transform, and determining the transform type based on a specified rule, etc.
[0337] In addition, the presence or absence of the application of MTS, IMTS, and SBT respectively can also be determined for each processing unit. For example, it can be determined for each of the sequence unit, picture unit, tile unit, slice unit, CTU unit, or CU unit.
[0338] In addition, the tool for selectively switching the transformation type in the present invention can also be renamed as a method for adaptively selecting the basis used in the transformation process, a selection process, or a process for selecting a basis. In addition, the tool for selectively switching the transformation type can also be renamed as a mode for adaptively selecting the transformation type.
[0339] Figure 16 It is a flowchart showing an example of the process performed by the transformation unit 106.
[0340] For example, the transformation unit 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined to perform an orthogonal transformation (Yes in step St_1), the transformation unit 106 selects the transformation type used for the orthogonal transformation from among multiple transformation types (step St_2). Next, the transformation unit 106 performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Then, the transformation unit 106 outputs the information indicating the selected transformation type to the entropy encoding unit 110 to encode this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transformation (No in step St_1), the transformation unit 106 outputs the information indicating that no orthogonal transformation is performed to the entropy encoding unit 110 to encode this information (step St_5). In addition, the determination of whether to perform an orthogonal transformation in step St_1 can be made, for example, based on the size of the transformation block, the prediction mode applied to the CU, etc. In addition, the information indicating the transformation type used for the orthogonal transformation may not be encoded, and an orthogonal transformation may be performed using a pre-specified transformation type.
[0341] Figure 17 It is a flowchart showing another example of the process performed by the transformation unit 106. In addition, Figure 17 The example shown is the same as the example shown in Figure 16 and is an example of an orthogonal transformation in the case of applying a method for selectively switching the transformation type used for the orthogonal transformation.
[0342] As an example, the first transformation type group may include DCT2, DST7, and DCT8. In addition, as an example, the second transformation type group may include DCT2. In addition, the transformation types included in the first transformation type group and the second transformation type group may be partially repeated or may be all different transformation types.
[0343] Specifically, the transform unit 106 determines whether the transform size is equal to or less than a specified value (step Su_1). Here, when it is determined that the size is equal to or less than the specified value (yes in step Su_1), the transform unit 106 orthogonally transforms the prediction residual of the current block using the transform types included in the first transform type group (step Su_2). Then, the transform unit 106 encodes the information indicating which one of the one or more transform types included in the first transform type group is used by outputting the information to the entropy encoding unit 110 (step Su_3). On the other hand, when it is determined that the transform size is not equal to or less than the specified value (no in step Su_1), the transform unit 106 orthogonally transforms the prediction residual of the current block using the second transform type group (step Su_4).
[0344] In step Su_3, the information indicating the transform type used for the orthogonal transform may be information representing a combination of the transform type applied to the vertical direction of the current block and the transform type applied to the horizontal direction. In addition, the first transform type group may include only one transform type, and the information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and the information indicating the transform type used in the orthogonal transform among the one or more transform types included in the second transform type group may also be encoded.
[0345] In addition, the transform type may be determined based only on the transform size. Furthermore, if the process of determining the transform type used for the orthogonal transform is based on the transform size, it is not limited to the determination of whether the transform size is equal to or less than the specified value.
[0346] [Quantization Unit]
[0347] The quantization unit 108 quantizes the transform coefficients output from the transform unit 106. Specifically, the quantization unit 108 scans the multiple transform coefficients of the current block in a specified scan order, and quantizes the transform coefficient based on the quantization parameter (QP) corresponding to the scanned transform coefficient. Then, the quantization unit 108 outputs the quantized multiple transform coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy encoding unit 110 and the inverse quantization unit 112.
[0348] The specified scan order is the order for quantization / inverse quantization of the transform coefficients. For example, the specified scan order is defined by the ascending order of frequencies (from low frequency to high frequency) or the descending order (from high frequency to low frequency).
[0349] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. That is, if the value of the quantization parameter increases, the error of the quantization coefficient (quantization error) increases.
[0350] In addition, in quantization, a quantization matrix is sometimes used. For example, multiple quantization matrices are sometimes used corresponding to frequency transformation sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance. In addition, quantization refers to digitizing values sampled at a predetermined interval by associating them with predetermined levels, and in this technical field, expressions such as rounding, truncation, or scaling are sometimes used.
[0351] As methods of using a quantization matrix, there are a method of using a quantization matrix directly set on the encoding device 100 side and a method of using a default quantization matrix (default matrix). On the encoding device 100 side, by directly setting the quantization matrix, a quantization matrix corresponding to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix. In addition, instead of directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix used in the quantization of the current block can be generated based on the default quantization matrix or the encoded quantization matrix.
[0352] On the other hand, there is also a method of performing quantization in such a way that the coefficients of high-frequency components and the coefficients of low-frequency components are the same without using a quantization matrix. In addition, this method is equivalent to a method of using a quantization matrix (flat matrix) in which all coefficients are the same value.
[0353] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level.
[0354] When using a quantization matrix, the quantization unit 108 scales, for example, the quantization width obtained based on quantization parameters and the like for each transform coefficient using the value of the quantization matrix. The quantization process without using a quantization matrix can also be a process of quantizing the transform coefficient based on the quantization width obtained based on quantization parameters and the like. In addition, in the quantization process without using a quantization matrix, a predetermined value common to all transform coefficients within the block can be multiplied by the quantization width.
[0355] Figure 18 It is a block diagram showing an example of the functional structure of the quantization unit 108.
[0356] The quantization unit 108 includes, for example, a differential quantization parameter generation unit 108a, a prediction quantization parameter generation unit 108b, a quantization parameter generation unit 108c, a quantization parameter storage unit 108d, and a quantization processing unit 108e.
[0357] Figure 19 It is a flowchart showing an example of quantization performed by the quantization unit 108.
[0358] As an example, the quantization unit 108 can be based onFigure 19 The flowchart shown performs quantization for each CU. Specifically, the quantization parameter generation unit 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (yes in step Sv_1), the quantization parameter generation unit 108c generates quantization parameters for the current block (step Sv_2) and stores the quantization parameters in the quantization parameter storage unit 108d (step Sv_3).
[0359] Next, the quantization processing unit 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). Then, the predicted quantization parameter generation unit 108b obtains quantization parameters for a processing unit different from the current block from the quantization parameter storage unit 108d (step Sv_5). The predicted quantization parameter generation unit 108b generates predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generation unit 108a calculates the difference between the quantization parameters for the current block generated by the quantization parameter generation unit 108c and the predicted quantization parameters for the current block generated by the predicted quantization parameter generation unit 108b (step Sv_7). By calculating this difference, differential quantization parameters are generated. The differential quantization parameter generation unit 108a outputs the differential quantization parameters to the entropy encoding unit 110, thereby encoding the differential quantization parameters (step Sv_8).
[0360] In addition, the differential quantization parameters can also be encoded at the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameters can be generated using the initial values of the quantization parameters and the differential quantization parameters.
[0361] In addition, the quantization unit 108 can include multiple quantizers, and dependent quantization that quantizes the transform coefficients using a quantization method selected from multiple quantization methods can also be applied.
[0362] [Entropy Encoding Unit]
[0363] Figure 20 is a block diagram showing an example of the functional structure of the entropy encoding unit 110.
[0364] The entropy encoding unit 110 performs entropy encoding on the quantized coefficients input from the quantization unit 108 and the prediction parameters input from the prediction parameter generation unit 130, thereby generating a stream. In this entropy encoding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used. Specifically, the entropy encoding unit 110 includes, for example, a binarization unit 110a, a context control unit 110b, and a binary arithmetic coding unit 110c. The binarization unit 110a performs binarization that transforms multi-value signals such as quantized coefficients and prediction parameters into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, etc. The context control unit 110b derives a context value corresponding to the characteristics of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. In the method for deriving this context value, for example, there are bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic coding unit 110c performs arithmetic coding on the binarized signal using the derived context value.
[0365] Figure 21 It is a diagram showing the process of CABAC in the entropy encoding unit 110.
[0366] First, in the CABAC in the entropy encoding unit 110, initialization is performed. In this initialization, initialization in the binary arithmetic coding unit 110c and setting of the initial context value are performed. Then, the binarization unit 110a and the binary arithmetic coding unit 110c sequentially perform binarization and arithmetic coding on the multiple quantized coefficients of the CTU, respectively. At this time, the context control unit 110b updates the context value each time arithmetic coding is performed. Then, the context control unit 110b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.
[0367] [Inverse quantization unit]
[0368] The inverse quantization unit 112 performs inverse quantization on the quantized coefficients input from the quantization unit 108. Specifically, the inverse quantization unit 112 performs inverse quantization on the quantized coefficients of the current block in a prescribed scanning order. And the inverse quantization unit 112 outputs the inverse quantized transform coefficients of the current block to the inverse transform unit 114.
[0369] [Inverse transform unit]
[0370] The inverse transform unit 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantization unit 112. Specifically, the inverse transform unit 114 restores the prediction residual of the current block by performing an inverse transform on the transform coefficients corresponding to the transform of the transform unit 106. Then, the inverse transform unit 114 outputs the restored prediction residual to the addition unit 116.
[0371] In addition, since information is usually lost through quantization in the restored prediction residual, it does not match the prediction error calculated by the subtraction unit 104. That is, the restored prediction residual usually contains a quantization error.
[0372] [Addition unit]
[0373] The addition unit 116 reconstructs the current block by adding the prediction residual input from the inverse transform unit 114 to the predicted image input from the prediction control unit 128. As a result, a reconstructed image is generated. Then, the addition unit 116 outputs the reconstructed image to the block memory 118 and the loop filter unit 120.
[0374] [Block memory]
[0375] The block memory 118 is, for example, a storage unit for storing blocks referred to in intra prediction and blocks within the current picture. Specifically, the block memory 118 stores the reconstructed image output from the addition unit 116.
[0376] [Frame memory]
[0377] The frame memory 122 is, for example, a storage unit for storing reference pictures used in inter prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter unit 120.
[0378] [Loop filter unit]
[0379] The loop filter unit 120 performs a loop filter process on the reconstructed image output from the addition unit 116 and outputs the reconstructed image after the filter process to the frame memory 122. Loop filtering refers to filtering used within the coding loop (in-loop filtering), and includes, for example, adaptive loop filtering (ALF), deblocking filtering (DF or DBF), and sample adaptive offset (SAO).
[0380] Figure 22 It is a block diagram showing an example of the functional structure of the loop filter unit 120.
[0381] For example Figure 22As shown in the figure, the loop filter unit 120 includes a deblocking filter processing unit 120a, an SAO processing unit 120b, and an ALF processing unit 120c. The deblocking filter processing unit 120a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 120b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 120c applies the above-described ALF processing to the reconstructed image after the SAO processing. Details of the ALF and the deblocking filter will be described later. The SAO processing is a process for improving the image quality by reducing ringing (a phenomenon in which pixel values around the edge are deformed in a fluctuating manner) and correcting the deviation of pixel values. In this SAO processing, for example, there are edge offset processing and band offset processing. In addition, the loop filter unit 120 may not include Figure 22 all the processing units disclosed, or may include only some of the processing units. In addition, the loop filter unit 120 may also be structured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 22 .
[0382] [Loop Filter Unit > Adaptive Loop Filter]
[0383] In the ALF, a least squares error filter for removing coding distortion is used. For example, for each 2×2 pixel sub-block within the current block, one filter selected from a plurality of filters based on the direction and activity of the gradient based on locality is used.
[0384] Specifically, first, sub-blocks (for example, 2×2 pixel sub-blocks) are classified into a plurality of classes (for example, 15 or 25 classes). The classification of the sub-blocks is performed, for example, based on the direction and activity of the gradient. In a specific example, using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4), the classification value C (for example, C = 5D + A) is calculated. And based on the classification value C, the sub-blocks are classified into a plurality of classes.
[0385] The gradient direction value D is derived, for example, by comparing the gradients in a plurality of directions (for example, horizontal, vertical, and two diagonal directions). In addition, the gradient activity value A is derived, for example, by adding the gradients in a plurality of directions and quantifying the addition result.
[0386] Based on the result of such classification, the filter to be used for the sub-block is determined from among a plurality of filters.
[0387] As the shape of the filter used in the ALF, for example, a circularly symmetric shape is used. Figures 23A to 23C is a diagram showing a plurality of examples of the shape of the filter used in the ALF. Figure 23A represents a 5×5 diamond-shaped filter,Figure 23B represents a 7×7 diamond-shaped filter, Figure 23C represents a 9×9 diamond-shaped filter. Information indicating the shape of the filter is usually signaled at the picture level. Additionally, the signaling of information indicating the shape of the filter need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0388] The on / off of ALF can also be determined, for example, at the picture level or CU level. For example, for luminance, it can be determined at the CU level whether to use ALF, and for chrominance difference, it can be determined at the picture level whether to use ALF. Information indicating the on / off of ALF is usually signaled at the picture level or CU level. Additionally, the signaling of information indicating the on / off of ALF need not be limited to the picture level or CU level and can also be at other levels (e.g., sequence level, slice level, tile level, or CTU level).
[0389] Additionally, as described above, one filter is selected from multiple filters and ALF processing is applied to the sub-block. For each of these multiple filters (e.g., up to 15 or 25 filters), the set of coefficients composed of the multiple coefficients used in the filter is usually signaled at the picture level. Additionally, the signaling of the set of coefficients need not be limited to the picture level and can also be at other levels (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0390] [Loop Filter > Cross Component Adaptive Loop Filter (or Cross Component Adaptive Loop Filter)]
[0391] Figure 23D is a diagram showing an example where the Y sample (the first component) is used for the CCALF of Cb and the CCALF of Cr (multiple components different from the first component). Figure 23E is a diagram showing a diamond-shaped filter.
[0392] One example of CC-ALF is by using a linear diamond-shaped filter ( Figure 23D , Figure 23E) It operates on the luminance channels of various color difference components. For example, filter coefficients are sent in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled with variable block sizes and is signaled by a context-encoded flag received for each block of samples. The block size and the CC-ALF enable flag are received at the slice level of each color difference component. The syntax and semantics of CC-ALF are provided in the Appendix. In this document, block sizes of 16x16, 32x32, 64x64, and 128x128 are supported (in color difference samples).
[0393] [Loop Filter> Joint Chroma Cross Component Adaptive Loop Filter]
[0394] Figure 23F is a diagram showing an example of JC-CCALF. Figure 23G is a diagram showing an example of the weight_index candidates of JC-CCALF.
[0395] One example of JC-CCALF uses only one CCALF filter, generates one CCALF filter output as the color difference adjustment signal for only one color component, and applies a properly weighted version of the same color difference adjustment signal to other color components. In this way, the complexity of the existing CCALF is approximately halved.
[0396] The weight value is encoded into a sign flag and a weight index. The weight index (denoted as weight_index) is encoded as 3 bits, specifying the magnitude of the JC-CCALF weight JcCcWeight. It cannot be the same as 0. The magnitude of JcCcWeight is determined as follows.
[0397] · When weight_index is 4 or less, JcCcWeight is equal to weight_index >> 2.
[0398] · In other cases, JcCcWeight is equal to 4 / (weight_index - 4).
[0399] The block-level on / off control for ALF filtering of Cb and Cr is separate. This is the same as CCALF, and two individual sets of block-level on / off control flags are encoded. Here, different from CCALF, the on / off control block sizes for Cb and Cr are the same, so only one block size variable is encoded.
[0400] [Loop Filtering Section > Deblocking Filter]
[0401] In the deblocking filter processing, the loop filtering section 120 reduces the distortion generated at the block boundary by filtering the block boundary of the reconstructed image.
[0402] Figure 24 is a block diagram showing an example of the detailed structure of the deblocking filter processing section 120a.
[0403] The deblocking filter processing section 120a includes, for example, a boundary determination section 1201, a filtering determination section 1203, a filtering processing section 1205, a processing determination section 1208, a filtering characteristic determination section 1207, and switches 1202, 1204, and 1206.
[0404] The boundary determination section 1201 determines whether there are pixels (i.e., target pixels) for which deblocking filter processing is to be performed near the block boundary. Then, the boundary determination section 1201 outputs its determination result to the switches 1202 and the processing determination section 1208.
[0405] When it is determined by the boundary determination section 1201 that target pixels exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1204. On the contrary, when it is determined by the boundary determination section 1201 that target pixels do not exist near the block boundary, the switch 1202 outputs the image before filtering processing to the switch 1206. In addition, the image before filtering processing is an image composed of target pixels and at least one surrounding pixel located around the target pixel.
[0406] The filtering determination section 1203 determines whether to perform deblocking filter processing on the target pixels based on the pixel values of at least one surrounding pixel located around the target pixel. Then, the filtering determination section 1203 outputs the determination result to the switch 1204 and the processing determination section 1208.
[0407] When it is determined by the filtering determination section 1203 that deblocking filter processing is to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the filtering processing section 1205. On the contrary, when it is determined by the filtering determination section 1203 that deblocking filter processing is not to be performed on the target pixels, the switch 1204 outputs the image before filtering processing obtained via the switch 1202 to the switch 1206.
[0408] When the image before filtering processing is obtained via the switches 1202 and 1204, the filtering processing section 1205 performs deblocking filter processing with the filtering characteristics determined by the filtering characteristic determination section 1207 on the target pixels. Then, the filtering processing section 1205 outputs the pixels after the filtering processing to the switch 1206.
[0409] Under the control of the processing determination unit 1208, the switch 1206 selectively outputs pixels that have not been deblocking-filtered and pixels that have been deblocking-filtered by the filtering processing unit 1205.
[0410] The processing determination unit 1208 controls the switch 1206 based on the respective determination results of the boundary determination unit 1201 and the filtering determination unit 1203. That is, when the boundary determination unit 1201 determines that the target pixel exists near the block boundary and the filtering determination unit 1203 determines that deblocking filtering is to be performed on the target pixel, the processing determination unit 1208 outputs the deblocking-filtered pixels from the switch 1206. In addition, in other cases than the above, the processing determination unit 1208 outputs the pixels that have not been deblocked / filtered from the switch 1206. By repeatedly outputting such pixels, the filtered image is output from the switch 1206. In addition, Figure 24 The structure shown is an example of the structure in the deblocking filtering unit 120a, and the deblocking filtering unit 120a may have other structures.
[0411] Figure 25 is a diagram showing an example of deblocking filtering having a filtering characteristic symmetric with respect to the block boundary.
[0412] In deblocking filtering, for example, using the pixel value and the quantization parameter, one of two deblocking filters with different characteristics, namely, a strong filter and a weak filter, is selected. In the strong filter, as Figure 25 shown, when there are pixels p0 to p2 and pixels q0 to q2 across the block boundary, the pixel values of the pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operations shown in the following equations.
[0413] q'0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8
[0414] q'1 = (p0 + q0 + q1 + q2 + 2) / 4
[0415] q'2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8
[0416] In addition, in the above equations, p0 to p2 and q0 to q2 are the pixel values of the pixels p0 to p2 and the pixels q0 to q2, respectively. In addition, q3 is the pixel value of the pixel q3 adjacent to the pixel q2 on the side opposite to the block boundary. In addition, on the right side of each of the above equations, the coefficients multiplied by the pixel values of the respective pixels used in the deblocking filtering are filtering coefficients.
[0417] Furthermore, in the deblocking filter processing, the clipping process may also be performed in such a way that the pixel value after the operation does not change when it exceeds the threshold value. In this clipping process, the threshold value determined according to the quantization parameter is used to clip the pixel value after the operation based on the above formula to "the pixel value before the operation ± 2 × the threshold value". Thereby, excessive smoothing can be prevented.
[0418] Figure 26 FIG. is an example of a block boundary for explaining the deblocking filter processing. Figure 27 FIG. is an example showing the BS value.
[0419] The block boundary for the deblocking filter processing is, for example, Figure 26 the boundary of the CU, PU, or TU of an 8×8 pixel block shown in the figure. The deblocking filter processing is performed, for example, in units of 4 rows or 4 columns. First, for Figure 26 the blocks P and Q shown in the figure, the Bs (Boundary Strength) value is determined as Figure 27 shown.
[0420] According to Figure 27 the Bs value, it can be determined whether to perform deblocking filter processing with different strengths even for block boundaries belonging to the same image. When the Bs value is 2, deblocking filter processing for the chrominance signal is performed. When the Bs value is 1 or more and satisfies the specified conditions, deblocking filter processing for the luminance signal is performed. In addition, the determination condition of the Bs value is not limited to Figure 27 the conditions shown in the figure, and can also be determined based on other parameters.
[0421] [Prediction unit (intra prediction unit / inter prediction unit / prediction control unit)]
[0422] Figure 28 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the encoding device 100. In addition, as an example, the prediction unit is composed of all or part of the constituent elements of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.
[0423] The prediction unit generates a prediction image of the current block (step Sb_1). In addition, in the prediction image, there are, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). Specifically, the prediction unit generates a prediction image of the current block using the reconstructed image that has already been obtained by generating a prediction image of other blocks, generating a prediction residual, generating quantization coefficients, restoring the prediction residual, and adding the prediction images.
[0424] The reconstructed image can be, for example, an image referring to a reference picture, or an image including the current block, i.e., an encoded block within the current picture (i.e., the other blocks described above). The encoded blocks within the current picture are, for example, adjacent blocks of the current block.
[0425] Figure 29 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device 100.
[0426] The prediction unit generates a prediction image by the first method (step Sc_1a), generates a prediction image by the second method (step Sc_1b), and generates a prediction image by the third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and can be, for example, an inter-frame prediction method, an intra-frame prediction method, and other prediction methods, respectively. In such prediction methods, the above-described reconstructed image can also be used.
[0427] Next, the prediction unit evaluates the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit evaluates these prediction images by calculating a cost C for each of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c and comparing the costs C of these prediction images. In addition, the cost C is calculated by an equation of an R-D optimization model, such as C = D + λ × R. In this equation, D is the encoding distortion of the prediction image and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. In addition, R is the bit rate of the stream. In addition, λ is, for example, a Lagrange undetermined multiplier.
[0428] Next, the prediction unit selects one of the prediction images respectively generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects a method or mode for obtaining the final prediction image. For example, the prediction unit selects the prediction image with the minimum cost C based on the cost C calculated for these prediction images. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 can also be performed based on the parameters used in the encoding process. The encoding device 100 can signal an information signal for determining the selected prediction image, method, or mode as a stream. This information can be, for example, a flag or the like. Thus, the decoding device 200 can generate a prediction image based on this information in the manner or mode selected in the encoding device 100. In addition, in Figure 29 In the example shown, after generating the prediction images by each method, the prediction unit selects any one of the prediction images. However, before generating these prediction images, the prediction unit can select a method or mode based on the parameters used in the above-described encoding process, and can generate a prediction image according to this method or mode.
[0429] For example, the first mode and the second mode are intra prediction and inter prediction respectively, and the prediction unit may select the final prediction image for the current block from the prediction images generated according to these prediction modes.
[0430] Figure 30 It is a flowchart showing another example of the processing performed by the prediction unit of the encoding device 100.
[0431] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a) and generates a prediction image by inter prediction (step Sd_1b). In addition, the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.
[0432] Next, the prediction unit evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above-mentioned cost C may also be used in this evaluation. Then, the prediction unit may select the prediction image that calculates the minimum cost C from the intra prediction image and the inter prediction image as the final prediction image for the current block (step Sd_3). That is, the prediction mode or pattern for generating the prediction image of the current block is selected.
[0433] [Intra Prediction Unit]
[0434] The intra prediction unit 124 performs intra prediction (also referred to as intra-picture prediction) of the current block with reference to the block in the current picture stored in the block memory 118, thereby generating a prediction image of the current block (i.e., an intra prediction image). Specifically, the intra prediction unit 124 generates an intra prediction image by performing intra prediction with reference to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction image to the prediction control unit 128.
[0435] For example, the intra prediction unit 124 performs intra prediction using one of a plurality of predefined intra prediction modes. The plurality of intra prediction modes usually include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0436] One or more non-directional prediction modes include, for example, the Planar (plane) prediction mode and the DC prediction mode defined by the H.265 / HEVC standard.
[0437] The plurality of directional prediction modes include, for example, 33-direction prediction modes defined by the H.265 / HEVC standard. In addition, the plurality of directional prediction modes may also include 32-direction prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). Figure 31This is a diagram showing all 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. The solid arrows represent 33 directions defined by the H.265 / HEVC standard, and the dashed arrows represent 32 additional directions (the 2 non-directional prediction modes are not shown in Figure 31 .
[0438] In various installation examples, in the intra prediction of chrominance blocks, the luminance blocks can also be referred to. That is, the chrominance components of the current block can also be predicted based on the luminance component of the current block. Such intra prediction is sometimes referred to as CCLM (cross-component linear model) prediction. The intra prediction mode of the chrominance block that refers to the luminance block in this way (for example, called the CCLM mode) can also be added as one of the intra prediction modes of the chrominance block.
[0439] The intra prediction unit 124 can also correct the pixel value after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions. The intra prediction accompanied by such correction is sometimes referred to as PDPC (position dependent intraprediction combination). The information indicating whether PDPC is used (for example, called the PDPC flag) is usually signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level and can also be at other levels (for example, sequence level, picture level, slice level, tile level, or CTU level).
[0440] Figure 32 This is a flowchart showing an example of the processing performed by the intra prediction unit 124.
[0441] The intra prediction unit 124 selects one intra prediction mode from multiple intra prediction modes (step Sw_1). Then, the intra prediction unit 124 generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit 124 determines the MPM (Most Probable Modes) (step Sw_3). The MPM consists of, for example, 6 intra prediction modes. Two of these 6 intra prediction modes can be the Planar prediction mode and the DC prediction mode, and the remaining 4 modes can be directional prediction modes. Then, the intra prediction unit 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0442] Here, when it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). Further, the MPM flag set to 1 and the information indicating the intra prediction mode are encoded as prediction parameters by the entropy encoding unit 110, respectively.
[0443] On the other hand, when it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction unit 124 does not set the MPM flag. Then, the intra prediction unit 124 generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). Further, the MPM flag set to 0 and the information indicating the intra prediction mode are encoded as prediction parameters by the entropy encoding unit 110, respectively. The information indicating the intra prediction mode represents any value from 0 to 60, for example.
[0444] [Inter - prediction unit]
[0445] The inter - prediction unit 126 performs inter - prediction (also called inter - picture prediction) of the current block with reference to a reference picture different from the current picture stored in the frame memory 122, thereby generating a predicted image (inter - prediction image). The inter - prediction is performed in units of the current block or a current sub - block within the current block. A sub - block is included in a block and is a unit smaller than the block. The size of the sub - block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub - block can also be switched in units of slices, tiles, or pictures, etc.
[0446] For example, the inter - prediction unit 126 performs motion search (motion estimation) within the reference picture for the current block or the current sub - block to find a reference block or sub - block that most matches the current block or the current sub - block. And the inter - prediction unit 126 obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or sub - block to the current block or sub - block. The inter - prediction unit 126 performs motion compensation (or motion prediction) based on this motion information, thereby generating an inter - prediction image of the current block or sub - block. And the inter - prediction unit 126 outputs the generated inter - prediction image to the prediction control unit 128.
[0447] The motion information used in motion compensation is signaled as an inter - prediction image in various forms. For example, a motion vector can also be signaled. As another example, the difference between a motion vector and a predicted motion vector (motion vector predictor) can also be signaled.
[0448] [List of reference pictures]
[0449] Figure 33 is a diagram showing an example of each reference picture, Figure 34 is a conceptual diagram showing an example of the list of reference pictures. The list of reference pictures is a list showing one or more reference pictures stored in the frame memory 122. In addition, in Figure 33 , the rectangle represents a picture, the arrow represents the reference relationship of the pictures, the horizontal axis represents time, I, P, and B in the rectangle represent an intra-frame predicted picture, a single predicted picture, and a bi-predicted picture respectively, and the numbers in the rectangle represent the decoding order. As Figure 33 shown, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As Figure 34 shown, the list of reference pictures is a list showing candidates for reference pictures. For example, one picture (or slice) can have more than one list of reference pictures. For example, if the current picture is a single predicted picture, one list of reference pictures is used, and if the current picture is a bi-predicted picture, two lists of reference pictures are used. In the examples of Figure 33 and Figure 34 , the picture B3 as the current picture currPic has two lists of reference pictures, namely the L0 list and the L1 list. When the current picture currPic is the picture B3, the candidates for the reference pictures of this current picture currPic are I0, P1, and B2, and each list of reference pictures (i.e., the L0 list and the L1 list) represents these pictures. The inter-frame prediction unit 126 or the prediction control unit 128 specifies which picture in each list of reference pictures is actually referred to through the reference picture index refidxLx. In Figure 34 , the reference pictures P1 and B2 are specified through the reference picture indices refIdxL0 and refIdxL1.
[0450] Such a list of reference pictures can be generated in units of sequence, picture, slice, tile, CTU, or CU. In addition, the reference picture indices of the reference pictures shown in the list of reference pictures that are referred to in inter-frame prediction can be encoded at the sequence level, picture level, slice level, tile level, CTU level, or CU level. In addition, in multiple inter-frame prediction modes, a common list of reference pictures can also be used.
[0451] [Basic process of inter-frame prediction]
[0452] Figure 35 is a flowchart showing the basic process of inter-frame prediction.
[0453] The inter-frame prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the prediction image as a prediction residual (step Se_4).
[0454] Here, in the generation of the prediction image, the inter-frame prediction unit 126 generates the prediction image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2). The selection of the candidate MV is performed, for example, by the inter-frame prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. In addition, in the candidate MV list, the previously derived MV can be added as a candidate MV. Further, in the derivation of the MV, the inter-frame prediction unit 126 may also determine the MV of the current block by further selecting at least one candidate MV from at least one candidate MV and determining the selected at least one candidate MV as the MV of the current block. Alternatively, the inter-frame prediction unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by each of the selected at least one candidate MV. In addition, the action of searching the region of the reference picture may also be referred to as motion estimation.
[0455] In addition, in the above example, steps Se_1 to Se_3 are performed by the inter-frame prediction unit 126. However, for example, the processing of step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.
[0456] In addition, a candidate MV list may be created for each process in each inter-frame prediction mode, or a common candidate MV list may be used in multiple inter-frame prediction modes. In addition, the processes of steps Se_3 and Se_4 respectively correspond to Figure 9 the processes of steps Sa_3 and Sa_4 shown. In addition, the process of step Se_3 corresponds to Figure 30 the process of step Sd_1b.
[0457] [Flow of MV Derivation]
[0458] Figure 36 is a flowchart showing an example of MV derivation.
[0459] The inter-frame prediction unit 126 may derive the MV of the current block in a mode of encoding motion information (e.g., MV). In this case, for example, the motion information may be encoded as a prediction parameter and signaled. That is, the encoded motion information is included in the stream.
[0460] Alternatively, the inter-frame prediction unit 126 may derive an MV in a mode where motion information is not encoded. In this case, the motion information is not included in the stream.
[0461] Here, the modes of MV derivation include the ordinary inter-frame mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes in which motion information is encoded include the ordinary inter-frame mode, the ordinary merge mode, and the affine mode (specifically, the affine inter-frame mode and the affine merge mode), etc. In addition, the motion information may include not only the MV, but also the predicted MV selection information described later. In addition, the modes in which motion information is not encoded include the FRUC mode, etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0462] Figure 37 is a flowchart showing another example of MV derivation.
[0463] The inter-frame prediction unit 126 may derive the MV of the current block in a mode where the differential MV is encoded. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. The differential MV is the difference between the MV of the current block and its predicted MV. In addition, the predicted MV is a predicted motion vector.
[0464] Alternatively, the inter-frame prediction unit 126 may derive an MV in a mode where the differential MV is not encoded. In this case, the encoded differential MV is not included in the stream.
[0465] Here, as described above, the modes of MV derivation include the ordinary inter-frame mode, the ordinary merge mode, the FRUC mode, the affine mode, etc., which will be described later. Among these modes, the modes in which the differential MV is encoded include the ordinary inter-frame mode and the affine mode (specifically, the affine inter-frame mode), etc. In addition, the modes in which the differential MV is not encoded include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode), etc. The inter-frame prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0466] [Modes of MV Derivation]
[0467] Figure 38A and Figure 38B is an example of a diagram showing the classification of each mode of MV derivation. For example, as Figure 38AAs shown, according to whether motion information is encoded and whether differential MVs are encoded, the MV derivation modes are classified into three major modes. The three modes are the inter-frame mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter-frame mode is the mode for performing motion search, and it is the mode for encoding motion information and differential MVs. For example, as Figure 38B shown, the inter-frame mode includes the affine inter-frame mode and the normal inter-frame mode. The merge mode is the mode that does not perform motion search, and it is the mode for selecting an MV from the surrounding encoded blocks and using that MV to derive the MV of the current block. This merge mode is basically the mode for encoding motion information and not encoding differential MVs. For example, as Figure 38B shown, the merge mode includes the normal merge mode (sometimes also referred to as the usual merge mode or the regular merge mode), the MMVD (Merge with Motion Vector Difference) mode, the CIIP (Combined inter merge / intraprediction) mode, the triangular mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the various modes included in the merge mode, differential MVs are encoded exceptionally. In addition, the above-mentioned affine merge mode and affine inter-frame mode are the modes included in the affine mode. The affine mode is the mode that assumes an affine transformation and derives the MV of the current block by using the MVs of the multiple sub-blocks that make up the current block as the MV of the current block. The FRUC mode is the mode for deriving the MV of the current block by searching between the encoded regions, and it is the mode in which neither motion information nor differential MVs are encoded. In addition, the details of these various modes will be described later.
[0468] In addition, Figure 38A and Figure 38B shown, the classification of each mode is an example and is not limited thereto. For example, when differential MVs are encoded in the CIIP mode, this CIIP mode is classified as the inter-frame mode.
[0469] [MV Derivation > Normal Inter-Frame Mode]
[0470] The normal inter-frame mode is an inter-frame prediction mode for deriving the MV of the current block by finding a block similar to the image of the current block from the region of the reference picture represented by the candidate MVs. In addition, in this normal inter-frame mode, differential MVs are encoded.
[0471] Figure 39 is a flowchart showing an example of inter-frame prediction based on the normal inter-frame mode.
[0472] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks located around the current block in time or space (step Sg_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0473] Next, the inter-frame prediction unit 126 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_1 as prediction MV candidates respectively in a predetermined priority order (step Sg_2). In addition, this priority order is predetermined for each of the N candidate MVs.
[0474] Next, the inter-frame prediction unit 126 selects one prediction MV candidate from the N prediction MV candidates as the prediction MV for the current block (step Sg_3). At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the selected prediction MV into the stream. That is, the inter-frame prediction unit 126 outputs the prediction MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0475] Next, the inter-frame prediction unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the inter-frame prediction unit 126 also encodes the difference value between the derived MV and the prediction MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs the differential MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. In addition, the encoded reference picture is a picture composed of a plurality of blocks reconstructed after encoding.
[0476] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes of steps Sg_1 to Sg_5 are executed for each block. For example, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a slice, the inter-frame prediction using the normal inter-frame mode for that slice ends. In addition, when the processes of steps Sg_1 to Sg_5 are respectively executed for all the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends. Furthermore, it may be that when the processes of steps Sg_1 to Sg_5 are not executed for all the blocks included in a slice but for some blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, it may be that when the processes of steps Sg_1 to Sg_5 are executed for some blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture ends.
[0477] In addition, the predicted image is the inter-frame prediction signal described above. Further, information indicating an inter-frame prediction mode (in the above example, the normal inter-frame mode) used in the generation of the predicted image, which is included in the coded signal, is coded as, for example, a prediction parameter.
[0478] In addition, the candidate MV list may also be used in common with lists used in other modes. Further, processing related to the candidate MV list may be applied to processing related to lists used in other modes. Processing related to the candidate MV list is, for example, extracting or selecting a candidate MV from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.
[0479] [MV derivation > Normal merge mode]
[0480] The normal merge mode is an inter-frame prediction mode in which a candidate MV is selected from the candidate MV list as the MV of the current block to derive the MV. In addition, the normal merge mode is a narrow sense of the merge mode, and is sometimes simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broad sense.
[0481] Figure 40 is a flowchart showing an example of inter-frame prediction based on the normal merge mode.
[0482] First, the inter-frame prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of coded blocks located around the current block in time or space (step Sh_1). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0483] Next, the inter-frame prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter-frame prediction unit 126 codes MV selection information for identifying the selected candidate MV into the stream. That is, the inter-frame prediction unit 126 outputs the MV selection information as a prediction parameter to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0484] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processes of steps Sh_1 to Sh_3 are performed on each block, for example. For example, when the processes of steps Sh_1 to Sh_3 are respectively performed on all the blocks included in a slice, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_1 to Sh_3 are respectively performed on all the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends. Further, the processes of steps Sh_1 to Sh_3 may be such that when they are not performed on all the blocks included in a slice but on a part of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processes of steps Sh_1 to Sh_3 may be such that when they are performed on a part of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.
[0485] In addition, information indicating the inter-frame prediction mode (in the above example, the normal merge mode) used in the generation of the predicted image included in the stream is encoded as, for example, a prediction parameter.
[0486] Figure 41 is a diagram for explaining an example of the MV derivation process of the current picture based on the normal merge mode.
[0487] First, the inter-frame prediction unit 126 generates a candidate MV list registering candidate MVs. As candidate MVs, there are: a spatially adjacent candidate MV, which is an MV possessed by a plurality of encoded blocks located in the spatial vicinity of the current block; a temporally adjacent candidate MV, which is an MV possessed by a nearby block obtained by projecting the position of the current block in the encoded reference picture; a combined candidate MV, which is an MV generated by combining the MV values of the spatially adjacent candidate MV and the temporally adjacent candidate MV; and a zero candidate MV, which is an MV with a value of zero, etc.
[0488] Next, the inter-frame prediction unit 126 determines one candidate MV as the MV of the current block by selecting one candidate MV from among the plurality of candidate MVs registered in the candidate MV list.
[0489] Moreover, in the entropy encoding unit 110, a signal indicating which candidate MV is selected, i.e., merge_idx, is described in the stream and encoded.
[0490] In addition, Figure 41 the candidate MVs registered in the candidate MV list described in is an example, and the number may be different from that in the figure, or the structure may not include some of the candidate MVs in the figure, or may be a structure with candidate MVs added other than the types of candidate MVs in the figure.
[0491] The MV of the current block exported through the normal merge mode can also be used to perform the subsequent DMVR (dynamic motion vector refreshing) to determine the final MV. In addition, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list in the same way as the normal merge mode, but encodes the differential MV. As Figure 38B shown, such an MMVD can also be classified as a merge mode together with the normal merge mode. In addition, the differential MV in the MMVD mode may not be the same as the differential MV used in the inter-frame mode. For example, the derivation of the differential MV in the MMVD mode may also be a process with a smaller processing amount than the derivation of the differential MV in the inter-frame mode.
[0492] In addition, it is also possible to overlap the predicted image generated in the inter-frame prediction with the predicted image generated in the intra-frame prediction to perform the CIIP (Combined inter merge / intra prediction) mode for generating the predicted image of the current block.
[0493] In addition, the candidate MV list may also be referred to as the candidate list. In addition, merge_idx is MV selection information.
[0494] [MV Derivation > HMVP Mode]
[0495] Figure 42 is a diagram for explaining an example of the MV derivation process of the current picture based on the HMVP mode.
[0496] In the normal merge mode, one candidate MV is selected from the candidate MV list generated by referring to the encoded block (e.g., CU), thereby determining the MV of, for example, the CU of the current block. Here, other candidate MVs can also be registered in the candidate MV list. The mode of registering such other candidate MVs is called the HMVP mode.
[0497] In the HMVP mode, separately from the candidate MV list of the normal merge mode, a FIFO (First-In First-Out) buffer for HMVP is used to manage the candidate MVs.
[0498] In the FIFO buffer, motion information such as the MV of the block processed in the past is sequentially saved from the new FIFO buffer. In the management of this FIFO buffer, whenever one block is processed, the MV of the latest block (i.e., the immediately preceding processed CU) is saved in the FIFO buffer, and instead, the MV of the earliest CU (i.e., the CU processed first) in the FIFO buffer is deleted from the FIFO buffer. In Figure 42In the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the earliest block.
[0499] Then, for example, the inter prediction unit 126 sequentially checks each MV managed in the FIFO buffer starting from HMVP1 to see if the MV is different from all the candidate MVs already registered in the candidate MV list in the ordinary merge mode. Also, the inter prediction unit 126 may, when it determines that the MV is different from all the candidate MVs, add the MV managed in the FIFO buffer as a candidate MV to the candidate MV list in the ordinary merge mode. At this time, the candidate MV registered in the FIFO buffer may be one or more.
[0500] In this way, by using the HMVP mode, not only can the MVs of spatially or temporally adjacent blocks of the current block be added to the candidates, but also the MVs of the blocks processed in the past can be added to the candidates. As a result, by expanding the variation of the candidate MVs in the ordinary merge mode, the possibility of improving the coding efficiency becomes higher.
[0501] In addition, the above MV may also be motion information. That is, the information stored in the candidate MV list and the FIFO buffer may include not only the value of the MV, but also information such as the information indicating the reference picture, the reference direction, and the number of pictures. In addition, the above block is, for example, a CU.
[0502] In addition, Figure 42 The candidate MV list and the FIFO buffer are an example, and the candidate MV list and the FIFO buffer may also be lists or buffers of different sizes from Figure 42 or a structure in which candidate MVs are registered in a different order from Figure 42 Here, the processing described is common to both the encoding device 100 and the decoding device 200.
[0503] In addition, the HMVP mode can also be applied to modes other than the ordinary merge mode. For example, motion information such as the MVs of the blocks processed in the affine mode in the past may be sequentially stored in a new FIFO buffer and used as candidate MVs. The mode in which the HMVP mode is applied in the affine mode may be referred to as the historical affine mode.
[0504] [MV derivation>FRUC mode]
[0505] Motion information may also be derived on the side of the decoding device 200 instead of being signaled from the side of the encoding device 100. For example, motion information may also be derived by performing a motion search on the side of the decoding device 200. In such a case, a motion search is performed on the side of the decoding device 200 without using the pixel values of the current block. Such a mode of performing a motion search on the side of the decoding device 200 includes a FRUC (frame rate up-conversion) mode or a PMMVD (pattern matched motion vector derivation) mode, etc.
[0506] Figure 43 An example of FRUC processing is shown. First, with reference to the MVs of each encoded block that is spatially or temporally adjacent to the current block, a list representing these MVs as candidate MVs is generated (i.e., it is a candidate MV list and may also be common to the candidate MV list in the normal merge mode) (step Si_1). Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, the evaluation value of each candidate MV included in the candidate MV list is calculated, and one candidate is selected as the best candidate MV based on this evaluation value. And, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. In addition, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral region of the position in the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and the evaluation value in the reference picture may be performed on the peripheral region of the best candidate MV, and if there is an MV with a better evaluation value, the best candidate MV is updated to this MV and used as the final MV for the current block. The update to an MV with a better evaluation value may not be implemented.
[0507] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processes of steps Si_1 to Si_5 are performed, for example, for each block. For example, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. In addition, when the processes of steps Si_1 to Si_5 are respectively performed for all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. In addition, the processes of steps Si_1 to Si_5 may be such that when they are not performed for all the blocks included in a slice but are performed for a part of the blocks, the inter-frame prediction using the FRUC mode for that slice ends. Similarly, the processes of steps Si_1 to Si_5 may be such that when they are performed for a part of the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends.
[0508] The same processing as that in the above block unit may also be performed in the case of processing in units of sub-blocks.
[0509] The evaluation value may also be calculated by various methods. For example, the reconstructed image of the region in the reference picture corresponding to the MV is compared with the reconstructed image of a specified region (for example, as shown below, this region may be a region of another reference picture or a region of an adjacent block of the current picture). Then, the difference in pixel values of the two reconstructed images may also be calculated for use as the evaluation value of the MV. In addition, it may also be that other information is used in addition to the difference value to calculate the evaluation value.
[0510] Next, the pattern matching will be described in detail. First, one candidate MV included in the candidate MV list (also referred to as the merge list) is selected as the starting point for the search based on pattern matching. As the pattern matching, the first pattern matching or the second pattern matching may be used. The first pattern matching and the second pattern matching are sometimes referred to as bilateral matching and template matching, respectively.
[0511] [MV Derivation>FRUC>Bilateral Matching]
[0512] In the first pattern matching, pattern matching is performed between two blocks along the motion trajectory of the current block in two different reference pictures. Therefore, in the first pattern matching, as the specified region for calculating the evaluation value of the candidate MV, the region in another reference picture along the motion trajectory of the current block is used.
[0513] Figure 44This is a diagram illustrating an example of the first pattern matching (bidirectional matching) between two blocks in two reference pictures along a motion trajectory. As Figure 44 shown, in the first pattern matching, by searching for the most matching pair among pairs of two blocks in two different reference pictures (Ref0, Ref1) along the motion trajectory of the current block (Cur block), two MVs (MV0, MV1) are derived. Specifically, for the current block, the difference between the reconstructed image at the specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at the specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval is derived, and the evaluation value is calculated using the obtained difference value. A candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.
[0514] Under the assumption of a continuous motion trajectory, the MVs (MV0, MV1) indicating the two reference blocks are proportional to the temporal distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, in the first pattern matching, mirror-symmetric bidirectional MVs are derived.
[0515] [MV Derivation > FRUC > Template Matching]
[0516] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture, such as the upper and / or left adjacent block) and a block in the reference picture. Thus, in the second pattern matching, a block adjacent to the current block in the current picture is used as the specified region for calculating the evaluation value of the above candidate MV.
[0517] Figure 45 This is a diagram illustrating an example of the pattern matching (template matching) between a template in the current picture and a block in the reference picture. As Figure 45 shown, in the second pattern matching, the MV of the current block is derived by searching for the block in the reference picture (Ref0) that most matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded region of both or one of the left adjacent and upper adjacent blocks and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and the evaluation value is calculated using the obtained difference value. A candidate MV with the best evaluation value among multiple candidate MVs can be selected as the best candidate MV.
[0518] Such information indicating whether the FRUC mode is adopted (for example, called a FRUC flag) is signaled at the CU level. In addition, when the FRUC mode is adopted (for example, when the FRUC flag is true), information indicating the method of style matching that can be adopted (first style matching or second style matching) is signaled at the CU level. In addition, the signaling of this information does not need to be limited to the CU level, and can also be other levels (for example, sequence level, picture level, slice level, brick level, CTU level or sub-block level).
[0519] [MV Export > Affine Mode]
[0520] The affine mode is a mode for generating MV using affine transform, and for example, MV can be derived in sub-block units based on MVs of a plurality of adjacent blocks. This mode is sometimes referred to as affine motion compensation prediction mode.
[0521] Figure 46A FIG. 1 is a diagram for explaining an example of deriving an MV in sub-block units based on MVs of a plurality of adjacent blocks. Figure 46A In the example, the current block includes 16 sub-blocks composed of 4×4 pixels. Here, the motion vector v0 of the upper left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper right corner control point of the current block is derived based on the MV of the adjacent sub-block. Then, according to the following formula (1A), the two motion vectors v0 and v1 are projected, and the motion vectors (v x , v y ).
[0522]
Formula 1
[0523]
[0524] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a predetermined weight coefficient.
[0525] Information indicating such an affine mode (e.g., called an affine flag) may be signaled at the CU level. Furthermore, the signaling of the information indicating the affine mode need not be limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).
[0526] In addition, such an affine mode may include several modes in which the MVs of the upper left and upper right control points are derived in different ways. For example, the affine mode includes two modes: an affine inter (also called an affine normal inter) mode and an affine merge mode.
[0527] Figure 46B This is an example diagram for explaining the derivation of the MV of the sub-block unit in the affine mode using three control points. In Figure 46B , the current block includes 16 sub-blocks of 4×4 pixels. Here, the motion vector v0 of the upper-left control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper-right control point of the current block is derived based on the MV of the adjacent block, and the motion vector v2 of the lower-left control point of the current block is derived based on the MV of the adjacent block. Then, according to the following formula (1B), the three motion vectors v0, v1, and v2 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.
[0528]
Equation 2
[0529]
[0530] Here, x and y respectively represent the horizontal position and vertical position of the sub-block center, and w and h represent predetermined weight coefficients. Alternatively, w can represent the width of the current block, and h can represent the height of the current block.
[0531] The affine modes using different numbers of control points (e.g., two and three) can also be switched at the CU level and signaled. Additionally, the information indicating the number of control points of the affine mode used at the CU level can be signaled at other levels (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0532] Furthermore, in such an affine mode with three control points, several modes with different derivation methods for the MVs of the upper-left, upper-right, and lower-left control points can also be included. For example, in the affine mode with three control points, similar to the affine mode with two control points, there are two modes: the affine inter-frame mode and the affine merge mode.
[0533] Moreover, in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and can also be other sizes. For example, the size of each sub-block can also be 8×8 pixels.
[0534] [MV Derivation > Affine Mode > Control Points]
[0535] Figure 47A , Figure 47B and Figure 47C are conceptual diagrams for explaining an example of the MV derivation of the control points in the affine mode.
[0536] In the affine mode, as Figure 47AAs shown, for example, a predicted motion vector (MV) for each control point of a current block is calculated based on a plurality of MVs corresponding to blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) encoded in an affine mode and adjacent to the current block. Specifically, these blocks are checked in the order of the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) to determine the first valid block encoded in the affine mode. The MV of the control point of the current block is calculated based on the plurality of MVs corresponding to the determined block.
[0537] For example, as shown, in the case where block A adjacent to the left side of the current block is encoded in an affine mode with two control points, motion vectors v3 and v4 projected onto positions of the upper left corner and the upper right corner of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are calculated.
[0538] For example, as shown, when block A adjacent to the left side of the current block is encoded in an affine mode with three control points, motion vectors v3, v4, and v5 projected onto positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left corner control point, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point of the current block are calculated.
[0539] In addition, the method for deriving the MVs shown can be used for deriving the MVs of each control point of the current block in step Sk_1 described later, and can also be used for deriving the predicted MVs of each control point of the current block in step Sj_1 described later.
[0540] And is a conceptual diagram for explaining another example of deriving the control point MVs in the affine mode.
[0541] is a diagram for explaining the affine mode with two control points.
[0542] In this affine mode, as As shown, the motion vectors (MVs) selected from the MVs of the encoded blocks A, B, and C adjacent to the current block are used as the motion vector v0 of the upper-left control point of the current block. Similarly, the MVs selected from the MVs of the encoded blocks D and E adjacent to the current block are used as the motion vector v1 of the upper-right control point of the current block.
[0543] Fig. is for explaining the affine mode with three control points.
[0544] In this affine mode, as shown, the MVs selected from the MVs of the encoded blocks A, B, and C adjacent to the current block are used as the motion vector v0 of the upper-left control point of the current block. Similarly, the MVs selected from the MVs of the encoded blocks D and E adjacent to the current block are used as the motion vector v1 of the upper-right control point of the current block. In addition, the MVs selected from the MVs of the encoded blocks F and G adjacent to the current block are used as the motion vector v2 of the lower-left control point of the current block.
[0545] In addition, and the MV derivation methods shown can be used for the derivation of the MVs of each control point of the current block in step Sk_1 described later, and can also be used for the derivation of the predicted MVs of each control point of the current block in step Sj_1 of described later.
[0546] Here, for example, in the case of signaling with different numbers of control points (e.g., two and three) in the affine mode at the CU level, etc., the number of control points may be different depending on the encoded block and the current block.
[0547] and are conceptual diagrams for explaining an example of the MV derivation method of control points in the case where the number of control points is different between the encoded block and the current block.
[0548] For example, as shown, the current block has three control points: upper-left, upper-right, and lower-left, and the block A adjacent to the left side of the current block is encoded in an affine mode with two control points. In this case, the motion vectors v3 and v4 projected onto the positions of the upper-left and upper-right corners of the encoded block containing block A are derived. Then, based on the derived motion vectors v3 and v4, the motion vector v0 of the upper-left control point and the motion vector v1 of the upper-right control point of the current block are calculated. Furthermore, based on the derived motion vectors v0 and v1, the motion vector v2 of the lower-left control point is calculated.
[0549] For example, as As shown, the current block has two control points, namely the upper left corner and the upper right corner. The block A adjacent to the left side of the current block is encoded in an affine mode with three control points. In this case, motion vectors v3, v4, and v5 projected to the positions of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block are calculated.
[0550] In addition, and the method for deriving the MV shown can be used for deriving the MV of each control point of the current block in step Sk_1 described later, and can also be used for deriving the predicted MV of each control point of the current block in step Sj_1 described later. shown in
[0551] [MV Derivation > Affine Mode > Affine Merge Mode]
[0552] is a flowchart showing an example of the affine merge mode.
[0553] In the affine merge mode, first, the inter-frame prediction unit 126 derives the MV of each control point of the current block (step Sk_1). The control points are, as shown, the upper left and upper right points of the current block, or, as shown, the upper left, upper right, and lower left points of the current block. At this time, the inter-frame prediction unit 126 can also encode the MV selection information used to identify the two or three derived MVs into the stream.
[0554] For example, in the case of using the method for deriving the MV Figures 47A to 47C shown, as Figure 47A shown, the inter-frame prediction unit 126 checks these blocks in the order of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and determines the initial valid block encoded in the affine mode.
[0555] The inter-frame prediction unit 126 uses the first valid block encoded in the determined affine mode to derive the MV of the control point. For example, when block A is determined and block A has two control points, as Figure 47BAs shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block based on the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block containing block A. For example, by projecting the motion vectors v3 and v4 of the upper left corner and the upper right corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point and the motion vector v1 of the upper right control point of the current block.
[0556] Alternatively, in the case where block A is determined and block A has three control points, as Figure 47C shown, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block based on the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block containing block A. For example, by projecting the motion vectors v3, v4, and v5 of the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block, the inter-frame prediction unit 126 calculates the motion vector v0 of the upper left control point, the motion vector v1 of the upper right control point, and the motion vector v2 of the lower left control point of the current block.
[0557] In addition, as shown above Figure 49A shown, block A can be determined, and in the case where block A has two control points, the MV of three control points can be calculated. Also, as shown above Figure 49B shown, block A can be determined, and in the case where block A has three control points, the MV of two control points can be calculated.
[0558] Next, the inter-frame prediction unit 126 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 126 calculates the MV of each of the multiple sub-blocks as an affine MV using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B) (step Sk_2). Then, the inter-frame prediction unit 126 uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-block (step Sk_3). When the processes of steps Sk_2 and Sk_3 are respectively executed for all the sub-blocks included in the current block, the process of generating the predicted image using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and the predicted image of the current block is generated.
[0559] In addition, in step Sk_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can, for example, also be a list containing candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be Figures 47A to 47C the MV derivation method shown Figure 48A andFigure 48B The method for exporting the MV shown Figure 49A and Figure 49B the method for exporting the MV shown, and any combination of other methods for exporting MVs.
[0560] In addition, the candidate MV list may also include candidate MVs of a mode that performs prediction in units of sub-blocks other than the affine mode.
[0561] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with 2 control points and candidate MVs of an affine merge mode with 3 control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode with 2 control points may be generated separately, and a candidate MV list including candidate MVs of an affine merge mode with 3 control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes with 2 control points and the affine merge mode with 3 control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be the MV of a valid block among these blocks.
[0562] In addition, as the MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent.
[0563] [MV Export>Affine Mode>Affine Inter-frame Mode]
[0564] Figure 51 is a flowchart showing an example of the affine inter-frame mode.
[0565] In the affine inter-frame mode, first, the inter-frame prediction unit 126 exports the predicted MVs (v0, v1) or (v0, v1, v2) of each of the 2 or 3 control points of the current block (step Sj_1). As Figure 46A or Figure 46B shown, the control points are the points at the upper left corner, upper right corner, or lower left corner of the current block.
[0566] For example, in the case of using the method for exporting the MV shown in Figure 48A and Figure 48B the inter-frame prediction unit 126 exports the predicted MVs (v0, v1) or (v0, v1, v2) of the control points of the current block by selecting the MV of a certain block among the encoded blocks near each control point of the current block shown in Figure 48A or Figure 48B At this time, the inter-frame prediction unit 126 encodes the prediction MV selection information for identifying the 2 or 3 selected predicted MVs into the stream.
[0567] For example, the inter-frame prediction unit 126 can determine which block's MV among the encoded blocks adjacent to the current block is to be used as the predicted MV for the control point by using cost evaluation or the like, and can record in the bitstream a flag indicating which predicted MV is selected. That is, the inter-frame prediction unit 126 outputs, via the prediction parameter generation unit 130, prediction MV selection information such as a flag as a prediction parameter to the entropy encoding unit 110.
[0568] Next, while updating the predicted MVs selected or derived in step Sj_1 respectively (step Sj_2), the inter-frame prediction unit 126 performs motion search (steps Sj_3 and Sj_4). That is, the inter-frame prediction unit 126 sets the MV of each sub-block corresponding to the predicted MV to be updated as the affine MV, and calculates it using the above formula (1A) or formula (1B) (step Sj_3). Then, the inter-frame prediction unit 126 performs motion compensation on each sub-block using these affine MVs and the encoded reference picture (step Sj_4). Whenever the predicted MV is updated in step Sj_2, the processes of steps Sj_3 and Sj_4 are executed for all blocks within the current block. As a result, in the motion search loop, the inter-frame prediction unit 126 determines, for example, the predicted MV that can obtain the minimum cost as the MV of the control point (step Sj_5). At this time, the inter-frame prediction unit 126 also encodes the difference value between the determined MV and the predicted MV as a differential MV into the stream. That is, the inter-frame prediction unit 126 outputs, via the prediction parameter generation unit 130, the differential MV as a prediction parameter to the entropy encoding unit 110.
[0569] Finally, the inter-frame prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0570] In addition, in step Sj_1, the above-mentioned candidate MV list can also be generated. The candidate MV list can be, for example, a list containing candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be Figures 47A to 47C the MV derivation method shown, Figure 48A and Figure 48B the MV derivation method shown, Figure 49A and Figure 49B the MV derivation method shown, and any combination of other MV derivation methods.
[0571] In addition, the candidate MV list can also contain candidate MVs of prediction modes that perform prediction in units of sub-blocks other than the affine mode.
[0572] In addition, as a candidate MV list, a candidate MV list including candidate MVs of an affine inter-frame mode having two control points and candidate MVs of an affine inter-frame mode having three control points may also be generated. Alternatively, a candidate MV list including candidate MVs of an affine inter-frame mode having two control points and a candidate MV list including candidate MVs of an affine inter-frame mode having three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of a mode of either an affine inter-frame mode having two control points or an affine inter-frame mode having three control points may be generated. The candidate MV may be, for example, the MV of the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or may be the MV of the valid block among these blocks.
[0573] In addition, as prediction MV selection information, an index indicating which candidate MV in the candidate MV list may also be sent out.
[0574] [MV Derivation>Triangle Mode]
[0575] In the above example, the inter-frame prediction unit 126 generates one rectangular prediction image for the current rectangular block. However, the inter-frame prediction unit 126 may generate a plurality of prediction images having a shape different from that of the rectangle for the current rectangular block, and generate a final rectangular prediction image by combining these plurality of prediction images. The shape different from the rectangle may also be a triangle, for example.
[0576] Figure 52A is a diagram for explaining the generation of two triangular prediction images.
[0577] The inter-frame prediction unit 126 performs motion compensation for the first partition of the triangle within the current block using the first MV of the first partition, thereby generating a triangular prediction image. Similarly, the inter-frame prediction unit 126 performs motion compensation for the second partition of the triangle within the current block using the second MV of the second partition, thereby generating a triangular prediction image. Then, the inter-frame prediction unit 126 combines these prediction images, thereby generating a rectangular prediction image identical to the current block.
[0578] In addition, as the prediction image of the first partition, a first rectangular prediction image corresponding to the current block may also be generated using the first MV. In addition, as the prediction image of the second partition, a second rectangular prediction image corresponding to the current block may also be generated using the second MV. The prediction image of the current block may be generated by weighted addition of the first prediction image and the second prediction image. In addition, the region where the weighted addition is performed may also be only a partial region sandwiching the boundary between the first partition and the second partition.
[0579] Figure 52BIt is a conceptual diagram showing an example of the first part of the first partition that overlaps with the second partition, and the first sample set and the second sample set that can be weighted as part of the correction process. For example, the first part can be one-fourth of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than zero. For example, N can be the integer 2. Figure 52B A rectangular partition representing a rectangular part with a width that is one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example in the center represents a rectangular partition with a height that is one-fourth of the height of the first partition. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The example on the right represents a triangular partition with a polygonal part having a height corresponding to 2 samples. Here, the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.
[0580] The first part can be the part of the first partition that overlaps with an adjacent partition. Figure 52C It is a conceptual diagram showing the first part of the first partition that is part of the first partition and overlaps with a part of an adjacent partition. For simplicity of explanation, a rectangular partition with a part that overlaps with a spatially adjacent rectangular partition is shown. Partitions with other shapes such as triangular partitions can be used, and the overlapping part can also overlap with spatially or temporally adjacent partitions.
[0581] In addition, an example of generating prediction images for two partitions respectively using inter-frame prediction is shown, but intra-frame prediction can also be used to generate a prediction image for at least one partition.
[0582] Figure 53 It is a flowchart showing an example of a triangular pattern.
[0583] In the triangular pattern, first, the inter-frame prediction unit 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame prediction unit 126 can encode the partition information, which is the information related to the division into each partition, as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 can output the partition information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0584] Next, the inter-frame prediction unit 126 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially located around the current block (step Sx_2). That is, the inter-frame prediction unit 126 creates a candidate MV list.
[0585] Then, the inter-frame prediction unit 126 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_1 (step Sx_3). At this time, the inter-frame prediction unit 126 may also encode the MV selection information for identifying the selected candidate MV as a prediction parameter into the stream. That is, the inter-frame prediction unit 126 may output the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0586] Next, the inter-frame prediction unit 126 performs motion compensation using the selected first MV and the encoded reference picture, thereby generating a first prediction image (step Sx_4). Similarly, the inter-frame prediction unit 126 performs motion compensation using the selected second MV and the encoded reference picture, thereby generating a second prediction image (step Sx_5).
[0587] Finally, the inter-frame prediction unit 126 performs weighted addition on the first prediction image and the second prediction image, thereby generating a prediction image of the current block (step Sx_6).
[0588] In addition, in Figure 52A the example shown, the first partition and the second partition are respectively triangles, but they may also be trapezoids, or they may be respectively of different shapes. Moreover, in Figure 52A the example shown, the current block is composed of two partitions, but it may also be composed of three or more partitions.
[0589] In addition, the first partition and the second partition may also be repeated. That is, the first partition and the second partition may also include the same pixel area. In this case, the prediction image in the first partition and the prediction image in the second partition may also be used to generate the prediction image of the current block.
[0590] In addition, in this example, an example in which prediction images are generated by inter-frame prediction in both two partitions is shown, but prediction images may also be generated by intra-frame prediction for at least one partition.
[0591] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different, or they may be the same candidate MV list.
[0592] In addition, the partitioning information may include an index indicating a partitioning direction for dividing at least the current block into a plurality of partitions. The MV selection information may also include an index indicating the selected first MV and an index indicating the selected second MV. One index may also represent a plurality of pieces of information. For example, one index summarizing a part or the whole of the partitioning information and a part or the whole of the MV selection information may be encoded.
[0593] [MV Derivation>ATMVP Mode]
[0594] Figure 54 FIG. is an example of the ATMVP mode for deriving MVs in units of sub-blocks.
[0595] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, candidate MVs in units of sub-blocks are registered in a candidate MV list for a normal merge mode.
[0596] Specifically, in the ATMVP mode, first, as Figure 54 shown, in the coded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block, a temporal MV reference block corresponding to the current block is determined. Then, for each sub-block within the current block, the MV used for encoding the region corresponding to the sub-block within the temporal MV reference block is determined. The MVs thus determined are included in the candidate MV list as candidate MVs for the sub-blocks of the current block. When selecting the candidate MVs for each sub-block from the candidate MV list, motion compensation is performed on the sub-block using the candidate MV as the MV of the sub-block. Thereby, a predicted image for each sub-block is generated.
[0597] In addition, in Figure 54 the example shown, as the surrounding MV reference block, the block adjacent to the lower left of the current block is used, but other blocks may also be used. In addition, the size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may also be switched in units of slices, bricks, or pictures, etc.
[0598] [Motion Search>DMVR]
[0599] Figure 55 FIG. is a diagram showing the relationship between the merge mode and DMVR.
[0600] The inter-frame prediction unit 126 derives the MV of the current block in the merge mode (step Sl_1). Next, the inter-frame prediction unit 126 determines whether to perform an MV search, that is, a motion search (step Sl_2). Here, when it is determined not to perform a motion search (No in step Sl_2), the inter-frame prediction unit 126 determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in the merge mode.
[0601] On the other hand, when it is determined in step Sl_1 to perform a motion search (Yes in step Sl_2), the inter-frame prediction unit 126 derives the final MV for the current block by searching the peripheral area of the reference picture represented by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by the DMVR.
[0602] Figure 56 It is a conceptual diagram for explaining an example of the DMVR for determining the MV.
[0603] First, for example, in the merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are determined based on the encoded picture in the L0 list, that is, the first reference picture (L0). Similarly, according to the candidate MV (L1), reference pixels are determined based on the encoded picture in the L1 list, that is, the second reference picture (L1). A template is generated by taking the average of these reference pixels.
[0604] Next, using this template, the peripheral areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched respectively, and the MV with the minimum cost is determined as the final MV of the current block. In addition, the cost can also be calculated using, for example, the difference values between the pixel values of the template and the pixel values of the search area, as well as the candidate MV values.
[0605] Even if it is not the processing itself described here, as long as it is a processing that can search the periphery of the candidate MV to derive the final MV, any processing can be used.
[0606] Figure 57 It is a conceptual diagram for explaining another example of the DMVR for determining the MV. Figure 57 The example shown is different from Figure 56 an example of the DMVR shown, and does not generate a template but calculates the cost.
[0607] First, the inter-frame prediction unit 126 searches the periphery of the reference blocks included in the reference pictures of the L0 list and the L1 list based on the candidate MV, that is, the initial MV, obtained from the candidate MV list. For example, as Figure 57As shown, the initial MV corresponding to the reference block of the L0 list is InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is InitMV_L1. In motion search, the inter-frame prediction unit 126 first sets the search position for the reference picture in the L0 list. The differential vector representing this set search position, specifically, the differential vector from the position represented by the initial MV (i.e., InitMV_L0) to this search position is MVd_L0. Then, the inter-frame prediction unit 126 determines the search position in the reference picture of the L1 list. This search position is represented by the differential vector from the position indicated by the initial MV (i.e., InitMV_L1) to this search position. Specifically, the inter-frame prediction unit 126 determines this differential vector as MVd_L1 by mirroring MVd_L0. That is, the inter-frame prediction unit 126 sets, in the reference pictures of the L0 list and the L1 list respectively, the positions that are symmetric with respect to the position represented by the initial MV as the search positions. The inter-frame prediction unit 126 calculates, for each search position, the sum of the absolute differences (SAD) of the pixel values within the block at this search position as the cost, and finds the search position with the minimum cost.
[0608] Figure 58A FIG. is an example showing the motion search in DMVR, Figure 58B is a flowchart showing an example of this motion search.
[0609] First, in Step1, the inter-frame prediction unit 126 calculates the cost of the search position (also referred to as the start point) represented by the initial MV and the 8 search positions around it. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than the start point is the minimum. Here, when it is determined that the cost of the search positions other than the start point is the minimum, the inter-frame prediction unit 126 moves to the search position with the minimum cost and performs the processing of Step2. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 126 skips the processing of Step2 and performs the processing of Step3.
[0610] In Step2, the inter-frame prediction unit 126 uses the search position moved according to the processing result of Step1 as the new start point and performs the same search as the processing of Step1. And the inter-frame prediction unit 126 determines whether the cost of the search positions other than this start point is the minimum. Here, if the cost of the search positions other than the start point is the minimum, the inter-frame prediction unit 126 performs the processing of Step4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 126 performs the processing of Step3.
[0611] In Step4, the inter-frame prediction unit 126 processes the search position of this start point as the final search position, and determines the difference between the position represented by the initial MV and this final search position as the differential vector.
[0612] In Step 3, the inter-frame prediction unit 126 determines the pixel position with fractional precision having the minimum cost based on the costs at four points above, below, left, and right of the start point in Step 1 or Step 2, and sets this pixel position as the final search position. This pixel position with fractional precision is determined by weighted addition of the vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) at the four points above, below, left, and right, using the costs of the search positions of the respective four points as weights. Then, the inter-frame prediction unit 126 determines the difference between the position indicated by the initial MV and this final search position as the difference vector.
[0613] [Motion Compensation > BIO / OBMC / LIC]
[0614] In motion compensation, there is a mode of generating a predicted image and correcting this predicted image. This mode is, for example, BIO, OBMC, and LIC described later.
[0615] Figure 59 It is a flowchart showing an example of the generation of a predicted image.
[0616] The inter-frame prediction unit 126 generates a predicted image (step Sm_1), and corrects this predicted image by any of the above modes (step Sm_2).
[0617] Figure 60 It is a flowchart showing another example of the generation of a predicted image.
[0618] The inter-frame prediction unit 126 derives the MV of the current block (step Sn_1). Next, the inter-frame prediction unit 126 generates a predicted image using this MV (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform a correction process (Yes in step Sn_3), the inter-frame prediction unit 126 generates a final predicted image by correcting this predicted image (step Sn_4). Also, in LIC described later, the luminance and color difference can also be corrected in step Sn_4. On the other hand, when it is determined not to perform a correction process (No in step Sn_3), the inter-frame prediction unit 126 outputs this predicted image as the final predicted image without correcting this predicted image (step Sn_5).
[0619] [Motion Compensation > OBMC]
[0620] Not only the motion information of the current block obtained through motion search can be used, but also the motion information of adjacent blocks can be used to generate an inter-frame prediction image. Specifically, an inter-frame prediction image can also be generated in units of sub-blocks within the current block by weighted addition of a prediction image based on the motion information obtained through motion search (within the reference picture) and a prediction image based on the motion information of adjacent blocks (within the current picture). Such inter-frame prediction (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or the OBMC mode.
[0621] In the OBMC mode, information indicating the size of the sub-blocks used for OBMC (e.g., referred to as the OBMC block size) can also be signaled at the sequence level. Also, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Additionally, the level at which these pieces of information are signaled need not be limited to the sequence level and the CU level, and can also be other levels (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0622] A more specific description of the OBMC mode is given. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for explaining the outline of the prediction image correction process based on OBMC.
[0623] First, as Figure 62 shown, using the MV assigned to the current block, a prediction image (Pred) based on normal motion compensation is obtained. In Figure 62 , the arrow "MV" points to the reference picture and indicates which block in the reference picture the current block in the current picture refers to for obtaining the prediction image.
[0624] Next, the MV (MV_L) that has been derived for the already encoded left adjacent block is applied (reused) to the current block to obtain a prediction image (Pred_L). The MV (MV_L) is represented by the arrow "MV_L" pointing from the current block to the reference picture. Then, by overlapping the two prediction images Pred and Pred_L, the first correction of the prediction image is performed. This has the effect of blending the boundaries between adjacent blocks.
[0625] Similarly, the Motion Vector (MV_U) that has been derived for the already encoded upper adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_U). The MV (MV_U) is represented by the arrow "MV_U" pointing from the current block to the reference picture. Then, the second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted images that have been corrected for the first time (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block where the boundaries with adjacent blocks are blended (smoothed).
[0626] In addition, in the above example, a two-path correction method using the left adjacent and upper adjacent blocks is used, but this correction method can also be a three-path or more path correction method that also uses the right adjacent and / or lower adjacent blocks.
[0627] Furthermore, the overlapping region can also be not the entire pixel region of the block, but only a partial region near the block boundary.
[0628] In addition, the prediction image correction process of OBMC has been described here, where the prediction image correction process of OBMC is used to obtain one predicted image Pred by overlapping one reference picture with the additional predicted images Pred_L and Pred_U. However, in the case of correcting the predicted image based on multiple reference images, the same process can be applied to each of the multiple reference pictures. In this case, through the OBMC image correction based on multiple reference pictures, after obtaining the corrected predicted images from each reference picture, the final predicted image is obtained by further overlapping the obtained multiple corrected predicted images.
[0629] In addition, in OBMC, the unit of the current block can be the PU unit or a sub-block unit obtained by further dividing the PU.
[0630] As a method for determining whether to apply OBMC, for example, there is a method of using a signal indicating whether to apply OBMC, i.e., obmc_flag. As a specific example, the encoding device 100 can also determine whether the current block belongs to a region with complex motion. When the encoding device 100 belongs to a region with complex motion, it sets the obmc_flag value to 1 and applies OBMC for encoding. When it does not belong to a region with complex motion, it sets the obmc_flag value to 0 and encodes the block without applying OBMC. On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, it switches whether to apply OBMC for decoding according to this value.
[0631] [Motion Compensation > BIO]
[0632] Next, a method for deriving an MV will be described. First, a mode for deriving an MV based on a model assuming uniform linear motion will be described. This mode is sometimes referred to as the BIO (bi - directional optical flow) mode. Additionally, the bi - directional optical flow can also be expressed as BDOF instead of BIO.
[0633] Figure 63 is a diagram for explaining a model assuming uniform linear motion. In Figure 63 , (v x , v y ) represents a velocity vector, τ0 and τ1 respectively represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.
[0634] At this time, under the assumption of uniform linear motion of the velocity vector (v x , v y ), (MVx0, MVy0) and (MVx1, MVy1) are respectively expressed as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), and the following optical flow equation (2) holds.
[0635]
Equation 3
[0636]
[0637] Here, I(k) represents the luminance value of the reference image k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the temporal differential of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Alternatively, based on the combination of this optical flow equation and Hermite interpolation, the motion vectors in block units obtained from a candidate MV list, etc. can be corrected in pixel units.
[0638] Additionally, an MV can also be derived on the side of the decoding device 200 by a method different from the derivation of the motion vector based on the model assuming uniform linear motion. For example, a motion vector can also be derived in sub - block units based on the MVs of multiple adjacent blocks.
[0639] Figure 64 is a flowchart showing an example of inter - frame prediction according to BIO. Additionally, Figure 65This is a diagram showing an example of the functional structure of the inter-frame prediction unit 126 that performs inter-frame prediction according to BIO.
[0640] As Figure 65 shown, the inter-frame prediction unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a predicted image correction unit 126f. Additionally, the memory 126a may be a frame memory 122.
[0641] The inter-frame prediction unit 126 uses two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) containing the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 126 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_1). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0642] Next, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to derive the interpolation image I of the current block 0 . Additionally, the interpolation image derivation unit 126b refers to the memory 126a and uses the motion vector M1 and the reference picture L1 to derive the interpolation image I of the current block 1 (step Sy_2). Here, the interpolation image I 0 is the image contained in the reference picture Ref0 derived for the current block, and the interpolation image I 1 is the image contained in the reference picture Ref1 derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Or, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each be an image larger than the current block. Additionally, the interpolation image I 0 and I 1 can include the predicted image derived by applying the motion vectors (M0, M1) and the reference pictures (L0, L1), as well as a motion compensation filter.
[0643] Additionally, the gradient image derivation unit 126c derives the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 , Iy 0 , Iy 1)(Step Sy_3). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivation unit 126c can also derive the gradient image by applying a gradient filter to the interpolated image, for example. The gradient image only needs to represent the spatial change amount of the pixel values along the horizontal direction or the vertical direction.
[0644] Next, the optical flow derivation unit 126d derives the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block, using the interpolated image (I 0 , I 1 ) and the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (Step Sy_4). The optical flow is a coefficient for correcting the spatial movement amount of pixels, and can also be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can be a 4x4 pixel sub-CU. In addition, the derivation of the optical flow can also be performed in other units such as pixel units instead of sub-block units.
[0645] Next, the inter-frame prediction unit 126 corrects the prediction image of the current block using the optical flow (vx, vy). For example, the correction value derivation unit 126e derives the correction value of the values of the pixels included in the current block using the optical flow (vx, vy) (Step Sy_5). Moreover, the prediction image correction unit 126f can also correct the prediction image of the current block using the correction value (Step Sy_6). In addition, the correction value can be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.
[0646] In addition, the processing flow of the BIO is not limited to Figure 64 the processing disclosed. It can either perform only a part of the processing disclosed in Figure 64 , or add or replace different processing, or execute it in a different processing order.
[0647] [Motion Compensation > LIC]
[0648] Next, an example of a mode for generating a prediction image (prediction) using LIC (local illumination compensation) will be described.
[0649] Figure 66A is a diagram for explaining an example of a prediction image generation method using the luminance correction process based on LIC. In addition, Figure 66BThis is a flowchart showing an example of a predictive image generation method using this LIC.
[0650] First, the inter-frame prediction unit 126 derives an MV from the encoded reference picture and obtains the reference image corresponding to the current block (step Sz_1).
[0651] Next, the inter-frame prediction unit 126 extracts information indicating how the luminance values change in the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions within the reference picture specified by the derived MV. Then, the inter-frame prediction unit 126 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_3).
[0652] The inter-frame prediction unit 126 generates a predicted image for the current block by performing a luminance correction process on the reference image within the reference picture specified by the MV by applying this luminance correction parameter (step Sz_4). That is, the prediction image, which is the reference image within the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the color difference can be corrected. That is, information indicating how the color difference changes can also be used to calculate a correction parameter for the color difference and perform a color difference correction process.
[0653] In addition, Figure 66A The shape of the peripheral reference region in is an example, and shapes other than this can also be used.
[0654] Furthermore, the process of generating a predicted image based on one reference picture has been described here, but the same applies when generating a predicted image based on multiple reference pictures. A predicted image can also be generated after performing a luminance correction process on the reference images obtained from each reference picture in the same manner as described above.
[0655] As a method for determining whether to apply the LIC, for example, there is a method of using a lic_flag as a signal indicating whether to apply the LIC. As a specific example, in the encoding device 100, it is determined whether the current block belongs to a region where a luminance change has occurred. If it belongs to a region where a luminance change has occurred, the value 1 is set as the lic_flag and encoding is performed by applying the LIC. If it does not belong to a region where a luminance change has occurred, the value 0 is set as the lic_flag and encoding is performed without applying the LIC. On the other hand, in the decoding device 200, it is also possible to decode the lic_flag described in the stream and switch whether to apply the LIC according to its value for decoding.
[0656] As another method for determining whether to apply LIC, for example, there is also a method of determining based on whether LIC is applied to neighboring blocks. As a specific example, when the current block is processed in the merge mode, the inter-frame prediction unit 126 determines whether the encoded neighboring blocks selected at the time of deriving the MV in the merge mode are encoded using LIC. Based on the result, the inter-frame prediction unit 126 switches whether to apply LIC for encoding. In addition, in the case of this example, the same processing also applies to the decoding device 200 side.
[0657] Use Figure 66A And Figure 66B LIC (luminance correction processing) has been described. Hereinafter, its detailed content will be described.
[0658] First, the inter-frame prediction unit 126 derives an MV for obtaining a reference image corresponding to the current block from a reference picture that is an encoded picture.
[0659] Next, for the current block, the inter-frame prediction unit 126 uses the luminance pixel values of the left and upper neighboring encoded peripheral reference regions and the luminance pixel values at the same positions in the reference picture specified by the MV to extract information indicating how the luminance values change between the reference picture and the current picture, and calculates a luminance correction parameter. For example, let the luminance pixel value of a certain pixel in the peripheral reference region within the current picture be p0, and let the luminance pixel value of the pixel at the same position in the peripheral reference region within the reference picture be p1. The inter-frame prediction unit 126 calculates coefficients A and B for optimizing A×p1 + B = p0 as the luminance correction parameter for multiple pixels in the peripheral reference region.
[0660] Next, the inter-frame prediction unit 126 generates a prediction image for the current block by performing a luminance correction process on the reference image in the reference picture specified by the MV using the luminance correction parameter. For example, let the luminance pixel value in the reference image be p2, and let the luminance pixel value of the prediction image after the luminance correction process be p3. The inter-frame prediction unit 126 generates a prediction image after the luminance correction process by calculating A×p2 + B = p3 for each pixel in the reference image.
[0661] In addition, a part of the Figure 66A shown peripheral reference region can also be used. For example, a region including a specified number of pixels respectively removed at intervals from the upper neighboring pixel and the left neighboring pixel can be used as the peripheral reference region. In addition, the peripheral reference region is not limited to the region adjacent to the current block, and can also be a region not adjacent to the current block. In addition, in Figure 66AIn the example shown, the surrounding reference area in the reference picture is the area specified by the MV of the current picture from the surrounding reference area in the current picture, but it can also be the area specified by other MVs. For example, the other MV can also be the MV of the surrounding reference area in the current picture.
[0662] In addition, the operation of the encoding device 100 is described here, but the operation of the decoding device 200 is the same.
[0663] Furthermore, LIC can be applied not only to luminance but also to color difference. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of them.
[0664] In addition, LIC can also be applied in units of sub-blocks. For example, correction parameters can be derived using the surrounding reference area of the current sub-block and the surrounding reference area of the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0665] [Prediction control unit]
[0666] The prediction control unit 128 selects one of the intra-predicted image (pixels or signals output from the intra-prediction unit 124) and the inter-predicted image (pixels or signals output from the inter-prediction unit 126), and outputs the selected predicted image to the subtraction unit 104 and the addition unit 116.
[0667] [Prediction parameter generation unit]
[0668] The prediction parameter generation unit 130 can output information related to the selection of the predicted image in intra-prediction, inter-prediction, and the prediction control unit 128 as prediction parameters to the entropy encoding unit 110. The entropy encoding unit 110 can generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters can also be used in the decoding device 200. The decoding device 200 can also receive and decode the stream, and perform the same processing as the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128. The prediction parameters can include the selected prediction signal (e.g., MV, prediction type, or prediction mode used by the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value based on or representing the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.
[0669] [Decoding device]
[0670] Next, a decoding device 200 that can decode the stream output from the above encoding device 100 will be described. Figure 67It is a block diagram showing an example of the functional structure of a decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes an encoded image, i.e., a stream, in block units.
[0671] As Figure 67 shown, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a segmentation determination unit 224. In addition, the intra prediction unit 216 and the inter prediction unit 218 are each configured as part of a prediction processing unit.
[0672] [Installation example of decoding device]
[0673] Figure 68 It is a block diagram showing an installation example of the decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, Figure 67 as shown, the multiple components of the decoding device 200 are implemented by Figure 68 the shown processor b1 and memory b2.
[0674] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 can also be a processor such as a CPU. In addition, the processor b1 can also be an aggregate of multiple electronic circuits. In addition, for example, the processor b1 can also function as Figure 67 multiple components of the decoding device 200 shown, etc., excluding the components for storing information.
[0675] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode a stream. The memory b2 can be either an electronic circuit or connected to the processor b1. In addition, the memory b2 can also be included in the processor b1. In addition, the memory b2 can also be an aggregate of multiple electronic circuits. In addition, the memory b2 can be a magnetic disk or an optical disc, etc., or can be represented as a storage or a recording medium, etc. In addition, the memory b2 can be either a non-volatile memory or a volatile memory.
[0676] For example, the memory b2 can store an image or a stream. In addition, a program for the processor b1 to decode a stream can also be stored in the memory b2.
[0677] In addition, for example, the memory b2 can also function as Figure 67The functions of the components for storing information among the multiple components of the decoding device 200 shown, etc. Specifically, the memory b2 can perform the functions of Figure 67 the block memory 210 and the frame memory 214 shown. More specifically, reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.) can be stored in the memory b2.
[0678] In addition, in the decoding device 200, all of the multiple components shown, etc. may not be installed, and all of the above-mentioned multiple processes may not be performed. Figure 67 A part of the multiple components shown, etc. may be included in other devices, or a part of the above-mentioned multiple processes may be performed by other devices. Figure 67
[0679]
[0680]
[0681] [Overall Process of Decoding Processing] Figure 69
[0682] It is a flowchart showing an example of the overall decoding process performed by the decoding device 200.
[0683] First, the segmentation determination unit 224 of the decoding device 200 determines the segmentation style of each of the multiple fixed-size blocks (128×128 pixels) included in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_1). This segmentation style is the segmentation style selected by the encoding device 100. Then, the decoding device 200 performs the processes of steps Sp_2 to Sp_6 on each of the multiple blocks constituting this segmentation style.
[0683] The entropy decoding unit 202 decodes the encoded quantization coefficients and prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).
[0684] Next, the inverse quantization unit 204 and the inverse transformation unit 206 restore the prediction residual of the current block by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (step Sp_3).
[0685] Next, a prediction processing unit including an intra prediction unit 216, an inter prediction unit 218, and a prediction control unit 220 generates a prediction image of the current block (step Sp_4).
[0686] Next, the addition unit 208 reconstructs the current block into a reconstructed image (also referred to as a decoded image block) by adding the prediction residual to the prediction image (step Sp_5).
[0687] Moreover, when generating the reconstructed image, the loop filter unit 212 filters the reconstructed image (step Sp_6).
[0688] Then, the decoding device 200 determines whether the decoding of the entire picture has been completed (step Sp_7). If it is determined that the decoding is not completed (No in step Sp_7), the processing from step Sp_1 is repeated.
[0689] In addition, the processing of steps Sp_1 to Sp_7 can be sequentially performed by the decoding device 200. Multiple processes of a part of these processes can be performed in parallel, or the order can be changed.
[0690] [Segmentation decision unit]
[0691] Figure 70 FIG. is a diagram showing the relationship between the segmentation decision unit 224 and other components. As an example, the segmentation decision unit 224 can also perform the following processing.
[0692] The segmentation decision unit 224 collects block information from, for example, the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoding unit 202. Moreover, the segmentation decision unit 224 can determine the segmentation pattern of a fixed-size block based on the block information and the parameters. And the segmentation decision unit 224 can also output information indicating the determined segmentation pattern to the inverse transformation unit 206, the intra prediction unit 216, and the inter prediction unit 218. The inverse transformation unit 206 can perform inverse transformation on the transformation coefficients based on the segmentation pattern indicated by the information from the segmentation decision unit 224. The intra prediction unit 216 and the inter prediction unit 218 can generate a prediction image based on the segmentation pattern indicated by the information from the segmentation decision unit 224.
[0693] [Entropy decoding unit]
[0694] Figure 71 FIG. is a block diagram showing an example of the functional structure of the entropy decoding unit 202.
[0695] The entropy decoding unit 202 generates quantized coefficients, prediction parameters, and parameters related to the segmentation pattern, etc. by performing entropy decoding on the stream. For example, CABAC is used in this entropy decoding. Specifically, the entropy decoding unit 202 includes, for example, a binary arithmetic decoding unit 202a, a context control unit 202b, and a de-binarization unit 202c. The binary arithmetic decoding unit 202a performs arithmetic decoding on the stream as a binary signal using the context value derived by the context control unit 202b. Similar to the context control unit 110b of the encoding device 100, the context control unit 202b derives a context value corresponding to the feature of the syntax element or the surrounding situation, that is, the occurrence probability of the binary signal. The de-binarization unit 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoding unit 202a into a multi-valued signal representing the above-mentioned quantized coefficients, etc. This de-binarization is performed in the same manner as the above-mentioned binarization.
[0696] The entropy decoding unit 202 outputs the quantized coefficients to the inverse quantization unit 204 in units of blocks. The entropy decoding unit 202 may also output the prediction parameters included in the stream (see Figure 1 ) to the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220. The intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220 can perform the same prediction processing as the processing performed by the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128 on the encoding device 100 side.
[0697] [Entropy Decoding Unit]
[0698] Figure 72 is a diagram showing the process of CABAC in the entropy decoding unit 202.
[0699] First, in the CABAC in the entropy decoding unit 202, initialization is performed. In this initialization, initialization in the binary arithmetic decoding unit 202c and setting of the initial context value are performed. Then, the binary arithmetic decoding unit 202c and the de-binarization unit 202c perform arithmetic decoding and de-binarization on the encoded data of the CTU, for example. At this time, the context control unit 202b updates the context value each time arithmetic decoding is performed. Then, the context control unit 202b saves the context value as post-processing. The saved context value is used, for example, as the initial value of the context value for the next CTU.
[0700] [Inverse Quantization Unit]
[0701] The inverse quantization unit 204 performs inverse quantization on the quantization coefficients of the current block that is input from the entropy decoding unit 202. Specifically, for the quantization coefficients of the current block, the inverse quantization unit 204 performs inverse quantization on each quantization coefficient based on the quantization parameter corresponding to the quantization coefficient. Further, the inverse quantization unit 204 outputs the inverse quantized quantization coefficients (i.e., transform coefficients) of the current block to the inverse transform unit 206.
[0702] Figure 73 FIG. is a block diagram showing an example of the functional configuration of the inverse quantization unit 204.
[0703] The inverse quantization unit 204 includes, for example, a quantization parameter generation unit 204a, a predicted quantization parameter generation unit 204b, a quantization parameter storage unit 204d, and an inverse quantization processing unit 204e.
[0704] Figure 74 FIG. is a flowchart showing an example of the inverse quantization performed by the inverse quantization unit 204.
[0705] As an example, the inverse quantization unit 204 can perform inverse quantization processing for each CU based on the Figure 74 flow shown. Specifically, the quantization parameter generation unit 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generation unit 204a obtains the differential quantization parameter of the current block from the entropy decoding unit 202 (step Sv_12).
[0706] Next, the predicted quantization parameter generation unit 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage unit 204d (step Sv_13). The predicted quantization parameter generation unit 204b generates a predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).
[0707] Then, the quantization parameter generation unit 204a adds the differential quantization parameter of the current block obtained from the entropy decoding unit 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generation unit 204b (step Sv_15). By this addition, the quantization parameter of the current block is generated. Further, the quantization parameter generation unit 204a stores the quantization parameter of the current block in the quantization parameter storage unit 204d (step Sv_16).
[0708] Next, the inverse quantization processing unit 204e inverse quantizes the quantization coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).
[0709] In addition, the differential quantization parameter can also be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can also be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the quantization parameter can be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0710] In addition, the inverse quantization unit 204 may include multiple inverse quantizers, and may also inverse-quantize the quantization coefficients using an inverse quantization method selected from multiple inverse quantization methods.
[0711] [Inverse transform unit]
[0712] The inverse transform unit 206 restores the prediction residual by inverse-transforming the transform coefficients that are the input from the inverse quantization unit 204.
[0713] For example, when the information read from the stream indicates the application of EMT or AMT (e.g., the AMT flag is true), the inverse transform unit 206 inverse-transforms the transform coefficients of the current block based on the information indicating the transform type read.
[0714] In addition, for example, when the information read from the stream indicates the application of NSST, the inverse transform unit 206 applies an inverse re-transformation to the transform coefficients.
[0715] Figure 75 It is a flowchart showing an example of the process performed by the inverse transform unit 206.
[0716] For example, the inverse transform unit 206 determines whether there is information indicating that orthogonal transformation is not performed in the stream (step St_11). Here, when it is determined that there is no such information (No in step St_11), the inverse transform unit 206 obtains the information indicating the transform type that has been decoded by the entropy decoding unit 202 (step St_12). Then, the inverse transform unit 206 determines the transform type used in the orthogonal transformation of the encoding device 100 based on this information (step St_13). And the inverse transform unit 206 performs inverse orthogonal transformation using the determined transform type (step St_14).
[0717] Figure 76 It is a flowchart showing another example of the process performed by the inverse transform unit 206.
[0718] For example, the inverse transform unit 206 determines whether the transform size is equal to or less than a specified value (step Su_11). Here, when it is determined that the transform size is equal to or less than the specified value (Yes in step Su_11), the inverse transform unit 206 obtains from the entropy decoding unit 202 information indicating which one of the one or more transform types included in the first transform type group is used by the encoding device 100 (step Su_12). Further, such information is decoded by the entropy decoding unit 202 and output to the inverse transform unit 206.
[0719] Based on this information, the inverse transform unit 206 determines the transform type to be used in the orthogonal transform in the encoding device 100 (step Su_13). Then, the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). On the other hand, when it is determined in step Su_11 that the transform size is not equal to or less than the specified value (No in step Su_11), the inverse transform unit 206 performs an inverse orthogonal transform on the transform coefficients of the current block using the second transform type group (step Su_15).
[0720] In addition, as an example, the inverse orthogonal transform performed by the inverse transform unit 206 can be implemented for each TU according to Figure 75 or Figure 76 the shown process. Alternatively, instead of decoding the information indicating the transform type used in the orthogonal transform, an inverse orthogonal transform can be performed using a pre-specified transform type. Specifically, the transform type is, for example, DST7 or DCT8, and in the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.
[0721] [Addition unit]
[0722] The addition unit 208 reconstructs the current block by adding the prediction residual, which is the input from the inverse transform unit 206, to the predicted image, which is the input from the prediction control unit 220. That is, a reconstructed image of the current block is generated. Then, the addition unit 208 outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.
[0723] [Block memory]
[0724] The block memory 210 is a storage unit for storing blocks within the current picture that are referred to in intra prediction. Specifically, the block memory 210 stores the reconstructed image output from the addition unit 208.
[0725] [Loop filter unit]
[0726] The loop filter unit 212 applies loop filtering to the reconstructed image generated by the addition unit 208, and outputs the filtered reconstructed image to the frame memory 214, the display device, and the like.
[0727] When the information indicating the ON / OFF of ALF read from the stream indicates that ALF is ON, one filter is selected from among a plurality of filters based on the direction and activity of the locality-based gradient, and the selected filter is applied to the reconstructed image.
[0728] Figure 77 FIG. is a block diagram showing an example of the functional configuration of the loop filter unit 212. In addition, the loop filter unit 212 has the same configuration as the loop filter unit 120 of the encoding apparatus 100.
[0729] The loop filter unit 212, for example, as Figure 77 shown, includes a deblocking filter processing unit 212a, an SAO processing unit 212b, and an ALF processing unit 212c. The deblocking filter processing unit 212a performs the above-described deblocking filter processing on the reconstructed image. The SAO processing unit 212b performs the above-described SAO processing on the reconstructed image after the deblocking filter processing. In addition, the ALF processing unit 212c applies the above-described ALF processing to the reconstructed image after the SAO processing. In addition, the loop filter unit 212 may not include Figure 77 all of the processing units disclosed, or may include only a part of the processing units. In addition, the loop filter unit 212 may also be configured to perform the above-described respective processes in an order different from the processing order disclosed in Figure 77 .
[0730] [Frame memory]
[0731] The frame memory 214 is a storage unit for storing reference pictures used in inter-frame prediction, and is sometimes referred to as a frame buffer. Specifically, the frame memory 214 stores the reconstructed image that has been filtered by the loop filter unit 212.
[0732] [Prediction unit (intra-frame prediction unit / inter-frame prediction unit / prediction control unit)]
[0733] Figure 78 FIG. is a flowchart showing an example of the processing performed by the prediction unit of the decoding apparatus 200. In addition, as an example, the prediction unit is composed of all or part of the constituent elements of the intra-frame prediction unit 216, the inter-frame prediction unit 218, and the prediction control unit 220. The prediction processing unit includes, for example, the intra-frame prediction unit 216 and the inter-frame prediction unit 218.
[0734] The prediction unit generates a prediction image for the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. Additionally, among the prediction signals, there are, for example, an intra prediction signal and an inter prediction signal. Specifically, the prediction unit generates a prediction image for the current block using a reconstructed image that has already been obtained by generating a prediction image for another block, restoring a prediction residual, and adding the prediction images. The prediction unit of the decoding device 200 generates the same prediction image as the prediction image generated by the prediction unit of the encoding device 100. That is, the methods for generating the prediction images used in these prediction units are mutually common or corresponding.
[0735] The reconstructed image can be, for example, an image of a reference picture, or an image of a decoded block within the current picture that contains the current block (i.e., the above-mentioned other block). The decoded block within the current picture is, for example, an adjacent block of the current block.
[0736] Figure 79 It is a flowchart showing another example of the processing performed by the prediction unit of the decoding device 200.
[0737] The prediction unit determines the method or mode for generating the prediction image (step Sr_1). For example, this method or mode can be determined based on, for example, prediction parameters, etc.
[0738] When it is determined that the first method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the first method (step Sr_2a). In addition, when it is determined that the second method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the second method (step Sr_2b). In addition, when it is determined that the third method is the mode for generating the prediction image, the prediction unit generates the prediction image according to the third method (step Sr_2c).
[0739] The first method, the second method, and the third method are different methods for generating the prediction image, and can be, for example, an inter prediction method, an intra prediction method, and other prediction methods. In such prediction methods, the above-mentioned reconstructed image can also be used.
[0740] Figure 80A and Figure 80B It is a flowchart showing another example of the processing performed by the prediction unit in the decoding device 200.
[0741] As an example, the prediction unit can also perform prediction processing according to the Figure 80A and Figure 80B shown process. Additionally, Figure 80A and Figure 80BThe intra-block copy shown belongs to one of the inter prediction modes, and is a mode in which the blocks included in the current picture are referred to as reference pictures or reference blocks. That is, in intra-block copy, pictures different from the current picture are not referred to. Additionally, Figure 80A The PCM mode shown belongs to one of the intra prediction modes and is a mode in which no transformation and quantization are performed.
[0742] [Intra Prediction Unit]
[0743] The intra prediction unit 216 performs intra prediction with reference to the blocks within the current picture stored in the block memory 210 based on the intra prediction mode decoded from the bitstream, thereby generating a predicted picture (i.e., intra prediction picture) of the current block. Specifically, the intra prediction unit 216 generates an intra prediction picture by performing intra prediction with reference to the pixel values (e.g., luminance values, chrominance difference values) of the blocks adjacent to the current block, and outputs the intra prediction picture to the prediction control unit 220.
[0744] Additionally, in the case where the intra prediction mode that refers to the luminance block is selected for the intra prediction of the chrominance difference block, the intra prediction unit 216 may also predict the chrominance difference component of the current block based on the luminance component of the current block.
[0745] Furthermore, in the case where the information decoded from the bitstream indicates the application of PDPC, the intra prediction unit 216 corrects the pixel values after intra prediction based on the gradients of the reference pixels in the horizontal / vertical directions.
[0746] Figure 81 is a diagram showing an example of the processing performed by the intra prediction unit 216 of the decoding apparatus 200.
[0747] The intra prediction unit 216 first determines whether an MPM flag indicating 1 exists in the bitstream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 exists (Yes in step Sw_11), the intra prediction unit 216 obtains from the entropy decoding unit 202 the information indicating the intra prediction mode selected in the encoding apparatus 100 among the MPMs (step Sw_12). Additionally, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Next, the intra prediction unit 216 determines the MPM (step Sw_13). The MPM is composed of, for example, six intra prediction modes. Then, the intra prediction unit 216 determines the intra prediction mode indicated by the information obtained in step Sw_12 from among the multiple intra prediction modes included in this MPM (step Sw_14).
[0748] On the other hand, when it is determined in step Sw_11 that there is no MPM flag representing 1 in the stream (No in step Sw_11), the intra prediction unit 216 acquires information representing the intra prediction mode selected in the encoding device 100 (step Sw_15). That is, the intra prediction unit 216 acquires from the entropy decoding unit 202 information representing the intra prediction mode selected in the encoding device 100 among one or more intra prediction modes not included in the MPM. In addition, this information is decoded by the entropy decoding unit 202 and output to the intra prediction unit 216. Then, the intra prediction unit 216 determines the intra prediction mode represented by the information acquired in step Sw_15 from among the one or more intra prediction modes not included in the MPM (step Sw_17).
[0749] The intra prediction unit 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0750] [Inter prediction unit]
[0751] The inter prediction unit 218 predicts the current block with reference to the reference pictures stored in the frame memory 214. The prediction is performed in units of the current block or sub-blocks within the current block. In addition, a sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can also be switched in units of slices, bricks, or pictures, etc.
[0752] For example, the inter prediction unit 218 performs motion compensation using the motion information (e.g., MV) read from the stream (e.g., the prediction parameters output from the entropy decoding unit 202), thereby generating an inter prediction image of the current block or sub-block, and outputting the inter prediction image to the prediction control unit 220.
[0753] When the information read from the stream indicates the application of the OBMC mode, the inter prediction unit 218 generates an inter prediction image using not only the motion information of the current block obtained through motion search but also the motion information of adjacent blocks.
[0754] In addition, when the information read from the stream indicates the application of the FRUC mode, the inter prediction unit 218 performs motion search according to the pattern matching method (bidirectional matching or template matching) read from the stream, thereby deriving the motion information. And the inter prediction unit 218 uses the derived motion information for motion compensation (prediction).
[0755] In addition, when the inter-frame prediction unit 218 applies the BIO mode, it derives the MV based on a model assuming uniform linear motion. In addition, when the information read from the stream indicates that the affine mode is applied, the inter-frame prediction unit 218 derives the MV in sub-block units based on the MVs of multiple adjacent blocks.
[0756] [Flow of MV derivation]
[0757] Figure 82 It is a flowchart showing an example of MV derivation in the decoding device 200.
[0758] The inter-frame prediction unit 218 determines, for example, whether to decode motion information (e.g., MV). For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the motion information, the inter-frame prediction unit 218 derives the MV of the current block in the mode of decoding this motion information. On the other hand, when it is determined not to decode the motion information, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the motion information.
[0759] Here, the MV derivation modes include the ordinary inter-frame mode, ordinary merge mode, FRUC mode, affine mode, etc., which will be described later. Among these modes, the modes of decoding motion information include the ordinary inter-frame mode, ordinary merge mode, and affine mode (specifically, affine inter-frame mode and affine merge mode), etc. In addition, the motion information can include not only the MV, but also the predicted MV selection information described later. In addition, the mode of not decoding the motion information includes the FRUC mode, etc. The inter-frame prediction unit 218 selects the mode for deriving the MV of the current block from these multiple modes, and uses the selected mode to derive the MV of the current block.
[0760] Figure 83 It is a flowchart showing another example of MV derivation in the decoding device 200.
[0761] The inter-frame prediction unit 218 determines, for example, whether to decode the differential MV. For example, the inter-frame prediction unit 218 can determine based on the prediction mode included in the stream, or can also determine based on other information included in the stream. Here, when it is determined to decode the differential MV, the inter-frame prediction unit 218 can derive the MV of the current block in the mode of decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.
[0762] On the other hand, when it is determined not to decode the differential MV, the inter-frame prediction unit 218 derives the MV in the mode of not decoding the differential MV. In this case, the encoded differential MV is not included in the stream.
[0763] Here, as described above, the export modes of the MV include the following ordinary inter-frame mode, ordinary merge mode, FRUC mode, and affine mode. Among these modes, the modes for encoding the differential MV include the ordinary inter-frame mode and the affine mode (specifically, the affine inter-frame mode). In addition, the modes that do not encode the differential MV include the FRUC mode, the ordinary merge mode, and the affine mode (specifically, the affine merge mode). The inter-frame prediction unit 218 selects a mode for exporting the MV of the current block from these multiple modes and exports the MV of the current block using the selected mode.
[0764] [MV Export > Ordinary Inter-frame Mode]
[0765] For example, when the information read from the stream indicates that the ordinary inter-frame mode is applied, the inter-frame prediction unit 218 exports the MV in the ordinary merge mode based on the information read from the stream and performs motion compensation (prediction) using the MV.
[0766] Figure 84 It is a flowchart showing an example of inter-frame prediction by the ordinary inter-frame mode in the decoding device 200.
[0767] The inter-frame prediction unit 218 of the decoding device 200 performs motion compensation for each block. At this time, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sg_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0768] Next, the inter-frame prediction unit 218 extracts N (N is an integer of 2 or more) candidate MVs from the plurality of candidate MVs obtained in step Sg_11 as prediction motion vector candidates (also referred to as prediction MV candidates) in a predetermined priority order (step Sg_12). In addition, this priority order may also be determined in advance for each of the N prediction MV candidates.
[0769] Next, the inter-frame prediction unit 218 decodes the prediction MV selection information from the input stream and uses the decoded prediction MV selection information to select one prediction MV candidate from the N prediction MV candidates as the prediction MV of the current block (step Sg_13).
[0770] Next, the inter-frame prediction unit 218 decodes the differential MV from the input stream and derives the MV of the current block by adding the difference value of the decoded differential MV to the selected prediction MV (step Sg_14).
[0771] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes of steps Sg_11 to Sg_15 are executed for each block. For example, when the processes of steps Sg_11 to Sg_15 are executed for all the blocks included in a slice respectively, the inter-frame prediction using the normal inter-frame mode for that slice ends. Also, when the processes of steps Sg_11 to Sg_15 are executed for all the blocks included in a picture respectively, the inter-frame prediction using the normal inter-frame mode for that picture ends. In addition, the processes of steps Sg_11 to Sg_15 may be such that when they are not executed for all the blocks included in a slice but for some of the blocks, the inter-frame prediction using the normal inter-frame mode for that slice ends. Similarly, when the processes of steps Sg_11 to Sg_15 are executed for some of the blocks included in a picture, the inter-frame prediction using the normal inter-frame mode for that picture may end.
[0772] [MV Derivation>Normal Merge Mode]
[0773] For example, in the case where the information read from the stream indicates the application of the normal merge mode, the inter-frame prediction unit 218 derives an MV in the normal merge mode and performs motion compensation (prediction) using the MV.
[0774] Figure 85 It is a flowchart showing an example of inter-frame prediction based on the normal merge mode in the decoding device 200.
[0775] The inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially around the current block (step Sh_11). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0776] Next, the inter-frame prediction unit 218 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_11 (step Sh_12). Specifically, the inter-frame prediction unit 218 obtains, for example, MV selection information included as a prediction parameter in the stream and selects the candidate MV identified by the MV selection information as the MV of the current block.
[0777] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processes of steps Sh_11 to Sh_13 are performed on each block, for example. For example, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a slice respectively, the inter-frame prediction using the normal merge mode for that slice ends. Also, when the processes of steps Sh_11 to Sh_13 are performed on all the blocks included in a picture respectively, the inter-frame prediction using the normal merge mode for that picture ends. Further, the processes of steps Sh_11 to Sh_13 may be such that when they are not performed on all the blocks included in a slice but on a part of the blocks, the inter-frame prediction using the normal merge mode for that slice ends. Similarly, the processes of steps Sh_11 to Sh_13 may be such that when they are performed on a part of the blocks included in a picture, the inter-frame prediction using the normal merge mode for that picture ends.
[0778] [MV Derivation > FRUC Mode]
[0779] For example, in the case where the information read from the stream indicates the application of the FRUC mode, the inter-frame prediction unit 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding device 100 side but is derived on the decoding device 200 side. For example, the decoding device 200 may also derive the motion information by performing a motion search. In this case, the decoding device 200 does not use the pixel values of the current block for the motion search.
[0780] Figure 86 It is a flowchart showing an example of the inter-frame prediction based on the FRUC mode in the decoding device 200.
[0781] First, the inter-frame prediction unit 218 refers to the MVs of each decoded block that is spatially or temporally adjacent to the current block, and generates a list representing these MVs as candidate MVs (i.e., it is a candidate MV list and can also be common with the candidate MV list in the normal merge mode) (step Si_11). Next, the inter-frame prediction unit 218 selects the best candidate MV from among the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter-frame prediction unit 218 calculates the evaluation value of each candidate MV included in the candidate MV list, and selects one candidate MV as the best candidate MV based on this evaluation value. Then, the inter-frame prediction unit 218 derives the MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Additionally, for example, the MV for the current block may also be derived by performing pattern matching in the peripheral area of the position in the reference picture corresponding to the selected best candidate MV. That is, for the area around the best candidate MV, a search using pattern matching and evaluation values in the reference picture is performed. Furthermore, in the case where there is an MV with a good evaluation value, the best candidate MV may also be updated to this MV and used as the final MV for the current block. It is also possible not to perform the update to an MV with a better evaluation value.
[0782] Finally, the inter-frame prediction unit 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processing of steps Si_11 to Si_15 is performed for each block, for example. For example, when the processing of steps Si_11 to Si_15 is performed for all the blocks included in a slice, the inter-frame prediction using the FRUC mode for that slice ends. Additionally, when the processing of steps Si_11 to Si_15 is performed for all the blocks included in a picture, the inter-frame prediction using the FRUC mode for that picture ends. It is also possible to perform the processing in the same manner as the above block unit in sub-block units.
[0783] [MV Derivation > Affine Merge Mode]
[0784] For example, in the case where the information read from the stream indicates the application of the affine merge mode, the inter-frame prediction unit 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using this MV.
[0785] Figure 87 It is a flowchart showing an example of inter-frame prediction based on the affine merge mode in the decoding apparatus 200.
[0786] In the affine merge mode, the inter-frame prediction unit 218 first derives the MV of each control point of the current block (step Sk_11). As Figure 46AAs shown, the control points are the points at the upper left corner and the upper right corner of the current block, or as Figure 46B shown, are the points at the upper left corner, the upper right corner, and the lower left corner of the current block.
[0787] For example, in the case of using the Figures 47A to 47C MV derivation method shown, as Figure 47A shown, the inter-frame prediction unit 218 checks these blocks in the order of the decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), and determines the first valid block decoded in the affine mode.
[0788] The inter-frame prediction unit 218 uses the first valid block decoded in the determined affine mode to derive the MV of the control point. For example, in the case where block A is determined and block A has two control points, as Figure 47B shown, the inter-frame prediction unit 218 projects the motion vectors v3 and v4 of the upper left corner and the upper right corner of the decoded block including block A onto the current block to calculate the motion vector v0 of the upper left corner control point of the current block and the motion vector v1 of the upper right corner control point. Thus, the MV of each control point is derived.
[0789] In addition, as Figure 49A shown, in the case where block A is determined and block A has two control points, it is also possible to calculate the MV of three control points, or as Figure 49B shown, determine block A, and in the case where block A has three control points, calculate the MV of two control points.
[0790] In addition, when the stream includes MV selection information as a prediction parameter, the inter-frame prediction unit 218 can also use this MV selection information to derive the MV of each control point of the current block.
[0791] Next, the inter-frame prediction unit 218 performs motion compensation on each of the multiple sub-blocks included in the current block. That is, the inter-frame prediction unit 218 calculates the MV of the sub-block as an affine MV (step Sk_12) for each of the multiple sub-blocks using two motion vectors v0 and v1 and the above formula (1A), or using three motion vectors v0, v1, and v2 and the above formula (1B). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine MVs and the decoded reference picture (step Sk_13). When the processes of steps Sk_12 and Sk_13 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a prediction image of the current block is generated.
[0792] In addition, in step Sk_11, the above-mentioned candidate MV list may also be generated. The candidate MV list may also be, for example, a list including candidate MVs derived using multiple MV derivation methods for each control point. The multiple MV derivation methods may be Figures 47A to 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in, and any combination of other MV derivation methods.
[0793] In addition, the candidate MV list may also include candidate MVs of a prediction mode that is performed in units of sub-blocks other than the affine mode.
[0794] In addition, as the candidate MV list, for example, a candidate MV list including candidate MVs of an affine merge mode with two control points and candidate MVs of an affine merge mode with three control points may be generated. Alternatively, a candidate MV list including candidate MVs of an affine merge mode with two control points and a candidate MV list including candidate MVs of an affine merge mode with three control points may be generated separately. Alternatively, a candidate MV list including candidate MVs of one of the affine merge modes with two control points and the affine merge mode with three control points may be generated.
[0795] [MV Derivation > Affine Inter-Frame Mode]
[0796] For example, when the information read from the stream indicates the application of the affine inter-frame mode, the inter-frame prediction unit 218 derives an MV in the affine inter-frame mode and performs motion compensation (prediction) using the MV.
[0797] Figure 88 is a flowchart showing an example of inter-frame prediction based on the affine inter-frame mode in the decoding device 200.
[0798] In the affine inter-frame mode, first, the inter-frame prediction unit 218 derives prediction MVs (v0, v1) or (v0, v1, v2) for each of the two or three control points of the current block (step Sj_11). The control points are, for example, as Figure 46A or Figure 46B shown, the upper left, upper right, or lower left point of the current block.
[0799] The inter-frame prediction unit 218 obtains prediction MV selection information included as prediction parameters in the stream, and uses the MV identified by the prediction MV selection information to derive the prediction MVs for the respective control points of the current block. For example, when using the Figure 48A and Figure 48B shown MV derivation methods, the inter-frame prediction unit 218 selects Figure 48Aor Figure 48B The motion vectors of the blocks identified by the prediction MV selection information in the decoded blocks near the respective control points of the current block shown in Figure 48B are used to derive the predicted motion vectors (v0, v1) or (v0, v1, v2) of the control points of the current block.
[0800] Next, the inter-frame prediction unit 218 obtains, for example, each differential motion vector included as a prediction parameter in the stream, and adds the predicted motion vectors of the respective control points of the current block and the differential motion vectors corresponding to the predicted motion vectors (step Sj_12). Thereby, the motion vectors of the respective control points of the current block are derived.
[0801] Next, the inter-frame prediction unit 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. That is, the inter-frame prediction unit 218 calculates the motion vector of each of the plurality of sub-blocks as an affine motion vector using two motion vectors v0 and v1 and the above-described formula (1A), or using three motion vectors v0, v1, and v2 and the above-described formula (1B) (step Sj_13). Then, the inter-frame prediction unit 218 performs motion compensation on the sub-block using these affine motion vectors and the decoded reference picture (step Sj_14). When the processes of steps Sj_13 and Sj_14 are respectively executed for all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. That is, motion compensation is performed on the current block, and a predicted image of the current block is generated.
[0802] In addition, in step Sj_11, the above-described candidate motion vector list may be generated in the same manner as in step Sk_11.
[0803] [MV Derivation > Triangle Mode]
[0804] For example, in the case where the information read from the stream indicates the application of the triangle mode, the inter-frame prediction unit 218 derives a motion vector in the triangle mode and performs motion compensation (prediction) using the motion vector.
[0805] Figure 89 is a flowchart showing an example of inter-frame prediction based on the triangle mode in the decoding apparatus 200.
[0806] In the triangle mode, first, the inter-frame prediction unit 218 divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter-frame prediction unit 218 may obtain, as a prediction parameter, partition information which is information related to the division into each partition, from the stream. Moreover, the inter-frame prediction unit 218 may divide the current block into the first partition and the second partition according to the partition information.
[0807] Next, the inter-frame prediction unit 218 first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). That is, the inter-frame prediction unit 218 creates a candidate MV list.
[0808] Then, the inter-frame prediction unit 218 respectively selects the candidate MV of the first partition and the candidate MV of the second partition as the first MV and the second MV from the plurality of candidate MVs obtained in step Sx_11 (step Sx_13). At this time, the inter-frame prediction unit 218 may also obtain MV selection information for identifying the selected candidate MVs from the stream as a prediction parameter. Then, the inter-frame prediction unit 218 may select the first MV and the second MV according to the MV selection information.
[0809] Next, the inter-frame prediction unit 218 performs motion compensation by using the selected first MV and the decoded reference picture to generate a first prediction image (step Sx_14). Similarly, the inter-frame prediction unit 218 performs motion compensation by using the selected second MV and the decoded reference picture to generate a second prediction image (step Sx_15).
[0810] Finally, the inter-frame prediction unit 218 generates a prediction image of the current block by performing weighted addition on the first prediction image and the second prediction image (step Sx_16).
[0811] [Motion Search > DMVR]
[0812] For example, in the case where the information read from the stream indicates the application of DMVR, the inter-frame prediction unit 218 performs motion search by DMVR.
[0813] Figure 90 It is a flowchart showing an example of motion search based on DMVR in the decoding device 200.
[0814] The inter-frame prediction unit 218 first derives the MV of the current block in the merge mode (step Sl_11). Next, the inter-frame prediction unit 218 searches the peripheral area of the reference picture represented by the MV derived in step Sl_11 to derive the final MV for the current block (step Sl_12). That is, the MV of the current block is determined by DMVR.
[0815] Figure 91 It is a flowchart showing a detailed example of motion search based on DMVR in the decoding device 200.
[0816] First, the inter-frame prediction unit 218 is in Figure 58AIn Step 1 shown above, the search positions (also referred to as start points) of the initial MV representation and the costs of 8 search positions around it are calculated. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, when it is determined that the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 moves to the search position with the minimum cost and performs Figure 58A the process of Step 2 shown above. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 skips Figure 58A the process of Step 2 shown above and performs the process of Step 3.
[0817] In Figure 58A Step 2 shown above, the inter-frame prediction unit 218 uses the search position moved according to the processing result of Step 1 as a new start point and performs the same search as the processing of Step 1. Further, the inter-frame prediction unit 218 determines whether the costs of search positions other than the start point are the minimum. Here, if the cost of a search position other than the start point is the minimum, the inter-frame prediction unit 218 performs the process of Step 4. On the other hand, if the cost of the start point is the minimum, the inter-frame prediction unit 218 performs the process of Step 3.
[0818] In Step 4, the inter-frame prediction unit 218 treats the search position of the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.
[0819] In Figure 58A Step 3 shown above, the inter-frame prediction unit 218 determines the pixel position with the minimum cost in decimal precision based on the costs of 4 points above, below, left, and right of the start point in Step 1 or Step 2, and uses this pixel position as the final search position. This pixel position in decimal precision is determined by weighted addition of the vectors of the 4 points above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)) with the costs of the search positions of the respective 4 points as weights. Then, the inter-frame prediction unit 218 determines the difference between the position indicated by the initial MV and the final search position as the difference vector.
[0820] [Motion Compensation > BIO / OBMC / LIC]
[0821] For example, in the case where the information read from the stream represents the application of the correction of the predicted image, when generating the predicted image, the inter-frame prediction unit 218 corrects the predicted image according to the correction mode. This mode is, for example, the above-mentioned BIO, OBMC, and LIC, etc.
[0822] Figure 92 is a flowchart showing an example of the generation of the predicted image in the decoding device 200.
[0823] The inter-frame prediction unit 218 generates a prediction image (step Sm_11), and corrects the prediction image by any of the above modes (step Sm_12).
[0824] Figure 93 It is a flowchart showing another example of the generation of a prediction image in the decoding device 200.
[0825] The inter-frame prediction unit 218 derives the MV of the current block (step Sn_11). Next, the inter-frame prediction unit 218 generates a prediction image using this MV (step Sn_12), and determines whether to perform a correction process (step Sn_13). For example, the inter-frame prediction unit 218 obtains the prediction parameters included in the stream, and determines whether to perform a correction process based on the prediction parameters. The prediction parameter is, for example, a flag indicating whether to apply each of the above modes. Here, when it is determined to perform a correction process (Yes in step Sn_13), the inter-frame prediction unit 218 generates a final prediction image by correcting the prediction image (step Sn_14). In addition, in the LIC, the luminance and chrominance differences of the prediction image can be corrected in step Sn_14. On the other hand, when it is determined not to perform a correction process (No in step Sn_13), the inter-frame prediction unit 218 outputs the prediction image without correcting it as the final prediction image (step Sn_15).
[0826] [Motion Compensation>OBMC]
[0827] For example, when the information read from the stream indicates the application of OBMC, when generating a prediction image, the inter-frame prediction unit 218 corrects the prediction image according to OBMC.
[0828] Figure 94 It is a flowchart showing an example of the correction of a prediction image based on OBMC in the decoding device 200. In addition, Figure 94 The flowchart of Figure 62 shows the process of correcting the prediction image using the current picture and the reference picture shown in
[0829] First, as Figure 62 shown, the inter-frame prediction unit 218 obtains a prediction image (Pred) based on normal motion compensation using the MV assigned to the current block.
[0830] Next, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_L) that has been derived for the decoded left adjacent block to the current block, and obtains a prediction image (Pred_L). Then, the inter-frame prediction unit 218 performs the first correction of the prediction image by overlapping the two prediction images Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0831] Similarly, the inter-frame prediction unit 218 applies (re-uses) the MV (MV_U) that has been derived for the decoded upper adjacent block to the current block, and obtains a predicted image (Pred_U). Then, the inter-frame prediction unit 218 performs a second correction of the predicted image by overlapping the predicted image Pred_U with the predicted images that have undergone the first correction (e.g., Pred and Pred_L). This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is the final predicted image of the current block that is blended (smoothed) with the boundaries of adjacent blocks.
[0832] [Motion Compensation > BIO]
[0833] For example, in the case where the information read from the stream indicates the application of BIO, when generating a predicted image, the inter-frame prediction unit 218 corrects the predicted image according to BIO.
[0834] Figure 95 It is a flowchart showing an example of the correction of the predicted image based on BIO in the decoding device 200.
[0835] As Figure 63 shown, the inter-frame prediction unit 218 uses two reference pictures (Ref0, Ref1) different from the picture (CurPic) containing the current block to derive two motion vectors (M0, M1). Then, the inter-frame prediction unit 218 uses these two motion vectors (M0, M1) to derive the predicted image of the current block (step Sy_11). Additionally, the motion vector M0 is the motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0836] Next, the inter-frame prediction unit 218 uses the motion vector M0 and the reference picture L0 to derive the interpolation image I 0 of the current block. Additionally, the inter-frame prediction unit 218 uses the motion vector M1 and the reference picture L1 to derive the interpolation image I 1 of the current block (step Sy_12). Here, the interpolation image I 0 is the image contained in the reference picture Ref0 that is derived for the current block, and the interpolation image I 1 is the image contained in the reference picture Ref1 that is derived for the current block. The interpolation image I 0 and the interpolation image I 1 can each be the same size as the current block. Or, in order to appropriately derive the gradient image described later, the interpolation image I 0 and the interpolation image I 1 can each also be an image larger than the current block. Furthermore, the interpolation image I 0 and I 1It may include a predicted image derived by applying motion vectors (M0, M1) and reference pictures (L0, L1), and a motion compensation filter.
[0837] In addition, the inter-frame prediction unit 218 derives the gradient image (Ix 0 and Ix 1 of the current block from the interpolation image I 0 and Ix 1 , Iy 0 , Iy 1 )(step Sy_13). In addition, the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Ix 0 , Ix 1 ). The inter-frame prediction unit 218 may also derive the gradient image by applying a gradient filter to the interpolation image, for example. The gradient image may be an image representing the spatial variation amount of pixel values along the horizontal direction or the vertical direction.
[0838] Next, the inter-frame prediction unit 218 uses the interpolation image (I 0 , I 1 ) and the gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) to derive the optical flow (vx, vy) as the above-mentioned velocity vector in units of a plurality of sub-blocks constituting the current block (step Sy_14). As an example, the sub-block may be a 4x4 pixel sub-CU.
[0839] Next, the inter-frame prediction unit 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame prediction unit 218 derives a correction value of the pixel values included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter-frame prediction unit 218 may correct the predicted image of the current block using the correction value (step Sy_16). In addition, the correction value may be derived in units of each pixel, or in units of a plurality of pixels or sub-blocks.
[0840] In addition, the processing flow of BIO is not limited to Figure 95 the disclosed processing. It may perform only a part of the disclosed processing, or may add or replace different processing, or may execute in a different processing order. Figure 95 the disclosed processing.
[0841] [Motion Compensation > LIC]
[0842] For example, in the case where the information read from the stream represents the application of LIC, when generating a predicted image, the inter-frame prediction unit 218 corrects the predicted image according to LIC.
[0843] Figure 96 It is a flowchart showing an example of the correction of the predicted image based on LIC in the decoding device 200.
[0844] First, the inter-frame prediction unit 218 uses the MV to obtain a reference image corresponding to the current block from the decoded reference pictures (step Sz_11).
[0845] Next, the inter-frame prediction unit 218 extracts information indicating how the luminance values change in the reference picture and the current picture for the current block (step Sz_12). As Figure 66A shown, this extraction is based on the luminance pixel values of the decoded left adjacent reference region (peripheral reference region) and the decoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the same position in the reference picture specified by the derived MV. Then, the inter-frame prediction unit 218 calculates a luminance correction parameter using the information indicating how the luminance values change (step Sz_13).
[0846] The inter-frame prediction unit 218 generates a predicted image for the current block by performing a luminance correction process of applying its luminance correction parameter to the reference image in the reference picture specified by the MV (step Sz_14). That is, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance or the color difference can be corrected.
[0847] [Prediction control unit]
[0848] The prediction control unit 220 selects one of the intra-frame predicted image and the inter-frame predicted image, and outputs the selected predicted image to the adder 208. Generally, the structures, functions, and processes of the prediction control unit 220, the intra-frame prediction unit 216, and the inter-frame prediction unit 218 on the decoding device 200 side can correspond to the structures, functions, and processes of the prediction control unit 128, the intra-frame prediction unit 124, and the inter-frame prediction unit 126 on the encoding device 100 side.
[0849] [Regarding the picture segmentation method]
[0850] Figure 97 It is a conceptual diagram showing the relationship between sub-pictures, slices, and tiles. Here, an example of a picture segmentation method using sub-pictures, slices, and tiles is shown. Specifically, in Figure 97 the example, the picture is divided into a plurality of tiles, a plurality of slices, and a plurality of sub-pictures.
[0851] The tiles are obtained by dividing the picture horizontally and vertically. In Figure 97 's example, the picture is divided into two parts horizontally and two parts vertically, and thus is divided into a total of 4 tiles.
[0852] Moreover, in Figure 97 's example, each tile is divided into more than 1 slice. Specifically, the upper left tile is divided into 1 slice, the upper right tile is divided into 4 slices, the lower left tile is divided into 2 slices, and the lower right tile is divided into 3 slices. In this case, each slice is a rectangle, so it is called a rectangular slice. A rectangular slice can be composed of multiple tiles that make up a rectangular area, or can be composed of more than 1 CTU row within 1 tile.
[0853] Moreover, in Figure 97 's example, the picture is divided into multiple sub-pictures, and the multiple sub-pictures are respectively units that aggregate more than 1 slice or more than 1 tile.
[0854] The sub-pictures are used for the following purposes, for example. That is, the sub-pictures can be extracted from the stream in units of sub-pictures, and only the extracted sub-pictures are transmitted as a stream and can also be decoded. In addition, the extracted sub-pictures can also be combined with the sub-pictures of other streams and reconstructed as 1 stream.
[0855] Figure 98 is a conceptual diagram showing the relationship between tiles and slices. In Figure 98 's example, the picture is divided into multiple tiles and multiple slices. Specifically, the picture is divided into four parts horizontally and four parts vertically, and thus is divided into a total of 16 tiles.
[0856] Furthermore, in Figure 98 's example, multiple tiles form 1 slice. Slice (1) is composed of 5 tiles, slice (2) is composed of 5 tiles, and slice (3) is composed of 6 tiles. At this time, each slice aggregates multiple tiles that are consecutive in raster scan order and determines them as 1 slice, so it is called a raster-scan slice.
[0857] [The First Form of Syntax Related to Sub-Pictures]
[0858] Figure 99 is the syntax structure diagram of the first form related to sub-pictures. Figure 99 Shows an example of the syntax structure for reading the sub-picture index. In Figure 99 , subpic_idx[i] is the sub-picture index for the tile with tile index i.
[0859] Specifically, when NumTilesInPic representing the number of tiles is 1 or more and subpics_present_flag representing whether there is information related to sub - pictures is true, the sub - picture indices are signaled for each tile.
[0860] In this form, a sub - picture is a rectangular area of one or more tiles within a picture. Additionally, it is determined that a tile is not included in multiple sub - pictures and a slice is not included in multiple sub - pictures. Figure 100A and Figure 100B represents a specific example of this form.
[0861] Figure 100A is a conceptual diagram showing the tile indices of each area in a picture. Specifically, Figure 100A shows a picture and multiple areas within the picture. Figure 100A The multiple areas shown are respectively tiles. In Figure 100A numerically represented tile indices are assigned to each tile.
[0862] Figure 100B is a conceptual diagram showing the sub - picture indices of each area in a picture. In Figure 100B a picture and multiple areas within the picture are also shown. Similar to Figure 100A the above, Figure 100B the multiple areas shown are respectively tiles. In Figure 100B numerically represented sub - picture indices are assigned to each tile.
[0863] For example, in the tile with tile index 0 and the tile with tile index 6, 0 is assigned as the sub - picture index. That is, the tile with tile index 0 and the tile with tile index 6 form one sub - picture with sub - picture index 0.
[0864] This syntax can also be signaled when NumTilesInPic is greater than 1.
[0865] [The second form of the syntax related to sub - pictures]
[0866] Figure 101 is the syntax structure diagram of the second form related to sub - pictures. In this example, assigning the same sub - picture index to multiple unconnected tiles is suppressed. Additionally, assigning the same sub - picture index to multiple tiles forming a non - rectangular area is suppressed.
[0867] Here, NumSubpicsInPic represents the number of sub - pictures in a picture. The number of sub - pictures in a picture can be obtained from the SPS or PPS.
[0868] In addition, top_left_tile_idx[i] is the tile index of the top-left tile in the sub-picture with sub-picture index i. In addition, bottom_right_tile_idx[i] is the tile index of the bottom-right tile in the sub-picture with sub-picture index i. In Figure 100A and Figure 100B 's example, these values are determined as follows.
[0869] top_left_tile_idx[0] = 0
[0870] bottom_right_tile_idx[0] = 6
[0871] top_left_tile_idx[1] = 1
[0872] bottom_right_tile_idx[1] = 7
[0873] top_left_tile_idx[2] = 2
[0874] bottom_right_tile_idx[2] = 8
[0875] …
[0876] That is, the tile index of the top-left tile in the sub-picture with sub-picture index 0 is 0. In addition, the tile index of the bottom-right tile in the sub-picture with sub-picture index 0 is 6.
[0877] In addition, the tile index of the top-left tile in the sub-picture with sub-picture index 1 is 1. In addition, the tile index of the bottom-right tile in the sub-picture with sub-picture index 1 is 7.
[0878] In addition, the tile index of the top-left tile in the sub-picture with sub-picture index 2 is 2. In addition, the tile index of the bottom-right tile in the sub-picture with sub-picture index 2 is 8.
[0879] This syntax can also be signaled when NumTilesInPic is greater than 1.
[0880] [The third form of the syntax related to sub-pictures]
[0881] Figure 102 is the syntax structure diagram of the third form related to sub-pictures. In this example, the sub-picture index is defined by bricks. In addition, in this example, subpic_idx[i] is the sub-picture index for the brick with brick index i.
[0882] Specifically, when NumBricksInPic representing the number of bricks is 1 or more and subpics_present_flag representing whether there is information related to sub - pictures is true, the sub - picture index is signaled for each brick.
[0883] In this form, a sub - picture is a rectangular area of one or more bricks within the picture. Additionally, it is determined that a brick is not included in multiple sub - pictures. Figures 103A to 103C Shows a specific example of this form.
[0884] Figure 103A It is a conceptual diagram showing the brick index of each area in the picture. Specifically, Figure 103A The picture and multiple areas within the picture are shown. Additionally, in Figure 103A the numerical value represents the brick index. One or more areas with the same numerical value form one brick. For example, the upper - left area and the area below it form one brick with a brick index of 0.
[0885] Figure 103B It is a conceptual diagram showing the tile index of each area in the picture. In Figure 103B the picture and multiple areas within the picture are also shown. Additionally, in Figure 103B the numerical value represents the tile index. One or more areas with the same numerical value form one tile. For example, the upper - left area and the area below it form one tile with a tile index of 0.
[0886] Figure 103C It is a conceptual diagram showing the sub - picture index of each area in the picture. In Figure 103C the picture and multiple areas within the picture are also shown. Additionally, in Figure 103C the numerical value represents the sub - picture index. One or more areas with the same numerical value form one sub - picture. For example, the upper - left area and the area below it form one sub - picture with a sub - picture index of 0.
[0887] In this example, multiple sub - pictures within the picture coincide with multiple bricks within the picture. Additionally, the same value as the brick index of each brick in the picture is assigned as the sub - picture index. Additionally, the tile with a tile index of 5 straddles the sub - picture with a sub - picture index of 5 and the sub - picture with a sub - picture index of 6.
[0888] This syntax can also be signaled when NumBricksInPic is greater than 1.
[0889] [The fourth form of the syntax related to sub - pictures]
[0890] Figure 104It is the syntactic structure diagram related to sub - pictures in the 4th form. In this example, it is inhibited to assign the same sub - picture index to multiple unconnected bricks. Additionally, it is inhibited to assign the same sub - picture index to multiple bricks that form a non - rectangular area.
[0891] Here, NumSubpicsInPic represents the number of sub - pictures in a picture. The number of sub - pictures in a picture can be obtained from the SPS or PPS.
[0892] Furthermore, top_left_brick_idx[i] is the brick index of the top - left brick in the sub - picture with sub - picture index i. Additionally, bottom_right_brick_idx[i] is the brick index of the bottom - right brick in the sub - picture with sub - picture index i. In Figures 103A to 103C the example of
[0893] top_left_brick_idx[0] = 0
[0894] bottom_right_brick_idx[0] = 0
[0895] top_left_brick_idx[1] = 1
[0896] bottom_right_brick_idx[1] = 1
[0897] top_left_brick_idx[2] = 2
[0898] bottom_right_brick_idx[2] = 2
[0899] …
[0900] That is, the brick index of the top - left brick in the sub - picture with sub - picture index 0 is 0. Additionally, the brick index of the bottom - right brick in the sub - picture with sub - picture index 0 is 0.
[0901] Additionally, the brick index of the top - left brick in the sub - picture with sub - picture index 1 is 1. Additionally, the brick index of the bottom - right brick in the sub - picture with sub - picture index 1 is 1.
[0902] Additionally, the brick index of the top - left brick in the sub - picture with sub - picture index 2 is 2. Additionally, the brick index of the bottom - right brick in the sub - picture with sub - picture index 2 is 2.
[0903] This syntax can also be signaled when NumBricksInPic is greater than 1.
[0904] [The 5th form of the syntax related to sub - pictures]
[0905] Figure 105 It is the syntactic structure diagram related to sub - pictures in the 5th form. In this example, the sub - picture index is defined by slices. Additionally, in this example, subpic_idx[i] is the sub - picture index for the slice with slice index i.
[0906] Specifically, when NumSlicesInPic representing the number of slices is 1 or more and subpics_present_flag representing whether there is information related to sub - pictures is true, the sub - picture index is signaled for each slice.
[0907] The sub - pictures in this form are rectangular regions of one or more slices within the picture. Additionally, a slice is determined not to be included in multiple sub - pictures. That is, one sub - picture index is assigned to one slice. When the sub - picture is valid, rectangular slices can also be used as slices. Figures 106A to 106C Shows a specific example of this form.
[0908] Figure 106A It is a conceptual diagram showing the slice indices of each region in the picture. Specifically, Figure 106A shows the picture and multiple regions within the picture. Additionally, in Figure 106A the numerical values represent slice indices. One or more regions with the same numerical value form one slice. For example, the upper - left region and the region below it form one slice with slice index 0.
[0909] Figure 106B It is a conceptual diagram showing the tile indices of each region in the picture. Figure 106B also shows the picture and multiple regions within the picture. Additionally, in Figure 106B the numerical values represent tile indices. One or more regions with the same numerical value form one tile. For example, the upper - left region and the region below it form one tile with tile index 0.
[0910] Figure 106C It is a conceptual diagram showing the sub - picture indices of each region in the picture. In Figure 106C the picture and multiple regions within the picture are also shown. Additionally, in Figure 106C the numerical values represent sub - picture indices. One or more regions with the same numerical value form one sub - picture. For example, the upper - left region and the region below it form one sub - picture with sub - picture index 0.
[0911] In this example, multiple sub - pictures within the picture coincide with multiple slices within the picture. Additionally, in each slice within the picture, the value same as the slice index of that slice is assigned as the sub - picture index. Additionally, the tile with tile index 5 straddles the sub - picture with sub - picture index 5 and the sub - picture with sub - picture index 6.
[0912] This syntax can also be signaled when NumSlicesInPic is greater than 1.
[0913] [The 6th form of the syntax related to sub - pictures]
[0914] Figure 107 It is the syntax structure diagram of the 6th form related to sub - pictures. In this example, assigning the same sub - picture index to unconnected multiple slices is suppressed. In addition, assigning the same sub - picture index to multiple slices that form a non - rectangular region is suppressed.
[0915] Here, NumSubpicsInPic represents the number of sub - pictures in the picture. The number of sub - pictures in the picture can be obtained from the SPS or PPS.
[0916] In addition, top_left_slice_idx[i] is the slice index of the top - left slice in the sub - picture with sub - picture index i. Also, bottom_right_slice_idx[i] is the slice index of the bottom - right slice in the sub - picture with sub - picture index i. In Figures 106A to 106C this example, these values are determined as follows.
[0917] top_left_slice_idx[0] = 0
[0918] bottom_right_slice_idx[0] = 0
[0919] top_left_slice_idx[1] = 1
[0920] bottom_right_slice_idx[1] = 1
[0921] top_left_slice_idx[2] = 2
[0922] bottom_right_slice_idx[2] = 2
[0923] …
[0924] That is, the slice index of the top - left slice in the sub - picture with sub - picture index 0 is 0. Also, the slice index of the bottom - right slice in the sub - picture with sub - picture index 0 is 0.
[0925] In addition, the slice index of the top - left slice in the sub - picture with sub - picture index 1 is 1. Also, the slice index of the bottom - right slice in the sub - picture with sub - picture index 1 is 1.
[0926] In addition, the slice index of the upper-left slice in the sub-picture with sub-picture index 2 is 2. In addition, the slice index of the lower-right slice in the sub-picture with sub-picture index 2 is 2.
[0927] This syntax can also be signaled when NumSlicesInPic is greater than 1.
[0928] [The 7th form of the syntax related to the sub-picture]
[0929] Figure 108 It is a syntax structure diagram related to the sub-picture determined by using a grid in the sequence parameter set. In this example, ...
Claims
1. An encoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, wherein, during operation, the circuit encodes the positions and shapes of the respective regions of a plurality of sub - pictures that make up a picture, determined according to constraint conditions, the constraint conditions (i) not allowing one tile among a plurality of tiles that make up the picture to be partially included in one sub - picture among the plurality of sub - pictures; (ii) not allowing one sub - picture among the plurality of sub - pictures to be partially included in one tile among the plurality of tiles; (iii) allowing two or more sub - pictures among the plurality of sub - pictures to be included in one tile among the plurality of tiles; (iv) allowing two or more tiles among the plurality of tiles to be included in one sub - picture among the plurality of sub - pictures.
2. The encoding device according to claim 1, wherein further in the constraint conditions, (i) not allowing one slice among a plurality of slices that make up the picture to be partially included in one sub - picture among the plurality of sub - pictures; (ii) not allowing one sub - picture among the plurality of sub - pictures to be partially included in one slice among the plurality of slices; (iii) not allowing two or more sub - pictures among the plurality of sub - pictures to be included in one slice among the plurality of slices; (iv) allowing two or more slices among the plurality of slices to be included in one sub - picture among the plurality of sub - pictures, and the circuit determines the plurality of tiles, the plurality of sub - pictures, and the plurality of slices according to the constraint conditions.
3. A decoding device, wherein, Comprising: a circuit; and a memory connected to the circuit, wherein, during operation, the circuit decodes the positions and shapes of the respective regions of a plurality of sub - pictures that make up a picture, determined according to constraint conditions, the constraint conditions (i) not allowing one tile among a plurality of tiles that make up the picture to be partially included in one sub - picture among the plurality of sub - pictures; (ii) not allowing one sub - picture among the plurality of sub - pictures to be partially included in one tile among the plurality of tiles; (iii) allowing two or more sub - pictures among the plurality of sub - pictures to be included in one tile among the plurality of tiles; (iv) allowing two or more tiles among the plurality of tiles to be included in one sub - picture among the plurality of sub - pictures.
4. The decoding device according to claim 3, wherein further in the constraint conditions, (i) not allowing one slice among a plurality of slices that make up the picture to be partially included in one sub - picture among the plurality of sub - pictures; (ii) not allowing one sub - picture among the plurality of sub - pictures to be partially included in one slice among the plurality of slices; (iii) not allowing two or more sub - pictures among the plurality of sub - pictures to be included in one slice among the plurality of slices; (iv) allowing two or more slices among the plurality of slices to be included in one sub - picture among the plurality of sub - pictures, and the circuit determines the plurality of tiles, the plurality of sub - pictures, and the plurality of slices according to the constraint conditions.
5. An encoding method, wherein encoding the positions and shapes of the respective regions of a plurality of sub - pictures that make up a picture, determined according to constraint conditions, the constraint conditions (i) It is not allowed that one tile among the multiple tiles constituting the picture is partially included in one sub-picture among the multiple sub-pictures; (ii) It is not allowed that one sub-picture among the multiple sub-pictures is partially included in one tile among the multiple tiles; (iii) It is allowed that more than two sub-pictures among the multiple sub-pictures are included in one tile among the multiple tiles; (iv) It is allowed that more than two tiles among the multiple tiles are included in one sub-picture among the multiple sub-pictures.
6. A decoding method, wherein, decode the positions and shapes of the respective regions of the multiple sub-pictures constituting the picture determined according to the constraint conditions, the constraint conditions (i) It is not allowed that one tile among the multiple tiles constituting the picture is partially included in one sub-picture among the multiple sub-pictures; (ii) It is not allowed that one sub-picture among the multiple sub-pictures is partially included in one tile among the multiple tiles; (iii) It is allowed that more than two sub-pictures among the multiple sub-pictures are included in one tile among the multiple tiles; (iv) It is allowed that more than two tiles among the multiple tiles are included in one sub-picture among the multiple sub-pictures.
7. A non-transitory computer-readable storage medium storing executable commands and a bitstream, wherein, the bitstream contains information indicating whether a picture includes multiple sub-pictures, and the executable commands cause a decoding device to perform a decoding process, in the decoding process, decode the positions and shapes of the respective regions of the multiple sub-pictures determined according to the constraint conditions, the constraint conditions (i) It is not allowed that one tile among the multiple tiles constituting the picture is partially included in one sub-picture among the multiple sub-pictures; (ii) It is not allowed that one sub-picture among the multiple sub-pictures is partially included in one tile among the multiple tiles; (iii) It is allowed that more than two sub-pictures among the multiple sub-pictures are included in one tile among the multiple tiles; (iv) It is allowed that more than two tiles among the multiple tiles are included in one sub-picture among the multiple sub-pictures.
Citation Information
Patent Citations
Method and device for decoding image by using partition unit including additional region
CA3105474A1