Encoding device, decoding device, encoding method, and decoding method
Patent Information
- Application Number
- KR1020227004729
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-03
- Filing Date
- 2020-09-02
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2040-09-02
Smart Images

Figure 112022015558866-PCT00121_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to video coding, and in particular, to a system, components, and methods for encoding and decoding moving images. Background Technology
[0002] Video coding technology is advancing from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec). Along with this advancement, it has become increasingly necessary to provide improvements and optimizations to video coding technology in order to handle the ever-increasing volume of digital video data for various applications. The present disclosure relates to new advancements, improvements, and optimizations in video coding.
[0003] In addition, Non-patent Document 1 relates to an example of a prior standard regarding the video coding technology described above. Prior art literature
[0004] H.265 (ISO / IEC 23008-2 HEVC) / HEVC (High Efficiency Video Coding) The problem to be solved
[0005] Regarding the above-mentioned encoding method, a new method is desired to improve encoding efficiency, improve image quality, reduce throughput, reduce circuit size, or to appropriately select elements or operations such as filters, blocks, sizes, motion vectors, reference pictures, or reference blocks.
[0006] The present disclosure provides a configuration or method that can contribute to one or more of, for example, an improvement in encoding efficiency, an improvement in image quality, a reduction in throughput, a reduction in circuit size, an improvement in processing speed, and the appropriate selection of elements or operations. Additionally, the present disclosure may include a configuration or method that can contribute to benefits other than those mentioned above. means of solving the problem
[0007] For example, an encoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, uses an Intra Block Copy (IBC) mode that references a processing completion area of a picture to which the processing target block belongs in generating a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from a History-based Motion Vector Predictor (HMVP) table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from the HMVP table to the vector candidate list, wherein the HMVP table contains a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, as the HMVP vector candidates. It is stored in a FIFO manner, and the processing target block is encoded using the vector candidate list.
[0008] In video coding technology, proposals for new methods are desired to improve coding efficiency, image quality, and reduce circuit size.
[0009] Each of the embodiments or parts of the configuration or method in the present disclosure enables at least one of, for example, improvement of encoding efficiency, improvement of image quality, reduction of encoding / decoding throughput, reduction of circuit size, or improvement of encoding / decoding processing speed. Alternatively, each of the embodiments or parts of the configuration or method in the present disclosure enables the appropriate selection of components / operations such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. Furthermore, the present disclosure includes the disclosure of configurations or methods capable of providing benefits other than those mentioned above. For example, configurations or methods that improve encoding efficiency while suppressing an increase in throughput.
[0010] The novel advantages and effects in one aspect of the present disclosure become apparent from the specification and drawings. These advantages and / or effects are each obtained by several embodiments and features described in the specification and drawings, but not all of them are required to be provided in order to obtain one or more advantages and / or effects.
[0011] In addition, the general or specific modes thereof may be realized as a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. Effects of the invention
[0012] A configuration or method according to one embodiment of the present disclosure may contribute to one or more of, for example, an improvement in encoding efficiency, an improvement in image quality, a reduction in throughput, a reduction in circuit size, an improvement in processing speed, and an appropriate selection of elements or operations. Additionally, a configuration or method according to one embodiment of the present disclosure may contribute to benefits other than those mentioned above. Brief explanation of the drawing
[0013] FIG. 1 is a schematic diagram showing an example of the configuration of a transmission system according to an embodiment. Figure 2 is a diagram showing an example of the hierarchical structure of data in a stream. Figure 3 is a drawing showing an example of the configuration of a slice. Figure 4 is a drawing showing an example of the composition of a tile. Figure 5 is a diagram showing an example of an encoding structure during scalable encoding. Figure 6 is a diagram showing an example of an encoding structure during scalable encoding. FIG. 7 is a block diagram showing an example of the functional configuration of an encoding device according to an embodiment. FIG. 8 is a block diagram showing an example of an implementation of an encoding device. Figure 9 is a flowchart showing an example of overall encoding processing by an encoding device. Figure 10 is a diagram showing an example of block division. FIG. 11 is a drawing showing an example of the functional configuration of a divided section. Figure 12 is a diagram showing an example of a division pattern. FIG. 13a is a diagram showing an example of a syntax tree of a partitioning pattern. Figure 13b is a diagram showing another example of a syntax tree of a partitioning pattern. Figure 14 is a table showing the transformation basis functions corresponding to each transformation type. Figure 15 is a diagram showing an example of an SVT. Figure 16 is a flowchart showing an example of processing by a conversion unit. Figure 17 is a flowchart showing another example of processing by the conversion unit. FIG. 18 is a block diagram showing an example of the functional configuration of the quantization unit. Figure 19 is a flowchart showing an example of quantization by a quantization unit. FIG. 20 is a block diagram showing an example of the functional configuration of an entropy encoding unit. Figure 21 is a diagram showing the flow of CABAC in the entropy encoding unit. FIG. 22 is a block diagram showing an example of the functional configuration of a loop filter section. FIG. 23a is a drawing showing an example of the shape of a filter used in an ALF (adaptive loop filter). FIG. 23b is a drawing showing another example of the shape of a filter used in ALF. FIG. 23c is a drawing showing another example of the shape of a filter used in ALF. FIG. 23d is a diagram showing an example in which a Y sample (first component) is used in CCALF of Cb and CCALF of Cr (multiple components different from the first component). FIG. 23e is a drawing showing a diamond-shaped filter. Figure 23f is a drawing showing an example of JC-CCALF. Figure 23g is a figure showing an example of a weight_index candidate for JC-CCALF. FIG. 24 is a block diagram showing an example of the detailed configuration of a loop filter section that functions as a DBF. FIG. 25 is a diagram showing an example of a deblocking filter having filter characteristics that are symmetric with respect to block boundaries. FIG. 26 is a diagram illustrating an example of a block boundary where deblocking filter processing is performed. Figure 27 is a diagram showing an example of a Bs value. FIG. 28 is a flowchart showing an example of processing performed in the prediction unit of an encoding device. FIG. 29 is a flowchart showing another example of processing performed in the prediction unit of an encoding device. FIG. 30 is a flowchart showing another example of processing performed in the prediction unit of an encoding device. FIG. 31 is a diagram showing an example of 67 intra prediction modes in intra prediction. Figure 32 is a flowchart showing an example of processing by an intra-prediction unit. FIG. 33 is a drawing showing an example of each reference picture. FIG. 34 is a conceptual diagram showing an example of a reference picture list. Figure 35 is a flowchart showing the flow of the basic processing of inter-prediction. Figure 36 is a flowchart showing an example of MV derivation. Figure 37 is a flowchart showing another example of MV derivation. FIG. 38a is a diagram showing an example of the classification of each mode of MV derivation. FIG. 38b is a diagram showing an example of the classification of each mode of MV derivation. Figure 39 is a flowchart showing an example of inter prediction by normal inter mode. Figure 40 is a flowchart showing an example of inter-prediction by normal merge mode. FIG. 41 is a diagram illustrating an example of MV derivation processing by normal merge mode. FIG. 42 is a diagram illustrating an example of MV derivation processing by HMVP mode. Figure 43 is a flowchart showing an example of FRUC (frame rate up conversion). FIG. 44 is a diagram illustrating an example of pattern matching (bilateral matching) between two blocks following a movement trajectory. FIG. 45 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. FIG. 46a is a diagram illustrating an example of deriving a sub-block unit MV in an affine mode using two control points. FIG. 46b is a diagram illustrating an example of deriving a sub-block unit MV in an affine mode using three control points. FIG. 47a is a conceptual diagram illustrating an example of deriving the MV of a control point in an affine mode. FIG. 47b is a conceptual diagram illustrating an example of deriving the MV of a control point in an affine mode. FIG. 47c is a conceptual diagram illustrating an example of deriving the MV of a control point in an affine mode. FIG. 48a is a diagram illustrating an affine mode having two control points. FIG. 48b is a diagram illustrating an affine mode having three control points. FIG. 49a is a conceptual diagram illustrating an example of a method for deriving the MV of a control point when the number of control points in the encoded complete block and the current block are different. FIG. 49b is a conceptual diagram illustrating another example of a method for deriving the MV of control points when the number of control points in the encoded complete block and the current block is different. Figure 50 is a flowchart showing an example of processing in affine merge mode. Figure 51 is a flowchart showing an example of processing an affine intermode. FIG. 52a is a diagram illustrating the generation of a predicted image of two triangles. FIG. 52b is a conceptual diagram showing the first part of the first partition, and examples of the first sample set and the second sample set. FIG. 52c is a conceptual diagram showing the first part of the first partition. FIG. 53 is a flowchart showing an example of a triangle mode. FIG. 54 is a diagram showing an example of an ATMVP mode in which MV is derived in sub-block units. Figure 55 is a diagram showing the relationship between merge mode and DMVR (dynamic motion vector refreshing). Figure 56 is a conceptual diagram illustrating an example of DMVR. FIG. 57 is a conceptual diagram illustrating another example of a DMVR for determining MV. FIG. 58a is a diagram showing an example of motion detection in DMVR. FIG. 58b is a flowchart showing an example of motion search in DMVR. Figure 59 is a flowchart showing an example of the generation of a predicted image. Figure 60 is a flowchart showing another example of the generation of a predicted image. Figure 61 is a flowchart illustrating an example of predictive image correction processing by OBMC (overlapped block motion compensation). Figure 62 is a conceptual diagram illustrating an example of predictive image correction processing by OBMC. Figure 63 is a diagram illustrating a model assuming uniform linear motion. Figure 64 is a flowchart showing an example of inter prediction based on BIO. FIG. 65 is a diagram showing an example of the functional configuration of an inter prediction unit that performs inter prediction based on BIO. FIG. 66a is a diagram illustrating an example of a method for generating a predicted image using luminance correction processing by local illumination compensation (LIC). FIG. 66b is a flowchart showing an example of a prediction image generation method using luminance correction processing by LIC. FIG. 67 is a block diagram showing the functional configuration of a decoding device according to an embodiment. FIG. 68 is a block diagram showing an example of the implementation of a decoding device. FIG. 69 is a flowchart showing an example of overall decoding processing by a decoding device. FIG. 70 is a diagram showing the relationship between the division decision unit and other components. FIG. 71 is a block diagram showing an example of the functional configuration of an entropy decoding unit. FIG. 72 is a diagram showing the flow of CABAC in the entropy decoding unit. FIG. 73 is a block diagram showing an example of the functional configuration of the inverse quantization unit. FIG. 74 is a flowchart showing an example of inverse quantization by an inverse quantization unit. FIG. 75 is a flowchart showing an example of processing by an inverse transformation unit. Figure 76 is a flowchart showing another example of processing by an inverse transformation unit. FIG. 77 is a block diagram showing an example of the functional configuration of a loop filter section. FIG. 78 is a flowchart showing an example of processing performed in the prediction unit of a decoder. FIG. 79 is a flowchart showing another example of processing performed in the prediction unit of a decoder. FIG. 80a is a flowchart showing part of another example of processing performed in the prediction section of a decoder. FIG. 80b is a flowchart showing the remainder of another example of processing performed in the prediction section of the decoder. FIG. 81 is a diagram showing an example of processing by the intra-prediction unit of a decoder. FIG. 82 is a flowchart showing an example of MV derivation in a decoding device. FIG. 83 is a flowchart showing another example of MV derivation in a decoding device. FIG. 84 is a flowchart showing an example of inter prediction by normal inter mode in a decoder. FIG. 85 is a flowchart showing an example of inter-prediction by normal merge mode in a decoder. FIG. 86 is a flowchart showing an example of inter-prediction by FRUC mode in a decoder. FIG. 87 is a flowchart showing an example of inter-prediction by an affine merge mode in a decoder. FIG. 88 is a flowchart showing an example of inter prediction by affine inter mode in a decoder. FIG. 89 is a flowchart showing an example of inter-prediction by triangle mode in a decoder. FIG. 90 is a flowchart showing an example of motion detection by DMVR in a decoder. FIG. 91 is a flowchart showing a detailed example of motion detection by DMVR in a decoder. FIG. 92 is a flowchart showing an example of the generation of a predicted image in a decoding device. FIG. 93 is a flowchart showing another example of the generation of a predicted image in a decoding device. FIG. 94 is a flowchart showing an example of correction of a predicted image by OBMC in a decoding device. FIG. 95 is a flowchart showing an example of correction of a predicted image by BIO in a decoding device. FIG. 96 is a flowchart showing an example of correction of a predicted image by LIC in a decoding device. Figure 97 is a diagram illustrating the Intra Block Copy (IBC) mode. FIG. 98 is a flowchart showing an example of the operation performed by the encoding device and the decoding device according to the first embodiment. FIG. 99 is a flowchart showing an example of the operation performed by the encoding device and the decoding device according to the second embodiment. FIG. 100 is a flowchart showing the operation performed by an encoding device. FIG. 101 is a flowchart showing the operation performed by the decoding device. FIG. 102 is an overall configuration diagram of a content supply system that realizes a content transmission service. FIG. 103 is a drawing showing an example of a display screen of a web page. FIG. 104 is a drawing showing an example of a display screen of a web page. FIG. 105 is a drawing showing an example of a smartphone. FIG. 106 is a block diagram showing an example of a smartphone configuration. Specific details for implementing the invention
[0014] [Introduction]
[0015] An encoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, uses an Intra Block Copy (IBC) mode that references a processing completion area of a picture to which the processing target block belongs in generating a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from a History-based Motion Vector Predictor (HMVP) table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from the HMVP table to the vector candidate list, wherein the HMVP table contains a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, as FIFO HMVP vector candidates. It is stored in a manner that the processing target block is encoded using the vector candidate list.
[0016] Accordingly, when the size of the block to be processed is below a threshold, the encoding device generates a vector candidate list without performing a first pruning process (hereinafter simply referred to as pruning process), thereby reducing throughput. Consequently, the encoding efficiency of the encoding device is improved.
[0017] For example, the circuit may generate a predicted image of the block to be processed using a second vector, encode the second vector using one of a plurality of vector candidates included in the vector candidate list, and in the first pruning process of generating the vector candidate list, for each of the plurality of HMVP vector candidates stored in the HMVP table, determine whether the HMVP vector candidate matches any one of the one or more vector candidates already registered in the vector candidate list, and if the HMVP vector candidate does not match any of the one or more vector candidates, generate the vector candidate list by registering the HMVP vector candidate in the vector candidate list.
[0018] As a result, the encoding device generates a vector candidate list without performing pruning when the size of the block to be processed is below a threshold, thereby reducing throughput. Consequently, the encoding efficiency of the encoding device is improved.
[0019] For example, the circuit may also update the HMVP table using a second vector candidate having information regarding the second vector, and in updating the HMVP table, determine whether the size of the block to be processed is less than or equal to the threshold; if the size of the block to be processed is less than or equal to the threshold, update the HMVP table without performing the second pruning process, and if the size of the block to be processed is greater than the threshold, perform the second pruning process and update the HMVP table.
[0020] Accordingly, when the size of the block to be processed is below the threshold, the encoding device updates the HMVP table without performing a second pruning process (hereinafter simply referred to as pruning process), thereby reducing the throughput. Therefore, the encoding device improves encoding efficiency.
[0021] For example, in the second pruning process of updating the HMVP table, the circuit may determine whether the second vector candidate matches any one of the plurality of HMVP vector candidates, and if the second vector candidate does not match any of the plurality of HMVP vector candidates, update the HMVP table by storing the second vector candidate in the HMVP table.
[0022] By doing so, the encoding device can register a more suitable vector candidate for the block to be processed in the vector candidate list.
[0023] For example, the size of the processing target block may be defined as the number of pixels within the processing target block. For example, the threshold may be 16 pixels.
[0024] Accordingly, when the number of pixels in the block to be processed is below a threshold (e.g., 16 pixels or less), the encoding device generates a vector candidate list and updates the HMVP table without performing pruning, thereby reducing throughput. Therefore, the encoding device improves encoding efficiency.
[0025] For example, the size of the block to be processed may be defined as at least one of the width and height of the block to be processed. For example, the threshold may be 4×4 pixels.
[0026] Accordingly, when at least one of the width and height of a block to be processed is below a threshold (for example, when both the width and height of a block to be processed are 4 pixels or less), the encoding device generates a vector candidate list and updates the HMVP table without performing pruning processing, thereby reducing throughput. Therefore, the encoding device improves encoding efficiency.
[0027] For example, the above vector candidate list does not need to be shared between the processing target block and a block adjacent to the processing target block.
[0028] As a result, the encoding device uses a vector candidate list generated for each block to be processed, thereby improving prediction accuracy.
[0029] For example, in generating a predicted image of the block to be processed, the circuit may, when it is decided to use the IBC mode among a plurality of prediction modes, use the IBC mode; when it is decided to use a prediction mode different from the IBC mode among the plurality of prediction modes, use the prediction mode different from the IBC mode; and when a prediction mode different from the IBC mode is used, perform the first pruning process and register the HMVP vector candidate from the HMVP table in the vector candidate list to generate the vector candidate list, and encode the block to be processed using the vector candidate list.
[0030] Accordingly, the encoding device can appropriately switch the conditions for performing pruning processing depending on the prediction mode used. Therefore, the encoding device improves encoding efficiency.
[0031] Additionally, a decoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein, in operation, the circuit uses an IBC mode that references a processing completion area of a picture to which the processing target block belongs in generating a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and the processing target Decode the block.
[0032] Accordingly, the decoder generates a vector candidate list without performing pruning when the size of the block to be processed is below a threshold, thereby reducing throughput. Consequently, the processing efficiency of the decoder is improved.
[0033] For example, the circuit may generate a predicted image of the block to be processed using a second vector, decode the second vector using one of a plurality of vector candidates included in the vector candidate list, and in the first pruning process of generating the vector candidate list, for each of the plurality of HMVP vector candidates stored in the HMVP table, determine whether the HMVP vector candidate matches any one of the one or more vector candidates already registered in the vector candidate list, and if the HMVP vector candidate does not match any of the one or more vector candidates, the vector candidate list may be generated by registering the HMVP vector candidate in the vector candidate list.
[0034] Accordingly, the decoder generates a vector candidate list without performing pruning when the size of the block to be processed is below a threshold, thereby reducing throughput. Consequently, the processing efficiency of the decoder is improved.
[0035] For example, the circuit may also update the HMVP table using a second vector candidate having information regarding the second vector, and in updating the HMVP table, determine whether the size of the block to be processed is less than or equal to the threshold; if the size of the block to be processed is less than or equal to the threshold, update the HMVP table without performing the second pruning process, and if the size of the block to be processed is greater than the threshold, perform the second pruning process and update the HMVP table.
[0036] Accordingly, when the size of the block to be processed is below the threshold, the decoder updates the HMVP table without performing pruning, thereby reducing throughput. Therefore, the decoder improves processing efficiency.
[0037] For example, in the second pruning process of updating the HMVP table, the circuit may determine whether the second vector candidate matches any one of the plurality of HMVP vector candidates, and if the second vector candidate does not match any of the plurality of HMVP vector candidates, update the HMVP table by storing the second vector candidate in the HMVP table.
[0038] Accordingly, the decoding device can register a more suitable vector candidate for the block to be processed in the vector candidate list.
[0039] For example, the size of the processing target block may be defined as the number of pixels within the processing target block. For example, the threshold may be 16 pixels.
[0040] Accordingly, when the number of pixels in the block to be processed is below a threshold (e.g., 16 pixels or less), the decoder does not perform pruning processing and instead generates a vector candidate list and updates the HMVP table, thereby reducing throughput. Therefore, the decoder improves processing efficiency.
[0041] For example, the size of the block to be processed may be defined as at least one of the width and height of the block to be processed. For example, the threshold may be 4×4 pixels.
[0042] Accordingly, when at least one of the width and height of the block to be processed is below a threshold (for example, when both the width and height of the block to be processed are 4 pixels or less), the decoder does not perform pruning processing and instead generates a vector candidate list and updates the HMVP table, thereby reducing throughput. Therefore, the decoder improves processing efficiency.
[0043] For example, the above vector candidate list does not need to be shared between the processing target block and a block adjacent to the processing target block.
[0044] As a result, the decoding device uses a vector candidate list generated for each block to be processed, thereby improving prediction accuracy.
[0045] For example, in generating a predicted image of the block to be processed, the circuit may, when it is decided to use the IBC mode among a plurality of prediction modes, use the IBC mode, when it is decided to use a prediction mode different from the IBC mode, when it is decided to use a prediction mode different from the IBC mode, when it is used a prediction mode different from the IBC mode, perform the first pruning process and register the HMVP vector candidate from the HMVP table in the vector candidate list to generate the vector candidate list, and decode the block to be processed using the vector candidate list.
[0046] Accordingly, the decoder can appropriately switch the conditions for performing pruning processing depending on the prediction mode used. Therefore, the processing efficiency of the decoder is improved.
[0047] In addition, an encoding method according to one aspect of the present disclosure, in generating a predicted image of a block to be processed, uses an IBC mode that references a processing completion area of a picture to which the block to be processed belongs, determines whether the size of the block to be processed, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the block to be processed is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the block to be processed is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and the block to be processed is encoded using the vector candidate list.
[0048] Accordingly, the device executing the encoding method generates a vector candidate list without performing pruning when the size of the block to be processed is below a threshold, thereby reducing throughput. Consequently, the device executing the encoding method improves encoding efficiency.
[0049] In addition, a decoding method according to one aspect of the present disclosure, in generating a predicted image of a block to be processed, uses an IBC mode that references a processing completion area of a picture to which the block to be processed belongs, determines whether the size of the block to be processed, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the block to be processed is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the block to be processed is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and decodes the block to be processed using the vector candidate list.
[0050] Accordingly, the device executing the decoding method generates a vector candidate list without performing pruning when the size of the block to be processed is below a threshold, thereby reducing throughput. Consequently, the device executing the decoding method improves processing efficiency.
[0051] [Definition of Terms]
[0052] Each term may be defined as follows, by way of example.
[0053] (1) Image
[0054] It is a unit of data composed of a set of pixels, consisting of a picture or a block smaller than a picture, and includes still images in addition to moving images.
[0055] (2) Picture
[0056] It is a processing unit of an image composed of a set of pixels, and is sometimes called a frame or field.
[0057] (3) Block
[0058] It is a processing unit of a set containing a specific number of pixels, and as listed in the examples below, the name is irrelevant. Furthermore, the shape is irrelevant; for example, it includes not only rectangles composed of M×N pixels and squares composed of M×M pixels, but also triangles, circles, and other shapes.
[0059] (Example of a block)
[0060] · Slice / Tile / Brick
[0061] · CTU / Super Block / Basic Partition Unit
[0062] · VPDU / Hardware processing partitioning unit
[0063] ·CU / Processing Block Unit / Prediction Block Unit (PU) / Orthogonal Transformation Block Unit (TU) / Unit
[0064] · Sub-block
[0065] (4) Pixels / Samples
[0066] It is the smallest unit point that constitutes an image, and includes not only pixels at integer positions but also pixels at fractional positions generated based on pixels at integer positions.
[0067] (5) Pixel value / Sample value
[0068] It is an intrinsic value of a pixel and includes luminance value, color difference value, RGB gradation, as well as depth value or 2 values of 0 and 1.
[0069] (6) Flag
[0070] In addition to 1 bit, cases involving multiple bits are included, for example, parameters or indices of 2 bits or more may be used. Also, not only 2 values using binary numbers may be used, but also multi-values using other bases may be used.
[0071] (7) signal
[0072] It is encoded and coded to transmit information, and includes analog signals that take continuous values in addition to discretized digital signals.
[0073] (8) Stream / Bit stream
[0074] It refers to a sequence of digital data or a flow of digital data. A stream / bit stream may be composed of multiple streams divided into multiple layers in addition to a single stream. Furthermore, in addition to cases where it is transmitted by serial communication through a single transmission path, cases where it is transmitted by packet communication through multiple transmission paths are also included.
[0075] (9) Cha / Cha
[0076] For scalar quantities, in addition to simple difference (xy), it is sufficient to include difference operations, including absolute difference (|xy|), square difference (x^2-y^2), square root of difference (√(xy)), weighted difference (ax-by : a, b is a constant), and offset difference (x-y + a : a is an offset).
[0077] (10) total
[0078] For scalar quantities, in addition to simple sum (x + y), it is sufficient to include sum operations, including absolute sum (|x + y|), sum of squares (x^2 + y^2), square root of sum (√(x + y)), weighted sum (ax + by : a, b is a constant), and offset sum (x + y + a : a is an offset).
[0079] (11) Based on
[0080] It also includes cases where elements other than the subject of reliance are added. In addition, it includes cases where a result is obtained via an intermediate result, in addition to cases where a result is obtained directly.
[0081] (12) using
[0082] It also includes cases where elements other than the target element are added. Furthermore, it includes cases where a result is obtained via an intermediate result, in addition to cases where a result is obtained directly.
[0083] (13) prohibit (prohibit, forbid)
[0084] It can be rephrased as not being permitted. Furthermore, what is not prohibited or is permitted does not necessarily imply an obligation.
[0085] (14) Limit (limit, restriction / restrict / restricted)
[0086] It can be rephrased as not being permitted. Furthermore, what is not prohibited or is permitted does not necessarily imply an obligation. Additionally, it is sufficient if only a portion is prohibited quantitatively or qualitatively, and cases of complete prohibition are also included.
[0087] (15) chroma
[0088] It is an adjective represented by the symbols Cb and Cr, which specifies that a sample array or a single sample represents one of two color difference signals related to primary colors. The term chrominance may also be used instead of the term chroma.
[0089] (16) Luminance (luma)
[0090] It is an adjective represented by a symbol or subscript Y or L that specifies that an array of samples or a single sample represents a black-and-white signal related to primary colors. Instead of the term luma, the term luminance may also be used.
[0091] [Explanation regarding the entry]
[0092] In drawings, the same reference number indicates the same or similar components. Also, the size and relative position of components in the drawings are not necessarily drawn to a fixed scale.
[0093] Embodiments will be described in detail below with reference to the drawings. Furthermore, the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection types of components, steps, relationships and sequences of steps, etc., shown in the embodiments below are examples and are not well known facts that limit the scope of the claims.
[0094] Hereinafter, embodiments of an encoding device and a decoding device are described. The embodiments are examples of encoding devices and decoding devices to which the processing and / or configurations described in each aspect of the present disclosure are applicable. The processing and / or configurations may also be implemented in encoding devices and decoding devices different from the embodiments. For example, regarding the processing and / or configurations applied to the embodiments, any one of the following may be implemented, for example.
[0095] (1) Any one of the plurality of components of the encoding device or decoding device of the embodiment described in each embodiment of the present disclosure may be substituted or combined with another component described in any one embodiment of the present disclosure.
[0096] (2) In the encoding device or decoding device of the embodiment, any modifications such as the addition, substitution, or deletion of functions or processing may be made to the functions or processing performed by some of the components of the plurality of components of the encoding device or decoding device. For example, any one function or processing may be substituted or combined with another function or processing described in any one of the embodiments of the present disclosure.
[0097] (3) In the method performed by the encoding device or decoding device of the embodiment, any modifications such as addition, substitution, and deletion may be made to some of the processing included in the method. For example, any one of the processing in the method may be substituted or combined with another processing described in any one of the embodiments of the present disclosure.
[0098] (4) Some of the components of the plurality of components constituting the encoding device or decoding device of the embodiment may be combined with the components described in any one of the embodiments of the present disclosure, may be combined with the components having a part of the function described in any one of the embodiments of the present disclosure, and may be combined with the components performing a part of the processing performed by the components described in each embodiment of the present disclosure.
[0099] (5) A component having part of the function of an encoding device or a decoder of an embodiment, or a component performing part of the processing of an encoding device or a decoder of an embodiment, may be combined or substituted with a component described in any one of the embodiments of the present disclosure, a component having part of the function described in any one of the embodiments of the present disclosure, or a component performing part of the processing described in any one of the embodiments of the present disclosure.
[0100] (6) In the method of the encoding device or decoding device of the embodiment, any one of the plurality of processes included in the method may be substituted or combined with the process described in any one of the embodiments of the present disclosure, or with any one of the same processes.
[0101] (7) Some of the processing included in the method of the encoding device or decoding device of the embodiment may be combined with the processing described in any one of the embodiments of the present disclosure.
[0102] (8) The method of implementing the processing and / or configuration described in each embodiment of the present disclosure is not limited to the encoding device or decoding device of the embodiment. For example, the processing and / or configuration may be implemented in a device used for a purpose different from the video encoding or video decoding disclosed in the embodiment.
[0103] [System Configuration]
[0104] FIG. 1 is a schematic diagram showing an example of the configuration of a transmission system according to the present embodiment.
[0105] A transmission system (Trs) is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system (Trs) includes, for example, an encoding device (100), a network (Nw), and a decoder (200), as shown in FIG. 1.
[0106] An image is input to the encoding device (100). The encoding device (100) generates a stream by encoding the input image and outputs the stream to a network (Nw). The stream includes, for example, an encoded image and control information for decoding the encoded image. The image is compressed by this encoding.
[0107] Additionally, the original image before encoding, which is input to the encoding device (100), is also referred to as the original image, original signal, or original sample. Additionally, the image may be a moving image or a still image. Additionally, the image is a higher-level concept such as a sequence, picture, and block, and is not subject to limitations in spatial and temporal domains unless otherwise specified. Additionally, the image is composed of pixels or an array of pixel values, and the signal or pixel value representing the image is also referred to as a sample. Additionally, the stream may be referred to as a bit stream, an encoded bit stream, a compressed bit stream, or an encoded signal. Additionally, the encoding device may be referred to as an image encoding device or a moving image encoding device, and the encoding method by the encoding device (100) may be referred to as an encoding method, an image encoding method, or a moving image encoding method.
[0108] The network (Nw) transmits the stream generated by the encoding device (100) to the decoding device (200). The network (Nw) may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or a combination thereof. The network (Nw) is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. In addition, the network (Nw) may be replaced by a storage medium that records the stream, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).
[0109] The decoding device (200) generates a decoded image, for example, an uncompressed image, by decoding a stream transmitted by the network (Nw). For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method by the encoding device (100).
[0110] Additionally, the decoding device may be called an image decoding device or a video decoding device, and the decoding method by the decoding device (200) may be called a decoding method, an image decoding method, or a video decoding method.
[0111] [Data Structure]
[0112] FIG. 2 is a diagram illustrating an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. This video sequence includes, for example, a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and a plurality of pictures, as shown in FIG. 2 (a).
[0113] A VPS includes, in a video image composed of multiple layers, encoding parameters common to multiple layers and encoding parameters related to multiple layers included in the video image, or to each layer.
[0114] The SPS includes parameters used for the sequence, that is, encoding parameters referenced by the decoder (200) to decode the sequence. For example, the encoding parameters may represent the width or height of the picture. Additionally, there may be multiple SPSs.
[0115] PPS includes parameters used for a picture, that is, encoding parameters referenced by a decoder (200) to decode each picture in a sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. Additionally, there may be multiple PPSs. Also, SPS and PPS are sometimes simply referred to as parameter sets.
[0116] The picture may include a picture header and one or more slices, as shown in FIG. 2(b). The picture header includes encoding parameters that a decoder (200) references to decode the one or more slices.
[0117] A slice includes a slice header and one or more bricks, as shown in (c) of FIG. 2. The slice header includes encoding parameters that a decoder (200) references to decode the one or more bricks.
[0118] The brick includes one or more Coding Tree Units (CTUs), as shown in (d) of FIG. 2.
[0119] Additionally, a picture may not include a slice, and instead include a tile group. In this case, the tile group includes one or more tiles. Also, a brick may include a slice.
[0120] A CTU is also called a super block or basic partition unit. As shown in FIG. 2(e), such a CTU includes a CTU header and one or more CUs (Coding Units). The CTU header includes encoding parameters that a decoder (200) references to decode one or more CUs.
[0121] A CU may be divided into multiple smaller CUs. Additionally, as shown in (f) of FIG. 2, a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residuals described later. Furthermore, a CU is basically identical to a PU (Prediction Unit) and a TU (Transform Unit), but for example, in the SBT described later, it may include multiple TUs smaller than the CU. Also, a CU may be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is a fixed unit that can be processed in one stage when performing pipeline processing in hardware, for example.
[0122] Additionally, the stream does not need to have any part of any of the layers shown in FIG. 2. Also, the order of these layers may be changed, and any one layer may be replaced with another layer. Also, a picture that is the subject of processing performed at the current time by a device such as an encoding device (100) or a decoding device (200) is called a current picture. If the processing is encoding, the current picture is identical to the picture to be encoded, and if the processing is decoding, the current picture is identical to the picture to be decoded. Also, a block such as a CU or CU that is the subject of processing performed at the current time by a device such as an encoding device (100) or a decoding device (200) is called a current block. If the processing is encoding, the current block is identical to the block to be encoded, and if the processing is decoding, the current block is identical to the block to be decoded.
[0123] [Picture composition slices / tiles]
[0124] To decode pictures in parallel, pictures may be organized into slice units or tile units.
[0125] A slice is a basic encoding unit that constitutes a picture. A picture consists of, for example, one or more slices. Also, a slice consists of one or more consecutive CTUs.
[0126] FIG. 3 is a diagram illustrating an example of the configuration of a slice. For example, a picture contains 11 × 8 CTUs and is also divided into four slices (slices 1 to 4). Slice 1 consists of, for example, 16 CTUs, slice 2 consists of, for example, 21 CTUs, slice 3 consists of, for example, 29 CTUs, and slice 4 consists of, for example, 22 CTUs. Here, each CTU within the picture belongs to one of the slices. The shape of the slice is a form in which the picture is divided in the horizontal direction. The boundary of the slice does not have to be at the edge of the screen, but can be any of the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs within the slice is, for example, raster scan order. In addition, the slice includes a slice header and encoded data. The slice header may describe the characteristics of the slice, such as the CTU address at the beginning of the slice and the slice type.
[0127] A tile is a unit of a rectangular area that makes up a picture. Each tile may be assigned a number called a TileId in raster or scan order.
[0128] FIG. 4 is a diagram showing an example of the configuration of tiles. For example, a picture includes 11×8 CTUs and is also divided into four rectangular tile areas (tiles 1 to 4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs within the picture are processed, for example, in a raster scan order. When tiles are used, for each of the multiple tiles, at least one CTU is processed, for example, in a raster scan order. For example, as shown in FIG. 4, the processing order of multiple CTUs included in Tile 1 is from the left end of the first column of Tile 1 to the right end of the first column of Tile 1, and then from the left end of the second column of Tile 1 to the right end of the second column of Tile 1.
[0129] In addition, one tile may include one or more slices, and one slice may include one or more tiles.
[0130] Additionally, the picture may be composed of tile sets. A tile set may include one or more tile groups and one or more tiles. The picture may be composed of only one of a tile set, a tile group, and a tile. For example, the order in which multiple tiles are scanned in raster order for each tile set is the basic encoding order of the tiles. A group of one or more tiles that have consecutive basic encoding orders within each tile set is called a tile group. Such a picture may be composed of a dividing unit (102) (see FIG. 7) described later.
[0131] [Scalable Encoding]
[0132] Figures 5 and 6 are diagrams illustrating an example of the configuration of a scalable stream.
[0133] As shown in FIG. 5, the encoding device (100) may generate a temporally and spatially scalable stream by encoding each of a plurality of pictures by dividing them into one of a plurality of layers. For example, the encoding device (100) realizes scalability in which an enhancement layer exists above the base layer by encoding the picture for each layer. Such encoding of each picture is called scalable encoding. Accordingly, the decoding device (200) can switch the quality of the image displayed by decoding the stream. That is, the decoding device (200) determines which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication band. As a result, the decoding device (200) can freely switch and decode the same content into low-resolution content and high-resolution content. For example, a user of the stream watches the video of the stream up to a certain point using a smartphone while on the move, and continues watching the video using a device such as an internet TV after returning home. Additionally, each of the aforementioned smartphone and device is equipped with a decoding device (200) having the same or different performance. In this case, if the device decodes up to the upper layer of the stream, the user can watch a high-quality video after returning home. As a result, the encoding device (100) does not need to generate multiple streams with different quality for the same content, thereby reducing the processing load.
[0134] Additionally, the enhancement layer may include meta-information based on statistical information of the image, etc. The decoder (200) may generate an image with improved quality by super-resolution of the picture of the base layer based on the meta-information. Super-resolution may be either an improvement in the SN ratio at the same resolution or an expansion of the resolution. The meta-information may include information for specifying linear or non-linear filter coefficients used for super-resolution processing, or information for specifying parameter values in filter processing, machine learning, or least squares operation used for super-resolution processing.
[0135] Alternatively, depending on the meaning of each object within the picture, the picture may be divided into tiles, etc. In this case, the decoding device (200) may decode only a portion of the picture area by selecting the tile to be decoded. Also, the attributes of the object (person, car, ball, etc.) and the location within the picture (coordinate location within the same picture, etc.) may be stored as meta information. In this case, the decoding device (200) may specify the location of the desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 6, the meta information is stored using a data storage structure different from image data, such as SEI in HEVC. This meta information indicates, for example, the location, size, or color of the main object.
[0136] In addition, meta-information may be stored in units composed of multiple pictures, such as streams, sequences, or random access units. Accordingly, the decoding device (200) can acquire the time at which a specific person appears in the video image, and by acquiring the time and picture unit information, it can identify the picture in which the object exists and the location of the object within that picture.
[0137] [Encoding device]
[0138] Next, an encoding device (100) according to an embodiment will be described. FIG. 7 is a block diagram showing an example of the functional configuration of an encoding device (100) according to an embodiment. The encoding device (100) encodes an image in blocks.
[0139] As shown in FIG. 7, the encoding device (100) is a device that encodes an image in block units and comprises a splitting unit (102), a subtraction unit (104), a conversion unit (106), a quantization unit (108), an entropy encoding unit (110), an inverse quantization unit (112), an inverse conversion unit (114), an addition unit (116), a block memory (118), a loop filter unit (120), a frame memory (122), an intra prediction unit (124), an inter prediction unit (126), a prediction control unit (128), and a prediction parameter generation unit (130). Additionally, the intra prediction unit (124) and the inter prediction unit (126) are each configured as part of a prediction processing unit.
[0140] [Example of Encoding Device Implementation]
[0141] FIG. 8 is a block diagram showing an example of implementation of an encoding device (100). The encoding device (100) includes a processor (a1) and a memory (a2). For example, a plurality of components of the encoding device (100) shown in FIG. 7 are implemented by the processor (a1) and memory (a2) shown in FIG. 8.
[0142] A processor (a1) is a circuit that performs information processing and is capable of accessing memory (a2). For example, the processor (a1) is a dedicated or general-purpose electronic circuit that encodes images. The processor (a1) may be a processor such as a CPU. Also, the processor (a1) may be an assembly of multiple electronic circuits. Also, for example, the processor (a1) may serve as a component of the encoding device (100) shown in FIG. 7, excluding the component for storing information.
[0143] Memory (a2) is a dedicated or general-purpose memory in which information for the processor (a1) to encode an image is stored. Memory (a2) may be an electronic circuit and may be connected to the processor (a1). Also, memory (a2) may be included in the processor (a1). Also, memory (a2) may be an assembly of multiple electronic circuits. Also, memory (a2) may be a magnetic disk or an optical disk, etc., and may be expressed as a storage or a recording medium. Also, memory (a2) may be a non-volatile memory or a volatile memory.
[0144] For example, the memory (a2) may store an image to be encoded, or a stream corresponding to the encoded image. Also, the memory (a2) may store a program for the processor (a1) to encode the image.
[0145] Also, for example, memory (a2) may serve as a component for storing information among the plurality of components of the encoding device (100) shown in FIG. 7. Specifically, memory (a2) may serve as a block memory (118) and frame memory (122) shown in FIG. 7. More specifically, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.) may be stored in memory (a2).
[0146] In addition, in the encoding device (100), not all of the plurality of components shown in FIG. 7 are implemented, and not all of the plurality of processes described above are performed. Some of the plurality of components shown in FIG. 7 may be included in another device, and some of the plurality of processes described above may be performed by another device.
[0147] Below, after explaining the overall processing flow of the encoding device (100), each component included in the encoding device (100) is described.
[0148] [Overall flow of encoding processing]
[0149] FIG. 9 is a flowchart showing an example of overall encoding processing by an encoding device (100).
[0150] First, the dividing unit (102) of the encoding device (100) divides the picture included in the original image into a plurality of fixed-size blocks (128×128 pixels) (step Sa_1). Then, the dividing unit (102) selects a dividing pattern for the fixed-size blocks (step Sa_2). That is, the dividing unit (102) further divides the fixed-size blocks into a plurality of blocks constituting the selected dividing pattern. Then, the encoding device (100) performs the processing of steps Sa_3 to Sa_9 for each of the plurality of blocks.
[0151] A prediction processing unit consisting of an intra prediction unit (124) and an inter prediction unit (126), and a prediction control unit (128) generate a prediction image of a current block (step Sa_3). Additionally, the prediction image is also referred to as a prediction signal, a prediction block, or a prediction sample.
[0152] Next, the subtraction unit (104) generates the difference between the current block and the predicted image as the predicted residual (step Sa_4). Also, the predicted residual is also called the prediction error.
[0153] Next, the conversion unit (106) and the quantization unit (108) generate a plurality of quantization coefficients by performing conversion and quantization on the predicted image (step Sa_5).
[0154] Next, the entropy encoding unit (110) generates a stream by performing encoding (specifically entropy encoding) on the plurality of quantization coefficients and prediction parameters regarding the generation of a prediction image (step Sa_6).
[0155] Next, the inverse quantization unit (112) and the inverse transformation unit (114) restore the predicted residual by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (step Sa_7).
[0156] Next, the adder (116) reconstructs the current block by adding the prediction image to the restored prediction residual (step Sa_8). By doing so, a reconstructed image is generated. Additionally, the reconstructed image is also called a reconstructed block, and in particular, the reconstructed image generated by the encoding device (100) is also called a local decoding block or a local decoding image.
[0157] When this reconstructed image is generated, the loop filter unit (120) performs filtering on the reconstructed image as needed (step Sa_9).
[0158] Then, the encoding device (100) determines whether the encoding of the entire picture is complete (step Sa_10), and if it determines that it is not complete (No in step Sa_10), it repeats the processing from step Sa_2.
[0159] In addition, in the example described above, the encoding device (100) selects one division pattern for a fixed-size block and performs encoding of each block according to the division pattern, but may also perform encoding of each block according to each of a plurality of division patterns. In this case, the encoding device (100) evaluates the cost for each of the plurality of division patterns and may, for example, select the stream obtained by encoding according to the division pattern with the smallest cost as the final output stream.
[0160] Additionally, the processing of these steps Sa_1 to Sa_10 may be performed sequentially by the encoding device (100), and some of the processing may be performed in parallel, or the order may be changed.
[0161] The encoding process by this encoding device (100) is a hybrid encoding using predictive encoding and transform encoding. Additionally, predictive encoding is performed by an encoding loop consisting of a subtraction unit (104), a transform unit (106), a quantization unit (108), an inverse quantization unit (112), an inverse transform unit (114), an adder unit (116), a loop filter unit (120), a block memory (118), a frame memory (122), an intra prediction unit (124), an inter prediction unit (126), and a prediction control unit (128). That is, the prediction processing unit consisting of the intra prediction unit (124) and the inter prediction unit (126) constitutes a part of the encoding loop.
[0162] [Divided Section]
[0163] The dividing unit (102) divides each picture included in the original image into multiple blocks and outputs each block to the subtraction unit (104). For example, the dividing unit (102) first divides the picture into blocks of a fixed size (e.g., 128×128 pixels). These fixed-size blocks may be called encoding tree units (CTU). Then, the dividing unit (102) divides each of the fixed-size blocks into blocks of a variable size (e.g., 64×64 pixels or less) based on, for example, recursive quad tree and / or binary tree block division. That is, the dividing unit (102) selects a division pattern. These variable-size blocks may be called encoding units (CU), prediction units (PU), or conversion units (TU). In addition, in various implementation examples, CU, PU, and TU do not need to be distinguished, and some or all blocks within the picture may become processing units of CU, PU, or TU.
[0164] FIG. 10 is a drawing showing an example of block division in an embodiment. In FIG. 10, solid lines represent block boundaries by quaternary tree block division, and dashed lines represent block boundaries by binary tree block division.
[0165] Here, block 10 is a square block of 128×128 pixels. This block 10 is first divided into four square blocks of 64×64 pixels (quadratic tree block division).
[0166] The 64×64 pixel square block at the top left is further divided vertically into two rectangular blocks each consisting of 32×64 pixels, and the 32×64 pixel rectangular block on the left is further divided vertically into two rectangular blocks each consisting of 16×64 pixels (binary tree block division). As a result, the 64×64 pixel square block at the top left is divided into two 16×64 pixel rectangular blocks 11 and 12 and a 32×64 pixel rectangular block 13.
[0167] The 64×64 pixel square block at the top right is horizontally divided into two rectangular blocks 14 and 15, each consisting of 64×32 pixels (binary tree block division).
[0168] The 64×64 pixel square block at the bottom left is divided into four square blocks, each consisting of 32×32 pixels (quadratic tree block division). Among the four square blocks, each consisting of 32×32 pixels, the block at the top left and the block at the bottom right are further divided. The 32×32 pixel square block at the top left is vertically divided into two rectangular blocks, each consisting of 16×32 pixels, and the 16×32 pixel rectangular block on the right is further horizontally divided into two square blocks, each consisting of 16×16 pixels (binary tree block division). The 32×32 pixel square block at the bottom right is horizontally divided into two rectangular blocks, each consisting of 32×16 pixels (binary tree block division). As a result, the 64×64 pixel square block at the bottom left is divided into a 16×32 pixel rectangular block 16, two 16×16 pixel square blocks 17 and 18, two 32×32 pixel square blocks 19 and 20, and two 32×16 pixel rectangular blocks 21 and 22.
[0169] Block 23, consisting of 64×64 pixels at the bottom right, is not divided.
[0170] As described above, in FIG. 10, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quad-tree and binary tree block partitioning. This partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.
[0171] In addition, in FIG. 10, one block was divided into four or two blocks (quadruple tree or binary tree block division), but the division is not limited to these. For example, one block may be divided into three blocks (ternary tree block division). A division including such ternary tree block division is sometimes called MBT (multi-type tree) division.
[0172] FIG. 11 is a drawing showing an example of the functional configuration of a dividing unit (102). As shown in FIG. 11, the dividing unit (102) may be equipped with a block dividing decision unit (102a). The block dividing decision unit (102a) may perform the following processing as an example.
[0173] The block division determination unit (102a) collects block information from, for example, a block memory (118) or a frame memory (122) and determines the division pattern described above based on the block information. The division unit (102) divides the original image according to the division pattern and outputs one or more blocks obtained by the division to the subtraction unit (104).
[0174] Additionally, the block division decision unit (102a) outputs a parameter representing the division pattern described above, for example, to the conversion unit (106), the inverse conversion unit (114), the intra prediction unit (124), the inter prediction unit (126), and the entropy encoding unit (110). The conversion unit (106) may convert the prediction residual based on the parameter, and the intra prediction unit (124) and the inter prediction unit (126) may generate a prediction image based on the parameter. Additionally, the entropy encoding unit (110) may perform entropy encoding on the parameter.
[0175] Parameters regarding the partitioning pattern may be described in the stream as follows, for example.
[0176] FIG. 12 is a diagram showing examples of division patterns. Division patterns include, for example, a 4-division (QT) in which a block is divided into two parts in the horizontal direction and a vertical direction, a 3-division (HT or VT) in which a block is divided in the same direction in a ratio of 1:2:1, a 2-division (HB or VB) in which a block is divided in the same direction in a ratio of 1:1, and no division (NS).
[0177] In addition, in the case of 4 divisions and no divisions, the division pattern does not have a block division direction, and in the case of 2 divisions and 3 divisions, the division pattern has division direction information.
[0178] FIGS. 13a and FIGS. 13b are diagrams illustrating an example of a syntax tree of a splitting pattern. In the example of FIG. 13a, first, there is information indicating whether to perform a split (S: Split flag), and next, there is information indicating whether to perform a 4-part split (QT: QT flag). Next, there is information indicating whether to perform a 3-part split or a 2-part split (TT: TT flag or BT: BT flag), and finally, there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag). Additionally, for each of the one or more blocks obtained by the splitting according to this splitting pattern, the splitting may be repeated using the same processing. That is, as an example, a determination may be recursively made as to whether to perform a division, whether to perform a 4-division, whether the division method is horizontal or vertical, and whether to perform a 3-division or a 2-division, and the result of the determination may be encoded into a stream according to the encoding order disclosed in the syntax tree shown in FIG. 13a.
[0179] Also, in the syntax tree shown in FIG. 13a, the information is arranged in the order of S, QT, TT, and Ver, but the information may also be arranged in the order of S, QT, Ver, and BT. That is, in the example of FIG. 13b, first, there is information indicating whether to perform a split (S: Split flag), and next, there is information indicating whether to perform a 4-part split (QT: QT flag). Next, there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag), and finally, there is information indicating whether to perform a 2-part split or a 3-part split (BT: BT flag or TT: TT flag).
[0180] In addition, the partitioning pattern described herein is an example, and you may use a partitioning pattern other than the one described, or use only a part of the partitioning pattern described.
[0181] [Decommissioning Department]
[0182] The subtraction unit (104) subtracts the predicted image (the predicted image input from the prediction control unit (128)) from the original image in block units that are input from the division unit (102) and divided by the division unit (102). That is, the subtraction unit (104) calculates the predicted residual of the current block. Then, the subtraction unit (104) outputs the calculated predicted residual to the conversion unit (106).
[0183] The original image is an input signal of the encoding device (100), and is, for example, a signal representing the image of each picture constituting the moving image (for example, a luminance (luma) signal and two chroma (chroma) signals).
[0184] [Conversion Section]
[0185] The conversion unit (106) converts the predicted residual in the spatial domain into a conversion coefficient in the frequency domain and outputs the conversion coefficient to the quantization unit (108). Specifically, the conversion unit (106) performs, for example, a predetermined discrete cosine transform (DCT) or discrete sine transform (DST) on the predicted residual in the spatial domain.
[0186] Additionally, the transformation unit (106) may adaptively select a transformation type from among a plurality of transformation types and use a transformation basis function corresponding to the selected transformation type to transform the prediction residual into a transformation coefficient. Such a transformation may be called an explicit multiple core transform (EMT) or an adaptive multiple transform (AMT). Also, the transformation basis function may simply be called a basis.
[0187] Multiple transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Additionally, these transformation types may be denoted as DCT2, DCT5, DCT8, DST1, and DST7, respectively. FIG. 14 is a table showing transformation basis functions corresponding to each transformation type. In FIG. 14, N represents the number of input pixels. The selection of a transformation type among these multiple transformation types may depend, for example, on the type of prediction (intra prediction and inter prediction, etc.) or on the intra prediction mode.
[0188] Information indicating whether to apply such EMT or AMT (e.g., called an EMT flag or an AMT flag) and information indicating the selected transformation type are typically signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).
[0189] Additionally, the transformation unit (106) may re-transform the transformation coefficients (i.e., the transformation result). Such re-transformation may be called an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transformation unit (106) performs re-transformation for each sub-block (e.g., a 4×4 pixel sub-block) included in the block of transformation coefficients corresponding to the intra-predicted residual. Information indicating whether to apply NSST and information regarding the transformation matrix used for NSST are typically signaled at the CU level. Furthermore, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).
[0190] Separable transformation and non-separable transformation may be applied to the transformation unit (106). Separable transformation is a method of performing multiple transformations by separating each direction according to the number of dimensions of the input, and non-separable transformation is a method of grouping two or more dimensions into one dimension when the input is multi-dimensional and performing transformations by grouping them.
[0191] For example, as an example of a non-separable transformation, if the input is a block of 4×4 pixels, it can be considered as a single array with 16 elements, and the transformation process is performed on that array using a 16×16 transformation matrix.
[0192] In addition, in another example of a non-separable transformation, a 4×4 pixel input block may be considered as a single array with 16 elements, and then a transformation (Hypercube Givens Transform) may be performed by performing Givens rotations multiple times on the array.
[0193] In the transformation in the transformation unit (106), the transformation type of the transformation basis function that transforms into the frequency domain according to the region within the CU may be switched. As an example, there is the SVT (Spatially Varying Transform).
[0194] Figure 15 is a diagram showing an example of an SVT.
[0195] In SVT, as shown in FIG. 15, the CU is divided into two equal parts in the horizontal or vertical direction, and only one of the regions is converted to the frequency domain. The conversion type may be set for each region, for example, DST7 and DCT8 are used. For example, among the two regions obtained by dividing the CU vertically, DST7 and DCT8 may be used for the region at position 0. Or, among the two regions, DST7 is used for the region at position 1. Similarly, among the two regions obtained by dividing the CU horizontally, DST7 and DCT8 are used for the region at position 0. Or, among the two regions, DST7 is used for the region at position 1. In the example shown in FIG. 15, conversion is performed on only one of the two regions within the CU and not on the other, but conversion may be performed on each of the two regions. In addition, the partitioning method may include not only 2-parts but also 4-parts. Furthermore, it can be made more flexible, such as by encoding information indicating the partitioning method and signaling it in the same way as CU partitioning. Also, SVT is sometimes referred to as SBT (Sub-block Transform).
[0196] The aforementioned AMT and EMT may also be referred to as MTS (Multiple Transform Selection). When applying MTS, a transformation type such as DST7 or DCT8 may be selected, and information indicating the selected transformation type may be encoded as index information for each CU. On the other hand, there is a process called IMTS (Implicit MTS) which selects the transformation type used for orthogonal transformation based on the shape of the CU without encoding index information. When applying IMTS, for example, if the shape of the CU is rectangular, the short side of the rectangle is transformed orthogonally using DST7 and the long side using DCT2. Also, for example, if the shape of the CU is square, orthogonal transformation is performed using DCT2 if MTS is valid within the sequence, and using DST7 if MTS is invalid. DCT2 and DST7 are examples, and other transformation types may be used, or different combinations of transformation types may be used. IMTS may be available only in the intra-prediction block, or it may be available in both the intra-prediction block and the inter-prediction block.
[0197] In the above description, three processes—MTS, SBT, and IMTS—were described as selection processes for selectively switching the conversion type used for orthogonal conversion. All three selection processes may be enabled, or only some of the selection processes may be selectively enabled. Whether to enable individual selection processes can be identified by flag information within the header, such as SPS. For example, if all three selection processes are enabled, one of the three selection processes is selected at the CU level to perform orthogonal conversion. Additionally, regarding the selection processes for selectively switching the conversion type, if the following four functions [1] to [4] can realize at least one function, a selection process different from the above three selection processes may be used, or each of the above three selection processes may be replaced with a separate process. Function [1] is a function that orthogonally converts the entire range within the CU and encodes information indicating the conversion type used for conversion. Function [2] is a function that performs an orthogonal transformation of the entire range of CU and determines the transformation type based on a predetermined rule without encoding information indicating the transformation type. Function [3] is a function that performs an orthogonal transformation of a part of the CU and encodes information indicating the transformation type used for the transformation. Function [4] is a function that performs an orthogonal transformation of a part of the CU and determines the transformation type based on a predetermined rule without encoding information indicating the transformation type used for the transformation.
[0198] In addition, the application of MTS, IMTS, and SBT may be determined for each processing unit. For example, the application may be determined for a sequence unit, picture unit, brick unit, slice unit, CTU unit, or CU unit.
[0199] Furthermore, the tool for selectively switching transformation types in the present disclosure may be described as a method for adaptively selecting a basis used in transformation processing, a selection process, or a process for selecting a basis. Additionally, the tool for selectively switching transformation types may be described as a mode for adaptively selecting transformation types.
[0200] FIG. 16 is a flowchart showing an example of processing by the conversion unit (106).
[0201] For example, the transformation unit (106) determines whether to perform an orthogonal transformation (step St_1). Here, if the transformation unit (106) determines that it will perform an orthogonal transformation (Yes in step St_1), it selects a transformation type to be used for the orthogonal transformation from a plurality of transformation types (step St_2). Next, the transformation unit (106) performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Then, the transformation unit (106) encodes the information by outputting information indicating the selected transformation type to the entropy encoding unit (110) (step St_4). On the other hand, if the transformation unit (106) determines that it will not perform an orthogonal transformation (No in step St_1), it encodes the information by outputting information indicating that it will not perform an orthogonal transformation to the entropy encoding unit (110) (step St_5). In addition, the determination of whether to perform an orthogonal transformation in step St_1 may be made based, for example, on the size of the transformation block, the prediction mode applied to the CU, etc. Also, information indicating the transformation type used for the orthogonal transformation may not be encoded, and the orthogonal transformation may be performed using a predefined transformation type.
[0202] FIG. 17 is a flowchart showing another example of processing by the conversion unit (106). Also, the example shown in FIG. 17 is an example of an orthogonal conversion in which a method of selectively switching the conversion type used for orthogonal conversion is applied, similar to the example shown in FIG. 16.
[0203] As an example, the first conversion type group may include DCT2, DST7, and DCT8. As another example, the second conversion type group may include DCT2. Furthermore, the conversion types included in the first conversion type group and the second conversion type group may partially overlap, or all may be different conversion types.
[0204] Specifically, the transformation unit (106) determines whether the transformation size is less than or equal to a predetermined value (step Su_1). Here, if it is determined that it is less than or equal to the predetermined value (Yes in step Su_1), the transformation unit (106) performs an orthogonal transformation of the predicted residual of the current block using a transformation type included in the first transformation type group (step Su_2). Next, the transformation unit (106) encodes the information by outputting information indicating which transformation type to use among one or more transformation types included in the first transformation type group to the entropy encoding unit (110) (step Su_3). Meanwhile, if the transformation unit (106) determines that the transformation size is not less than or equal to the predetermined value (No in step Su_1), it performs an orthogonal transformation of the predicted residual of the current block using the second transformation type group (step Su_4).
[0205] In step Su_3, the information indicating the transformation type used for orthogonal transformation may be information indicating a combination of the transformation type applied to the vertical direction of the current block and the transformation type applied to the horizontal direction. Additionally, the first transformation type group may include only one transformation type, and the information indicating the transformation type used for orthogonal transformation may not be encoded. The second transformation type group may include multiple transformation types, and among one or more transformation types included in the second transformation type group, the information indicating the transformation type used for orthogonal transformation may be encoded.
[0206] In addition, the transformation type may be determined based solely on the transformation size. Furthermore, if the process determines the transformation type used for orthogonal transformation based on the transformation size, it is not limited to determining whether the transformation size is less than or equal to a predetermined value.
[0207] [Quantumization Department]
[0208] The quantization unit (108) quantizes the conversion coefficients output from the conversion unit (106). Specifically, the quantization unit (108) scans a plurality of conversion coefficients of the current block in a predetermined scanning order and quantizes the conversion coefficients based on a quantization parameter (QP) corresponding to the scanned conversion coefficients. Then, the quantization unit (108) outputs the quantized plurality of conversion coefficients of the current block (hereinafter referred to as quantization coefficients) to the entropy encoding unit (110) and the inverse quantization unit (112).
[0209] A predetermined scanning order is a sequence for the quantization / inverse quantization of the transformation coefficients. For example, a predetermined scanning order is defined as an ascending order of frequency (from low frequency to high frequency) or a descending order (from high frequency to low frequency).
[0210] A quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the error of the quantization coefficient (quantization error) increases.
[0211] In addition, quantization matrices are sometimes used for quantization. For example, several types of quantization matrices may be used corresponding to frequency conversion sizes such as 4×4 and 8×8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and chrominance. Furthermore, quantization refers to the process of digitizing values sampled at predetermined intervals by corresponding them to predetermined levels, and in this technical field, terms such as rounding, scaling, or scaling are sometimes used.
[0212] As a method of using a quantization matrix, there is a method of using a quantization matrix directly set on the side of the encoding device (100) and a method of using a default quantization matrix (default matrix). On the side of the encoding device (100), by directly setting the quantization matrix, a quantization matrix according to the characteristics of the image can be set. However, in this case, there is a disadvantage that the amount of encoding increases due to the encoding of the quantization matrix. In addition, instead of using the default quantization matrix or the encoded quantization matrix as is, a quantization matrix used for quantizing the current block may be generated based on the default quantization matrix or the encoded quantization matrix.
[0213] On the other hand, there is also a method that quantizes both high-frequency and low-frequency coefficients equally without using a quantization matrix. Furthermore, this method is equivalent to using a quantization matrix where all coefficients have the same value (a flat matrix).
[0214] The quantization matrix may be encoded, for example, at the sequence level, picture level, slice level, brick level, or CTU level.
[0215] When the quantization unit (108) uses a quantization matrix, for example, for each conversion coefficient, the quantization width, etc. obtained from the quantization parameter, etc. is scaled using the value of the quantization matrix. Quantization processing performed without using a quantization matrix may be a process of quantizing conversion coefficients based on the quantization width obtained from the quantization parameter, etc. In addition, in quantization processing performed without using a quantization matrix, the quantization width may be multiplied by a predetermined value that is common to all conversion coefficients within the block.
[0216] FIG. 18 is a block diagram showing an example of the functional configuration of the quantization unit (108).
[0217] The quantization unit (108) comprises, for example, a difference quantization parameter generation unit (108a), a prediction quantization parameter generation unit (108b), a quantization parameter generation unit (108c), a quantization parameter memory unit (108d), and a quantization processing unit (108e).
[0218] FIG. 19 is a flowchart showing an example of quantization by the quantization unit (108).
[0219] For example, the quantization unit (108) may perform quantization for each CU according to the flowchart shown in FIG. 19. Specifically, the quantization parameter generation unit (108c) determines whether to perform quantization (step Sv_1). Here, if it is determined that quantization is to be performed (Yes in step Sv_1), the quantization parameter generation unit (108c) generates quantization parameters for the current block (step Sv_2) and stores the quantization parameters in the quantization parameter memory unit (108d) (step Sv_3).
[0220] Next, the quantization processing unit (108e) quantizes the transformation coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). Then, the prediction quantization parameter generation unit (108b) obtains a quantization parameter of a processing unit different from the current block from the quantization parameter memory unit (108d) (step Sv_5). The prediction quantization parameter generation unit (108b) generates a prediction quantization parameter of the current block based on the obtained quantization parameter (step Sv_6). The difference quantization parameter generation unit (108a) calculates the difference between the quantization parameter of the current block generated by the quantization parameter generation unit (108c) and the prediction quantization parameter of the current block generated by the prediction quantization parameter generation unit (108b) (step Sv_7). By calculating this difference, a difference quantization parameter is generated. The difference quantization parameter generation unit (108a) encodes the difference quantization parameter by outputting the difference quantization parameter to the entropy encoding unit (110) (step Sv_8).
[0221] Additionally, the differential quantization parameter may be encoded at the sequence level, picture level, slice level, brick level, or CTU level. Also, the initial value of the quantization parameter may be encoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the quantization parameter may be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0222] Additionally, the quantization unit (108) may be equipped with a plurality of quantizers, and may apply dependent quantization that quantizes the conversion coefficients using a quantization method selected from a plurality of quantization methods.
[0223] [Entropy Encoding Section]
[0224] FIG. 20 is a block diagram showing an example of the functional configuration of the entropy encoding unit (110).
[0225] The entropy encoding unit (110) generates a stream by performing entropy encoding on the quantization coefficients input from the quantization unit (108) and the prediction parameters input from the prediction parameter generation unit (130). For example, CABAC (Context-based Adaptive Binary Arithmetic Coding) is used for the entropy encoding. Specifically, the entropy encoding unit (110) is equipped with, for example, a binarization unit (110a), a context control unit (110b), and a binary arithmetic encoding unit (110c). The binarization unit (110a) performs binarization to convert multi-valued signals, such as quantization coefficients and prediction parameters, into binary signals. Examples of binarization methods include Truncated Rice Binarization, Exponential Golomb codes, Fixed Length Binarization, etc. The context control unit (110b) derives a context value, that is, a probability of occurrence of a binary signal, based on the characteristics of the syntax element or surrounding conditions. Methods for deriving this context value include, for example, bypass, syntax element reference, upper / left adjacent block reference, hierarchical information reference, and others. The binary arithmetic encoding unit (110c) performs arithmetic encoding on the binary signal using the derived context value.
[0226] FIG. 21 is a diagram showing the flow of CABAC in the entropy encoding unit (110).
[0227] First, initialization is performed on the CABAC in the entropy encoding unit (110). In this initialization, initialization in the binary arithmetic encoding unit (110c) and the setting of the initial context value are performed. Then, the binary encoding unit (110a) and the binary arithmetic encoding unit (110c) perform binary encoding and arithmetic encoding in sequence for each of the multiple quantization coefficients of, for example, the CTU. At this time, the context control unit (110b) updates the context value whenever arithmetic encoding is performed. Then, the context control unit (110b) resets the context value as a post-processing step. This reset context value is used, for example, as the initial value of the context value for the next CTU.
[0228] [Inverse Quantumization Unit]
[0229] The inverse quantization unit (112) inversely quantizes the quantization coefficients input from the quantization unit (108). Specifically, the inverse quantization unit (112) inversely quantizes the quantization coefficients of the current block in a predetermined scanning order. Then, the inverse quantization unit (112) outputs the inversely quantized conversion coefficients of the current block to the inverse conversion unit (114).
[0230] [Inverse Transformation Section]
[0231] The inverse transform unit (114) restores the predicted residual by inversely transforming the transform coefficients input from the inverse quantization unit (112). Specifically, the inverse transform unit (114) restores the predicted residual of the current block by performing an inverse transform corresponding to the transformation by the transform unit (106) on the transform coefficients. Then, the inverse transform unit (114) outputs the restored predicted residual to the adder (116).
[0232] In addition, the restored prediction residual typically does not match the prediction error calculated by the subtraction unit (104) because information is lost due to quantization. That is, the restored prediction residual typically contains a quantization error.
[0233] [Additional]
[0234] The adder (116) reconstructs the current block by adding the prediction residual input from the inverse transform unit (114) and the prediction image input from the prediction control unit (128). As a result, a reconstructed image is generated. Then, the adder (116) outputs the reconstructed image to the block memory (118) and the loop filter unit (120).
[0235] [Block Memory]
[0236] The block memory (118) is, for example, a block referenced in intra prediction and is a memory unit for storing blocks within the current picture. Specifically, the block memory (118) stores the reconstructed image output from the adder (116).
[0237] [Frame Memory]
[0238] The frame memory (122) is a memory unit for storing, for example, a reference picture used for inter prediction, and is also called a frame buffer. Specifically, the frame memory (122) stores a reconstructed image filtered by the loop filter unit (120).
[0239] [Loop Filter Section]
[0240] The loop filter unit (120) performs loop filtering processing on the reconstructed image output from the adder unit (116) and outputs the filtered reconstructed image to the frame memory (122). A loop filter is a filter (in-loop filter) used within an encoding loop, and includes, for example, an adaptive loop filter (ALF), a deblocking filter (DF or DBF), and a sample adaptive offset (SAO).
[0241] FIG. 22 is a block diagram showing an example of the functional configuration of a loop filter section (120).
[0242] The loop filter unit (120) comprises, for example as shown in FIG. 22, a deblocking filter processing unit (120a), an SAO processing unit (120b), and an ALF processing unit (120c). The deblocking filter processing unit (120a) performs the deblocking filter processing described above on the reconstructed image. The SAO processing unit (120b) performs the SAO processing described above on the reconstructed image after the deblocking filter processing. Additionally, the ALF processing unit (120c) applies the ALF processing described above to the reconstructed image after the SAO processing. Details regarding the ALF and deblocking filters will be described later. SAO processing is a process that improves image quality by reducing ringing (a phenomenon in which pixel values are deformed to ripple around edges) and correcting pixel value misalignment. Examples of this SAO processing include edge offset processing and band offset processing. Additionally, the loop filter unit (120) may not have to have all the processing units disclosed in FIG. 22, and may have only some of the processing units. Also, the loop filter unit (120) may be configured to perform each of the above-described processing in a sequence different from the processing sequence disclosed in FIG. 22.
[0243] [Loop Filter Section > Adaptive Loop Filter]
[0244] In ALF, a least squares error filter is applied to remove encoding distortion, and for example, for each 2×2 pixel sub-block within the current block, one filter selected from a plurality of filters is applied based on the direction and activity of the local gradient.
[0245] Specifically, first, a subblock (e.g., a 2×2 pixel subblock) is classified into multiple classes (e.g., 15 or 25 classes). The classification of the subblock is performed, for example, based on the gradient direction and activity level. In a specific example, a classification value C (e.g., C = 5D + A) is calculated using a gradient direction value D (e.g., 0 to 2 or 0 to 4) and a gradient activity value A (e.g., 0 to 4). Then, based on the classification value C, the subblock is classified into multiple classes.
[0246] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Also, the gradient activation value A is derived, for example, by adding gradients in multiple directions and quantizing the result of the addition.
[0247] Based on these classification results, a filter for the sub-block is determined from among multiple filters.
[0248] For example, a circularly symmetric shape is used as the shape of the filter used in ALF. FIGS. 23a to 23c are drawings showing multiple examples of the shapes of filters used in ALF. FIG. 23a shows a 5×5 diamond-shaped filter, FIG. 23b shows a 7×7 diamond-shaped filter, and FIG. 23c shows a 9×9 diamond-shaped filter. Information indicating the shape of the filter is typically signaled at the picture level. In addition, the signaling of information indicating the shape of the filter does not need to be limited to the picture level and may be at other levels (e.g., sequence level, slice level, brick level, CTU level, or CU level).
[0249] The on / off status of ALF may be determined, for example, at the picture level or the CU level. For example, whether to apply ALF to luminance may be determined at the CU level, and whether to apply ALF to chrominance may be determined at the picture level. Information indicating the on / off status of ALF is typically signaled at the picture level or the CU level. Furthermore, the signaling of information indicating the on / off status of ALF is not limited to the picture level or the CU level, and may be at other levels (for example, the sequence level, slice level, brick level, or CTU level).
[0250] Also, as described above, one filter is selected from among a plurality of filters and ALF processing is performed on the subblock. For each of the plurality of filters (e.g., 15 or 25 filters), a set of coefficients consisting of a plurality of coefficients used in the filter is typically signaled at the picture level. Furthermore, the signaling of the set of coefficients is not limited to the picture level and may be at other levels (e.g., sequence level, slice level, brick level, CTU level, CU level, or subblock level).
[0251] [Loop Filter > Cross Component Adaptive Loop Filter]
[0252] FIG. 23d is a diagram showing an example in which a Y sample (first component) is used in CCALF of Cb and CCALF of Cr (multiple components different from the first component). FIG. 23e is a diagram showing a diamond-shaped filter.
[0253] One example of CC-ALF operates by applying a linear diamond filter (Fig. 23d, Fig. 23e) to the luminance channel of each chrominance component. For example, the filter coefficients are transmitted via APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and notified by a context encoding complete flag received at each block of samples. The block size and the CC-ALF enable flag are received at the slice level of each chrominance component. The syntax and semantics of CC-ALF are provided in the Appendix. In this document, block sizes of 16×16, 32×32, 64×64, and 128×128 (for chrominance samples) are supported.
[0254] [Loop Filter > Joint Chroma Cross Component Adaptive Loop Filter]
[0255] Figure 23f is a diagram showing an example of JC-CCALF. Figure 23g is a diagram showing an example of a weight_index candidate of JC-CCALF.
[0256] One example of JC-CCALF uses only one CCALF filter to generate a single CCALF filter output as a chromatic difference adjustment signal for only one chromatic component, and applies an appropriately weighted version of the same chromatic difference adjustment signal to other chromatic components. In this way, the complexity of the conventional CCALF is generally halved.
[0257] The weight is encoded in the sign flag and the weight index. The weight index (represented as weight_index) is encoded in 3 bits and specifies the size of the JC-CCALF weight JcCcWeight. It cannot be equal to 0. The size of JcCcWeight is determined as follows.
[0258] · If weight_index is 4 or less, JcCcWeight is equal to weight_index>>2.
[0259] · In other cases, JcCcWeight is equal to 4 / (weight_index-4).
[0260] The block-level on / off control of ALF filtering for Cb and Cr is as follows. This is the same as CCALF, where two separate sets of block-level on / off control flags are encoded. Here, unlike CCALF, only one block size variable is encoded because the on / off control block sizes for Cb and Cr are the same.
[0261] [Loop Filter Section > Deblocking Filter]
[0262] In the deblocking filter processing, the loop filter section (120) reduces deformation occurring at the block boundaries by performing filter processing on the block boundaries of the reconstructed image.
[0263] FIG. 24 is a block diagram showing an example of the detailed configuration of a deblocking filter processing unit (120a).
[0264] The deblocking filter processing unit (120a) comprises, for example, a boundary determination unit (1201), a filter determination unit (1203), a filter processing unit (1205), a processing determination unit (1208), a filter characteristic determination unit (1207), and switches (1202, 1204 and 1206).
[0265] The boundary determination unit (1201) determines whether the pixel being processed by the deblocking filter (i.e., the target pixel) exists near the block boundary. Then, the boundary determination unit (1201) outputs the determination result to the switch (1202) and the processing determination unit (1208).
[0266] If the boundary determination unit (1201) determines that the target pixel exists near the block boundary, the switch (1202) outputs the image before filter processing to the switch (1204). Conversely, if the boundary determination unit (1201) determines that the target pixel does not exist near the block boundary, the switch (1202) outputs the image before filter processing to the switch (1206). Additionally, the image before filter processing is an image consisting of a target pixel and at least one surrounding pixel located near the target pixel.
[0267] The filter determination unit (1203) determines whether to perform deblocking filter processing on the target pixel based on the pixel values of at least one surrounding pixel surrounding the target pixel. Then, the filter determination unit (1203) outputs the determination result to the switch (1204) and the processing determination unit (1208).
[0268] If the filter determination unit (1203) determines that the switch (1204) performs deblocking filter processing on the target pixel, the switch (1204) outputs the image before filter processing obtained through the switch (1202) to the filter processing unit (1205). Conversely, if the filter determination unit (1203) determines that the switch (1204) does not perform deblocking filter processing on the target pixel, the switch (1204) outputs the image before filter processing obtained through the switch (1202) to the switch (1206).
[0269] When the filter processing unit (1205) acquires an image prior to filter processing through switches (1202 and 1204), it performs deblocking filter processing on the target pixel having filter characteristics determined by the filter characteristic determining unit (1207). Then, the filter processing unit (1205) outputs the pixel after filter processing to the switch (1206).
[0270] The switch (1206) selectively outputs pixels that are not deblocked by the filter processing unit (1208) and pixels that are deblocked by the filter processing unit (1205) according to the control of the processing judgment unit (1208).
[0271] The processing decision unit (1208) controls the switch (1206) based on the respective determination results of the boundary determination unit (1201) and the filter determination unit (1203). That is, the processing decision unit (1208) outputs the deblocking filter-processed pixel from the switch (1206) when the boundary determination unit (1201) determines that the target pixel exists near the block boundary and the filter determination unit (1203) determines that deblocking filter processing is performed on the target pixel. In addition, except for the cases described above, the processing decision unit (1208) outputs the pixel that has not been deblocking filter-processed from the switch (1206). By repeating this output of pixels, an image after filter processing is output from the switch (1206). Additionally, the configuration shown in FIG. 24 is an example of the configuration in the deblocking filter processing unit (120a), and the deblocking filter processing unit (120a) may have other configurations.
[0272] FIG. 25 is a diagram showing an example of a deblocking filter having filter characteristics that are symmetric with respect to block boundaries.
[0273] In the deblocking filter processing, for example, using pixel values and quantization parameters, one of two deblocking filters with different characteristics, namely a strong filter and a weak filter, is selected. In the strong filter, as shown in FIG. 25, when pixels p0 to p2 and pixels q0 to q2 exist with block boundaries in between, the pixel values of each of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the operation shown in the following equation.
[0274] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0275] q'1=(p0+q0+q1+q2+2) / 4
[0276] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0277] In addition, in the above-described equation, p0~p2 and q0~q2 are the pixel values of pixels p0~p2 and pixels q0~q2, respectively. Also, q3 is the pixel value of pixel q3 adjacent to the block boundary and opposite side of pixel q2. Also, in the right-hand side of each of the above-described equations, the coefficient multiplied by the pixel value of each pixel used for deblocking filter processing is the filter coefficient.
[0278] Additionally, in the deblocking filter processing, clipping may be performed so that the pixel value after the operation does not change beyond a threshold. In this clipping process, the pixel value after the operation according to the above-described equation is clipped to “pixel value before operation ± 2 × threshold” using a threshold determined from the quantization parameter. This prevents excessive smoothing.
[0279] FIG. 26 is a diagram illustrating an example of a block boundary where deblocking filter processing is performed. FIG. 27 is a diagram showing an example of a Bs value.
[0280] The block boundary where the deblocking filter processing is performed is, for example, the boundary of the CU, PU, or TU of the 8×8 pixel block shown in FIG. 26. The deblocking filter processing is performed in units of, for example, 4 rows or 4 columns. First, for blocks P and Q shown in FIG. 26, the Bs (Boundary Strength) value is determined as shown in FIG. 27.
[0281] Depending on the Bs value in FIG. 27, it may be determined whether to perform deblocking filter processing of different intensities even for block boundaries belonging to the same image. Deblocking filter processing for the chrominance signal is performed when the Bs value is 2. Deblocking filter processing for the luminance signal is performed when the Bs value is 1 or greater and a predetermined condition is satisfied. In addition, the determination condition for the Bs value is not limited to that shown in FIG. 27 and may be determined based on other parameters.
[0282] [Prediction Unit (Intra Prediction Unit · Inter Prediction Unit · Prediction Control Unit)]
[0283] FIG. 28 is a flowchart illustrating an example of processing performed in the prediction unit of an encoding device (100). In addition, as an example, the prediction unit is composed of all or part of the components of an intra prediction unit (124), an inter prediction unit (126), and a prediction control unit (128). The prediction processing unit includes, for example, an intra prediction unit (124) and an inter prediction unit (126).
[0284] The prediction unit generates a prediction image of the current block (step Sb_1). Additionally, the prediction image may include, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). Specifically, the prediction unit generates a prediction image of the current block using a reconstruction image already obtained by performing the generation of a prediction image for another block, the generation of a prediction residual, the generation of quantization coefficients, the restoration of the prediction residual, and the addition of the prediction images.
[0285] The reconstructed image may, for example, be an image of a reference picture, or an image of a fully encoded block (i.e., another block described above) within a current picture that includes the current block. The fully encoded block within the current picture is, for example, an adjacent block of the current block.
[0286] FIG. 29 is a flowchart showing another example of processing performed in the prediction section of the encoding device (100).
[0287] The prediction unit generates a prediction image using a first method (step Sc_1a), generates a prediction image using a second method (step Sc_1b), and generates a prediction image using a third method (step Sc_1c). The first method, the second method, and the third method are different methods for generating a prediction image, and may each be, for example, an inter-prediction method, an intra-prediction method, and a prediction method other than those. The above-described reconstructed image may be used as these prediction methods.
[0288] Next, the prediction unit evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, respectively (step Sc_2). For example, the prediction unit calculates a cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, respectively, and evaluates the predicted images by comparing the cost C of the predicted images. Additionally, the cost C is calculated by an equation of the RD optimization model, for example, C = D + λ × R. In this equation, D is the encoding variant of the predicted image and is represented, for example, by the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. Also, R is the bit rate of the stream. Also, λ is, for example, an undetermined Lagrange multiplier.
[0289] Next, the prediction unit selects one of the prediction images generated in each of steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). That is, the prediction unit selects a method or mode to obtain a final prediction image. For example, the prediction unit selects the prediction image with the smallest cost C based on the cost C calculated for the prediction images. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 may be performed based on parameters used in the processing of encoding. The encoding device (100) may signal information to specify the selected prediction image, method, or mode into a stream. The information may be, for example, a flag. By doing so, the decoder (200) can generate a prediction image according to the method or mode selected by the encoding device (100) based on the information. In addition, in the example shown in FIG. 29, the prediction unit selects one of the prediction images after generating a prediction image in each method. However, before generating the prediction images, the prediction unit may select a method or mode based on the parameters used in the processing of the encoding described above, and generate a prediction image according to that method or mode.
[0290] For example, the first method and the second method are intra prediction and inter prediction, respectively, and the prediction unit may select a final prediction image for the current block from the prediction images generated according to these prediction methods.
[0291] FIG. 30 is a flowchart showing another example of processing performed in the prediction section of the encoding device (100).
[0292] First, the prediction unit generates a prediction image by intra prediction (step Sd_1a) and generates a prediction image by inter prediction (step Sd_1b). Additionally, the prediction image generated by intra prediction is also referred to as the intra prediction image, and the prediction image generated by inter prediction is also referred to as the inter prediction image.
[0293] Next, the prediction unit evaluates the intra prediction image and the inter prediction image, respectively (step Sd_2). The cost C described above may be used for this evaluation. Then, the prediction unit may select the prediction image from the intra prediction image and the inter prediction image for which the smallest cost C is calculated as the final prediction image of the current block (step Sd_3). That is, a prediction method or mode for generating the prediction image of the current block is selected.
[0294] [Intra Prediction Department]
[0295] The intra prediction unit (124) generates a predicted image of a current block (i.e., an intra prediction image) by performing intra prediction (also called in-frame prediction) of the current block by referencing a block within the current picture stored in the block memory (118). Specifically, the intra prediction unit (124) generates an intra prediction image by performing intra prediction by referencing pixel values (e.g., luminance values, chrominance values) of a block adjacent to the current block, and outputs the intra prediction image to the prediction control unit (128).
[0296] For example, the intra prediction unit (124) performs intra prediction using one of a plurality of predetermined intra prediction modes. The plurality of intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes.
[0297] One or more non-directional prediction modes include, for example, the Planar prediction mode and the DC prediction mode defined in the H.265 / HEVC standard.
[0298] Multiple directional prediction modes include, for example, the 33 directional prediction modes defined by the H.265 / HEVC standard. Additionally, multiple directional prediction modes may include 32 additional directional prediction modes in addition to the 33 directions (a total of 65 directional prediction modes). FIG. 31 is a diagram showing a total of 67 intra prediction modes (2 non-directional prediction modes and 65 directional prediction modes) in intra prediction. Solid arrows indicate the 33 directions defined by the H.265 / HEVC standard, and dashed arrows indicate the additional 32 directions (the 2 non-directional prediction modes are not shown in FIG. 31).
[0299] In various implementation examples, a luminance block may be referenced in the intra prediction of a chrominance block. That is, the chrominance component of the current block may be predicted based on the luminance component of the current block. Such intra prediction is sometimes called CCLM (cross-component linear model) prediction. An intra prediction mode of a chrominance block that references such a luminance block (for example, called a CCLM mode) may be added as one of the intra prediction modes of a chrominance block.
[0300] The intra prediction unit (124) may correct the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction. Intra prediction involving such correction may be called PDPC (position dependent intra prediction combination). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is typically signaled at the CU level. In addition, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).
[0301] FIG. 32 is a flowchart showing an example of processing by the intra prediction unit (124).
[0302] The intra prediction unit (124) selects one intra prediction mode from a plurality of intra prediction modes (step Sw_1). Then, the intra prediction unit (124) generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra prediction unit (124) determines the Most Probable Modes (MPM) (step Sw_3). The MPM consists of, for example, six intra prediction modes. Two of the six intra prediction modes may be a Planar prediction mode and a DC prediction mode, and the remaining four modes may be directional prediction modes. Then, the intra prediction unit (124) determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0303] Here, if it is determined that the selected intra prediction mode is included in the MPM (Yes in step Sw_4), the intra prediction unit (124) sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode among the MPMs (step Sw_6). In addition, the MPM flag set to 1 and the information indicating the intra prediction mode are each encoded by the entropy encoding unit (110) as prediction parameters.
[0304] Meanwhile, if it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra prediction unit (124) sets the MPM flag to 0 (step Sw_7). Alternatively, the intra prediction unit (124) does not set the MPM flag. Then, the intra prediction unit (124) generates information indicating the selected intra prediction mode among one or more intra prediction modes not included in the MPM (step Sw_8). Additionally, the MPM flag set to 0 and the information indicating the intra prediction mode are each encoded by the entropy encoding unit (110) as prediction parameters. The information indicating the intra prediction mode represents, for example, any one of 0 to 60.
[0305] [Inter Prediction Department]
[0306] The inter prediction unit (126) generates a predicted image (inter prediction image) by performing inter prediction of the current block (also called inter-frame prediction) by referencing a reference picture stored in the frame memory (122) that is different from the current picture. Inter prediction is performed on the unit of the current block or the current sub-block within the current block. The sub-block is included in the block and is a unit smaller than the block. The size of the sub-block may be 4×4 pixels, 8×8 pixels, or any other size. The size of the sub-block may be converted into units such as slices, bricks, or pictures.
[0307] For example, the inter prediction unit (126) performs motion estimation within a reference picture for a current block or a current sub-block and finds the reference block or sub-block that most matches the current block or current sub-block. Then, the inter prediction unit (126) obtains motion information (e.g., motion vector) that compensates for movement or change from the reference block or sub-block to the current block or sub-block. Based on the motion information, the inter prediction unit (126) performs motion compensation (or motion prediction) and generates an inter prediction image of the current block or sub-block. The inter prediction unit (126) outputs the generated inter prediction image to the prediction control unit (128).
[0308] The motion information used for motion compensation may be signaled in various forms as inter-prediction images. For example, motion vectors may be signaled. As another example, the difference between a motion vector and a motion vector predictor may be signaled.
[0309] [Reference Picture List]
[0310] FIG. 33 is a diagram showing an example of each reference picture, and FIG. 34 is a conceptual diagram showing an example of a reference picture list. A reference picture list is a list representing one or more reference pictures stored in a frame memory (122). In FIG. 33, a rectangle represents a picture, an arrow represents the reference relationship of the picture, the horizontal axis represents time, I, P, and B inside the rectangle represent an intra-predicted picture, a single-predicted picture, and a pair-predicted picture, respectively, and a number inside the rectangle represents the decoding order. As shown in FIG. 33, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As shown in FIG. 34, a reference picture list is a list representing candidates for reference pictures, and, for example, one picture (or slice) may have one or more reference picture lists. For example, if the current picture is a single-predicted picture, one reference picture list is used, and if the current picture is a pairwise-predicted picture, two reference picture lists are used. In the example of FIG. 33 and FIG. 34, picture B3, which is the current picture currPic, has two reference picture lists, L0 list and L1 list. When the current picture currPic is picture B3, the candidates for the reference picture of that current picture currPic are I0, P1, and B2, and each reference picture list (i.e., L0 list and L1 list) represents these pictures. The inter prediction unit (126) or prediction control unit (128) specifies whether to actually reference any picture among each reference picture list by the reference picture index refidxLx. In FIG. 34, reference pictures P1 and B2 are specified by the reference picture indices refIdxL0 and refIdxL1.
[0311] Such reference picture lists may be generated in sequence units, picture units, slice units, brick units, CTU units, or CU units. Additionally, among the reference pictures shown in the reference picture list, the reference picture index representing the reference picture referenced in inter prediction may be encoded at the sequence level, picture level, slice level, brick level, CTU level, or CU level. Additionally, a common reference picture list may be used for multiple inter prediction modes.
[0312] [Basic Flow of Inter-Prediction]
[0313] Figure 35 is a flowchart showing the basic flow of inter-prediction.
[0314] The inter-prediction unit (126) first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit (104) generates the difference between the current block and the prediction image as a prediction residual (step Se_4).
[0315] Here, the inter prediction unit (126) generates the predicted image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3) in the generation of the predicted image. Also, the inter prediction unit (126) determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2) in the determination of the MV. The selection of the candidate MV is performed, for example, by the inter prediction unit (126) generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Additionally, a previously derived MV may be added to the candidate MV list as a candidate MV. In addition, in deriving the MV, the inter prediction unit (126) may determine the selected at least one candidate MV as the MV of the current block by selecting at least one additional candidate MV from at least one candidate MV. Alternatively, the inter prediction unit (126) may determine the MV of the current block by searching for the area of the reference picture indicated by the candidate MV for each of the selected at least one candidate MV. In addition, searching for the area of the reference picture may be referred to as motion estimation.
[0316] In addition, in the example described above, steps Se_1 to Se_3 are performed by the inter prediction unit (126), and, for example, processing of step Se_1 or step Se_2, etc., may be performed by other components included in the encoding device (100).
[0317] In addition, a candidate MV list may be created for each processing in each inter prediction mode, or a common candidate MV list may be used for multiple inter prediction modes. Also, the processing of steps Se_3 and Se_4 corresponds to the processing of steps Sa_3 and Sa_4 shown in FIG. 9, respectively. Also, the processing of step Se_3 corresponds to the processing of step Sd_1b in FIG. 30.
[0318] [Flow for Deriving MV]
[0319] Figure 36 is a flowchart showing an example of MV derivation.
[0320] The inter prediction unit (126) may derive the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information may be encoded and signaled as a prediction parameter. That is, the encoded motion information is included in the stream.
[0321] Alternatively, the inter prediction unit (126) may derive the MV in a mode that does not encode motion information. In this case, motion information is not included in the stream.
[0322] Here, the modes for deriving MV include the normal inter mode, normal merge mode, FRUC mode, and affine mode described later. Among these modes, the modes that encode motion information include the normal inter mode, normal merge mode, and affine mode (specifically, affine inter mode and affine merge mode). In addition, the motion information may include not only MV but also the predicted MV selection information described later. Also, the modes that do not encode motion information include the FRUC mode. The inter prediction unit (126) selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.
[0323] Figure 37 is a flowchart showing another example of MV derivation.
[0324] The inter-prediction unit (126) may derive the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded and signaled as a prediction parameter. That is, the encoded differential MV is included in the stream. This differential MV is the difference between the MV of the current block and the predicted MV. Also, the predicted MV is a predicted motion vector.
[0325] Alternatively, the inter prediction unit (126) may derive the MV in a mode that does not encode the differential MV. In this case, the encoded differential MV is not included in the stream.
[0326] Here, as described above, the modes for deriving MV include the normal inter, normal merge mode, FRUC mode, and affine mode described later. Among these modes, the modes for encoding differential MV include the normal inter mode and affine mode (specifically, affine inter mode). Also, the modes for not encoding differential MV include the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode). The inter prediction unit (126) selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.
[0327] [Mode of MV Derivation]
[0328] FIGS. 38a and FIG. 38b are diagrams illustrating an example of the classification of each mode of MV derivation. For example, as shown in FIG. 38a, the modes of MV derivation are broadly classified into three modes depending on whether motion information is encoded and whether differential MV is encoded. The three modes are the inter mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter mode is a mode that performs motion search and is a mode that encodes motion information and differential MV. For example, as shown in FIG. 38b, the inter mode includes the affine inter mode and the normal inter mode. The merge mode is a mode that does not perform motion search and is a mode that selects an MV from a surrounding encoded complete block and uses that MV to derive the MV of the current block. This merge mode is, basically, a mode that encodes motion information and does not encode differential MV. For example, as shown in FIG. 38b, the merge mode includes a normal merge mode (sometimes called a normal merge mode or regular merge mode), an MMVD (Merge with Motion Vector Difference) mode, a CIIP (Combined inter merge / intra prediction) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, among the modes included in the merge mode, in the MMVD mode, the difference MV is exceptionally encoded. In addition, the aforementioned affine merge mode and affine inter mode are modes included in the affine mode. The affine mode is a mode that derives the MV of each of a plurality of sub-blocks constituting the current block as the MV of the current block by assuming an affine transformation. The FRUC mode is a mode that derives the MV of the current block by performing a search between encoded complete regions, and is a mode that does not encode either motion information or difference MV. In addition, details of each of these modes will be described later.
[0329] In addition, the classification of each mode shown in FIGS. 38a and FIGS. 38b is an example and is not limited thereto. For example, when differential MV is encoded in CIIP mode, the CIIP mode is classified as an inter mode.
[0330] [MV Derivation > Normal Inter Mode]
[0331] Normal inter mode is an inter prediction mode that derives the MV of the current block by finding a block similar to the image of the current block from the region of the reference picture represented by the candidate MV. Also, in this normal inter mode, the differential MV is encoded.
[0332] Figure 39 is a flowchart showing an example of inter prediction by normal inter mode.
[0333] The inter prediction unit (126) first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of encoded completed blocks located around the current block in time or space (step Sg_1). That is, the inter prediction unit (126) creates a list of candidate MVs.
[0334] Next, the inter prediction unit (126) selects each of the N candidate MVs (N is an integer greater than or equal to 2) among the multiple candidate MVs obtained in step Sg_1 as predicted MV candidates and extracts them according to a predetermined priority (step Sg_2). In addition, the priority is predetermined for each of the N candidate MVs.
[0335] Next, the inter prediction unit (126) selects one prediction MV candidate from among the N prediction MV candidates as the prediction MV of the current block (step Sg_3). At this time, the inter prediction unit (126) encodes prediction MV selection information into a stream to identify the selected prediction MV. That is, the inter prediction unit (126) outputs the prediction MV selection information as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0336] Next, the inter prediction unit (126) derives the MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter prediction unit (126) also encodes the difference value between the derived MV and the predicted MV into the stream as the difference MV. That is, the inter prediction unit (126) outputs the difference MV as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130). Additionally, the encoded reference picture is a picture composed of multiple blocks reconstructed after encoding.
[0337] Finally, the inter prediction unit (126) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processing of steps Sg_1 to Sg_5 is executed for each block. For example, if the processing of steps Sg_1 to Sg_5 is executed for each of all blocks included in a slice, the inter prediction using normal inter mode for that slice is terminated. Also, if the processing of steps Sg_1 to Sg_5 is executed for each of all blocks included in a picture, the inter prediction using normal inter mode for that picture is terminated. Additionally, if the processing of steps Sg_1 to Sg_5 is not executed for all blocks included in a slice but is executed for some blocks, the inter prediction using normal inter mode for that slice may be terminated. If the processing of steps Sg_1 through Sg_5 is executed for some blocks included in the picture, the inter prediction using normal inter mode for that picture may be terminated.
[0338] Furthermore, the predicted image is the aforementioned inter prediction signal. Also, information indicating the inter prediction mode (normal inter mode in the example above) used to generate the predicted image, which is included in the encoded signal, is encoded, for example, as a prediction parameter.
[0339] Additionally, the candidate MV list may be shared with the list used in other modes. Also, processing regarding the candidate MV list may be applied to processing regarding the list used in other modes. This processing regarding the candidate MV list is, for example, extraction or selection of candidate MVs from the candidate MV list, rearrangement of candidate MVs, or deletion of candidate MVs.
[0340] [MV Derivation > Normal Merge Mode]
[0341] Normal merge mode is an inter-prediction mode that derives an MV by selecting a candidate MV from a list of candidate MVs as the MV of the current block. Additionally, normal merge mode is a merge mode in the narrow sense and is sometimes simply called merge mode. In the present embodiment, normal merge mode and merge mode are distinguished, and merge mode is used in a broad sense.
[0342] Figure 40 is a flowchart showing an example of inter-prediction by normal merge mode.
[0343] The inter prediction unit (126) first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of encoded completed blocks located around the current block in time or space (step Sh_1). That is, the inter prediction unit (126) creates a list of candidate MVs.
[0344] Next, the inter prediction unit (126) derives the MV of the current block by selecting one candidate MV from among the multiple candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter prediction unit (126) encodes MV selection information to identify the selected candidate MV into a stream. That is, the inter prediction unit (126) outputs the MV selection information as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0345] Finally, the inter prediction unit (126) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). The processing of steps Sh_1 to Sh_3 is executed for each block, for example. For example, when the processing of steps Sh_1 to Sh_3 is executed for each of all blocks included in the slice, the inter prediction using normal merge mode for that slice is terminated. Also, when the processing of steps Sh_1 to Sh_3 is executed for each of all blocks included in the picture, the inter prediction using normal merge mode for that picture is terminated. Additionally, the processing of steps Sh_1 to Sh_3 is not executed for all blocks included in the slice, but is executed for some blocks, and the inter prediction using normal merge mode for that slice is terminated. If the processing of steps Sh_1 to Sh_3 is executed for some blocks included in the picture, the inter prediction using normal merge mode for that picture may be terminated.
[0346] In addition, information indicating the inter-prediction mode (normal merge mode in the example above) used to generate the prediction image included in the stream is encoded, for example, as a prediction parameter.
[0347] FIG. 41 is a diagram illustrating an example of the MV derivation process of a current picture by normal merge mode.
[0348] First, the inter prediction unit (126) generates a list of candidate MVs in which candidate MVs are registered. Candidate MVs include spatially adjacent candidate MVs, which are MVs of multiple encoded complete blocks located spatially around the current block; temporally adjacent candidate MVs, which are MVs of blocks in the vicinity of the current block projected in the encoded complete reference picture; combined candidate MVs, which are MVs generated by combining the MV values of spatially adjacent candidate MVs and temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero.
[0349] Next, the inter prediction unit (126) determines the 1 candidate MV as the MV of the current block by selecting one candidate MV from among the multiple candidate MVs registered in the candidate MV list.
[0350] Additionally, the entropy encoding unit (110) encodes the merge_idx, a signal indicating which candidate MV was selected, by describing it in the stream.
[0351] In addition, the candidate MV registered in the candidate MV list described in FIG. 41 is an example and may be a number different from the number in the drawing, a configuration that does not include some types of candidate MV in the drawing, or a configuration that adds candidate MVs other than the types of candidate MV in the drawing.
[0352] The final MV may be determined by performing the dynamic motion vector refreshing (DMVR) described later using the MV of the current block derived by the normal merge mode. Additionally, in the normal merge mode, the differential MV is not encoded, but in the MMVD mode, the differential MV is encoded. The MMVD mode selects one candidate MV from the candidate MV list, similar to the normal merge mode, but encodes the differential MV. Such an MMVD may be classified as a merge mode along with the normal merge mode, as shown in FIG. 38b. Furthermore, the differential MV in the MMVD mode does not have to be the same as the differential MV used in the inter mode, and for example, the derivation of the differential MV in the MMVD mode may be a process with a smaller throughput compared to the derivation of the differential MV in the inter mode.
[0353] Additionally, a CIIP (Combined inter merge / intra prediction) mode may be performed to generate a prediction image of the current block by superimposing the prediction image generated by inter prediction and the prediction image generated by intra prediction.
[0354] Also, the list of candidate MVs may be referred to as the candidate list. Also, merge_idx is the MV selection information.
[0355] [MV Derivation > HMVP Mode]
[0356] FIG. 42 is a diagram illustrating an example of the MV derivation process of a current picture by HMVP mode.
[0357] In normal merge mode, the MV of the current block, e.g., CU, is determined by selecting one candidate MV from a list of candidate MVs generated by referencing a completed encoding block (e.g., CU). Here, other candidate MVs may be registered in the list of candidate MVs. The mode in which other candidate MVs are registered is called HMVP mode.
[0358] In HMVP mode, candidate MVs are managed using a FIFO (First-In First-Out) buffer for HMVP, separate from the candidate MV list of normal merge mode.
[0359] In the FIFO buffer, movement information such as the MVs of blocks processed in the past is stored in order from newest to oldest. In the management of this FIFO buffer, whenever a block is processed, the MV of the newest block (i.e., the CU processed immediately before) is stored in the FIFO buffer, and instead, the MV of the oldest CU (i.e., the CU processed first) in the FIFO buffer is removed from the FIFO buffer. In the example shown in FIG. 42, HMVP1 is the MV of the newest block, and HMVP5 is the MV of the oldest block.
[0360] And, for example, the inter prediction unit (126) checks, for each MV managed in the FIFO buffer, whether the MV is different from all candidate MVs already registered in the candidate MV list of normal merge mode in order from HMVP1. And, if the inter prediction unit (126) determines that it is different from all candidate MVs, it may add the MV managed in the FIFO buffer as a candidate MV to the candidate MV list of normal merge mode. At this time, the candidate MV registered from the FIFO buffer may be one or multiple.
[0361] By using the HMVP mode in this way, it becomes possible to add not only the MVs of blocks spatially or temporally adjacent to the current block, but also the MVs of blocks processed in the past to the candidate list. As a result, the variation of candidate MVs in the normal merge mode is expanded, thereby increasing the likelihood of improving coding efficiency.
[0362] In addition, the aforementioned MV may be motion information. That is, the information stored in the candidate MV list and the FIFO buffer may include not only the value of the MV, but also information indicating the referenced picture, the referenced direction, and the number of frames. Also, the aforementioned block is, for example, a CU.
[0363] In addition, the candidate MV list and FIFO buffer of FIG. 42 are examples, and the candidate MV list and FIFO buffer may be a list or buffer of a different size than FIG. 42, or a configuration in which candidate MVs are registered in a different order than FIG. 42. Also, the processing described herein is common to both the encoding device (100) and the decoding device (200).
[0364] In addition, HMVP mode can be applied to modes other than normal merge mode. For example, movement information such as MVs of blocks previously processed in affine mode can be stored in a FIFO buffer in order from newest to newest and used as candidate MVs. A mode in which HMVP mode is applied to an affine mode may be called history affine mode.
[0365] [MV Derivation > FRUC Mode]
[0366] Motion information may be derived from the decoding device (200) side without being signaled from the encoding device (100) side. For example, motion information may be derived by performing motion search on the decoding device (200) side. In this case, motion search is performed on the decoding device (200) side without using the pixel values of the current block. Modes for performing motion search on the decoding device (200) side include FRUC (frame rate up-conversion) mode or PMMVD (pattern matched motion vector derivation) mode.
[0367] An example of FRUC processing is shown in FIG. 43. First, the MVs of each encoded complete block spatially or temporally adjacent to the current block are referenced, and a list is generated in which those MVs are represented as candidate MVs (i.e., a candidate MV list, which may be common to the candidate MV list of normal merge mode) (Step Si_1). Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (Step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and one candidate MV is selected as the best candidate MV based on that evaluation value. Then, based on the selected best candidate MV, an MV for the current block is derived (Step Si_4). Specifically, for example, the selected best candidate MV is derived as is as the MV for the current block. For example, an MV for the current block may be derived by performing pattern matching on the surrounding area of a location within the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and evaluation values in the reference picture is performed on the surrounding area of the best candidate MV, and if there is an MV with a good evaluation value, the best candidate MV may be updated to that MV and used as the final MV for the current block. It is not necessary to update the MV with a better evaluation value.
[0368] Finally, the inter prediction unit (126) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). The processing of steps Si_1 to Si_5 is executed for each block, for example. For example, when the processing of steps Si_1 to Si_5 is executed for each of all blocks included in the slice, the inter prediction using FRUC mode for that slice is terminated. Also, when the processing of steps Si_1 to Si_5 is executed for each of all blocks included in the picture, the inter prediction using FRUC mode for that picture is terminated. Additionally, when the processing of steps Si_1 to Si_5 is not executed for all blocks included in the slice but is executed for some blocks, the inter prediction using FRUC mode for that slice may be terminated. When the processing of steps Si_1 to Si_5 is executed for some blocks included in the picture, the inter prediction using FRUC mode for that picture may be terminated.
[0369] Processing may also be performed at the sub-block level in the same way as the block level described above.
[0370] The evaluation value may be calculated by various methods. For example, a reconstructed image of an area within a reference picture corresponding to the MV may be compared with a reconstructed image of a predetermined area (such area may be, for example, an area of another reference picture or an area of an adjacent block of the current picture as shown below). Then, the difference between the pixel values of the two reconstructed images may be calculated and used as the evaluation value of the MV. Additionally, the evaluation value may be calculated by using other information in addition to the difference value.
[0371] Next, pattern matching will be explained in detail. First, one candidate MV included in the candidate MV list (also called the merge list) is selected as the starting point for the search by pattern matching. As the pattern matching, a first pattern matching or a second pattern matching may be used. The first pattern matching and the second pattern matching are sometimes called bilateral matching and template matching, respectively.
[0372] [MV Derivation > FRUC > Bilateral Matching]
[0373] In the first pattern matching, pattern matching is performed between two blocks within two different reference pictures that follow the motion trajectory of the current block. Accordingly, in the first pattern matching, an area within another reference picture that follows the motion trajectory of the current block is used as a predetermined area for calculating the evaluation value of the aforementioned candidate MV.
[0374] FIG. 44 is a diagram illustrating an example of a first pattern matching (bilateral matching) between two blocks in two reference pictures following a motion trajectory. As shown in FIG. 44, in the first pattern matching, two MVs (MV0, MV1) are derived by searching for the most matching pair among two block pairs in two different reference pictures (Ref0, Ref1) that follow the motion trajectory of a current block (Cur block). Specifically, for the current block, a difference is derived between a reconstructed image at a designated position in a first encoded reference picture (Ref0) designated as a candidate MV and a reconstructed image at a designated position in a second encoded reference picture (Ref1) designated as a symmetric MV scaled by the display time interval of the candidate MV, and an evaluation value is calculated using the obtained difference value. Among the multiple candidate MVs, the candidate MV that has the best evaluation value is selected as the best candidate MV.
[0375] Under the assumption of a continuous motion trajectory, the MV(MV0, MV1) referring to two reference blocks is proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is located temporally between the two reference pictures and the temporal distance from the current picture to the two reference pictures is the same, then in the first pattern matching, a bidirectional MV that is mirror-symmetric is derived.
[0376] [MV Derivation > FRUC > Template Matching]
[0377] In the second pattern matching (template matching), pattern matching is performed between a template in the current picture (a block adjacent to the current block in the current picture (e.g., an upper and / or left adjacent block)) and a block in the reference picture. Accordingly, in the second pattern matching, a block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the aforementioned candidate MV.
[0378] FIG. 45 is a diagram illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. As shown in FIG. 45, in the second pattern matching, the MV of the current block is derived by searching within the reference picture (Ref0) for the block that most matches the block adjacent to the current block (Cur block) within the current picture (Cur Pic). Specifically, for the current block, the difference between the reconstructed image of the encoded completed region of both the left adjacent and the upper adjacent, or either one thereof, and the reconstructed image at an equivalent position within the encoded completed reference picture (Ref0) designated as the candidate MV is derived, and an evaluation value is calculated using the obtained difference value. Among the multiple candidate MVs, the candidate MV that has the best evaluation value is selected as the best candidate MV.
[0379] Information indicating whether to apply such FRUC mode (e.g., called a FRUC flag) may be signaled at the CU level. Also, when FRUC mode is applied (e.g., when the FRUC flag is true), information indicating an applicable pattern matching method (a first pattern matching or a second pattern matching) may be signaled at the CU level. Additionally, the signaling of this information is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).
[0380] [MV Derivation > Affine Mode]
[0381] An affine mode is a mode that generates an MV using an affine transformation, and, for example, may derive an MV on a sub-block basis based on the MVs of multiple adjacent blocks. This mode is sometimes called an affine motion compensation prediction mode.
[0382] FIG. 46a is a diagram illustrating an example of deriving a sub-block unit MV based on the MVs of multiple adjacent blocks. In FIG. 46a, the current block includes, for example, a sub-block composed of 16 4×4 pixels. Here, the motion vector v0 of the upper-left corner control point of the current block is derived based on the MVs of the adjacent blocks, and likewise, the motion vector v1 of the upper-right corner control point of the current block is derived based on the MVs of the adjacent sub-blocks. Then, by projecting the two motion vectors v0 and v1 according to the following equation (1A), the motion vector (v) of each sub-block within the current block is obtained. x , v y ) is derived.
[0383]
[0384] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents a predetermined weighting factor.
[0385] Information indicating such an affine mode (e.g., called an affine flag) may be signaled at the CU level. In addition, the signaling of information indicating such an affine mode is not limited to the CU level and may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or sub-block level).
[0386] In addition, these affine modes may include several modes in which the derivation methods of the MV of the top-left and top-right corner control points are different. For example, there are two modes in the affine mode: the affine inter (also called affine normal inter) mode and the affine merge mode.
[0387] FIG. 46b is a diagram illustrating an example of deriving a sub-block unit MV in an affine mode using three control points. In FIG. 46b, the current block includes, for example, a sub-block consisting of 16 4×4 pixels. Here, the motion vector v0 of the upper-left corner control point of the current block is derived based on the MV of the adjacent block. Similarly, the motion vector v1 of the upper-right corner control point of the current block is derived based on the MV of the adjacent block, and the motion vector v2 of the lower-left corner control point of the current block is derived based on the MV of the adjacent block. Then, by projecting the three motion vectors v0, v1, and v2 according to the following equation (1B), the motion vector (v) of each sub-block within the current block is obtained. x , v y ) is derived.
[0388]
[0389] Here, x and y represent the horizontal and vertical positions of the sub-block center, respectively, and w and h represent predetermined weighting factors. w may represent the width of the current block, and h may represent the height of the current block.
[0390] Affine modes using different numbers of control points (e.g., 2 and 3) may be switched to the CU level and signaled. Additionally, information indicating the number of control points of the affine mode used at the CU level may be signaled at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or subblock level).
[0391] In addition, the affine mode having these three control points may include several modes in which the derivation methods of the MV of the upper-left, upper-right, and lower-left corner control points are different. For example, the affine mode having three control points has two modes, an affine inter mode and an affine merge mode, just like the affine mode having two control points described above.
[0392] In addition, in affine mode, the size of each sub-block included in the current block is not limited to 4×4 pixels and may be of a different size. For example, the size of each sub-block may be 8×8 pixels.
[0393] [MV Derivation > Affine Mode > Control Point]
[0394] FIGS. 47a, FIGS. 47b, and FIGS. 47c are conceptual diagrams for explaining an example of MV derivation of a control point in an affine mode.
[0395] In affine mode, as shown in FIG. 47a, for example, among the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, the predicted MV of each control point of the current block is calculated based on a plurality of MVs corresponding to the blocks encoded in affine mode. Specifically, these blocks are examined in the order of encoded block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left), and the first valid block encoded in affine mode is identified. Based on the plurality of MVs corresponding to this identified block, the MV of the control point of the current block is calculated.
[0396] For example, as shown in FIG. 47b, when block A adjacent to the left of the current block is encoded in an affine mode having two control points, motion vectors v3 and v4 are derived by projecting them onto the positions of the upper left corner and upper right corner of the encoded block containing block A. Then, from the derived motion vectors v3 and v4, the motion vector v0 of the upper left corner control point and the motion vector v1 of the upper right corner control point of the current block are derived.
[0397] For example, as shown in FIG. 47c, when block A adjacent to the left of the current block is encoded in an affine mode having three control points, motion vectors v3, v4, and v5 are derived by projecting them onto the positions of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. Then, from the derived motion vectors v3, v4, and v5, motion vector v0 of the upper left corner control point, motion vector v1 of the upper right corner control point, and motion vector v2 of the lower left corner control point are derived.
[0398] In addition, the method for deriving MV shown in FIGS. 47a to 47c may be used to derive the MV of each control point of the current block in step Sk_1 shown in FIG. 50, which will be described later, and may also be used to derive the predicted MV of each control point of the current block in step Sj_1 shown in FIG. 51, which will be described later.
[0399] FIGS. 48a and FIGS. 48b are conceptual diagrams for explaining another example of deriving a control point MV in an affine mode.
[0400] FIG. 48a is a diagram illustrating an affine mode having two control points.
[0401] In this affine mode, as shown in FIG. 48a, the MV selected from each of the coded complete blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the top-left corner control point of the current block. Likewise, the MV selected from each of the coded complete blocks D and E adjacent to the current block is used as the motion vector v1 of the top-right corner control point of the current block.
[0402] FIG. 48b is a diagram illustrating an affine mode having three control points.
[0403] In this affine mode, as shown in FIG. 48b, the MV selected from each of the coded complete blocks A, B, and C adjacent to the current block is used as the motion vector v0 of the top-left corner control point of the current block. Similarly, the MV selected from each of the coded complete blocks D and E adjacent to the current block is used as the motion vector v1 of the top-right corner control point of the current block. Additionally, the MV selected from each of the coded complete blocks F and G adjacent to the current block is used as the motion vector v2 of the bottom-left corner control point of the current block.
[0404] In addition, the method for deriving MV shown in FIG. 48a and FIG. 48b may be used to derive the MV of each control point of the current block in step Sk_1 shown in FIG. 50, which will be described later, and may also be used to derive the predicted MV of each control point of the current block in step Sj_1 of FIG. 51, which will be described later.
[0405] Here, for example, when switching affine modes of different numbers of control points (e.g., 2 and 3) to the CU level to signal, there may be cases where the number of control points in the coded complete block and the current block are different.
[0406] FIGS. 49a and FIGS. 49b are conceptual diagrams for explaining an example of a method for deriving the MV of a control point when the number of control points in the encoded complete block and the current block is different.
[0407] For example, as shown in FIG. 49a, the current block has three control points at the top-left corner, the top-right corner, and the bottom-left corner, and block A adjacent to the left of the current block is encoded in an affine mode having two control points. In this case, motion vectors v3 and v4 are derived by projecting them onto the top-left corner and top-right corner positions of the encoded block containing block A. Then, from the derived motion vectors v3 and v4, the motion vector v0 of the top-left corner control point and the motion vector v1 of the top-right corner control point of the current block are derived. Additionally, from the derived motion vectors v0 and v1, the motion vector v2 of the bottom-left corner control point is derived.
[0408] For example, as shown in FIG. 49b, the current block has two control points at the top-left corner and the top-right corner, and block A adjacent to the left of the current block is encoded in an affine mode having three control points. In this case, motion vectors v3, v4, and v5 are derived by projecting them onto the top-left corner, top-right corner, and bottom-left corner positions of the encoded block containing block A. Then, from the derived motion vectors v3, v4, and v5, the motion vector v0 of the top-left corner control point and the motion vector v1 of the top-right corner control point of the current block are derived.
[0409] In addition, the method for deriving MV shown in FIG. 49a and FIG. 49b may be used to derive the MV of each control point of the current block in step Sk_1 shown in FIG. 50, which will be described later, and may also be used to derive the predicted MV of each control point of the current block in step Sj_1 of FIG. 51, which will be described later.
[0410] [MV Derivation > Affine Mode > Affine Merge Mode]
[0411] FIG. 50 is a flowchart showing an example of an affine merge mode.
[0412] In the affine merge mode, first, the inter prediction unit (126) derives the MV for each control point of the current block (step Sk_1). The control points are the points at the top left corner and the top right corner of the current block as shown in FIG. 46a, or the points at the top left corner, the top right corner, and the bottom left corner of the current block as shown in FIG. 46b. At this time, the inter prediction unit (126) may encode MV selection information into a stream to identify the two or three derived MVs.
[0413] For example, when using the method of deriving MV shown in FIG. 47a to FIG. 47c, the inter prediction unit (126) examines these blocks in the order of encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), as shown in FIG. 47a, and identifies the first valid block encoded in affine mode.
[0414] The inter prediction unit (126) derives the MV of the control point using the first valid block encoded in a specified affine mode. For example, if block A is specified and block A has two control points, as shown in FIG. 47b, the inter prediction unit (126) calculates the motion vector v0 of the left-top corner control point and the motion vector v1 of the right-top corner control point of the current block from the motion vectors v3 and v4 of the left-top corner and right-top corner of the encoded block containing block A. For example, the inter prediction unit (126) calculates the motion vector v0 of the left-top corner control point and the motion vector v1 of the right-top corner control point of the current block by projecting the motion vectors v3 and v4 of the left-top corner and right-top corner of the encoded block onto the current block.
[0415] Alternatively, if block A is specified and block A has three control points, as shown in FIG. 47c, the inter prediction unit (126) calculates the motion vector v0 of the upper left corner control point of the current block, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point from the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. For example, the inter prediction unit (126) calculates the motion vector v0 of the upper left corner control point of the current block, the motion vector v1 of the upper right corner control point, and the motion vector v2 of the lower left corner control point by projecting the motion vectors v3, v4, and v5 of the upper left corner, upper right corner, and lower left corner of the encoded block onto the current block.
[0416] In addition, as shown in FIG. 49a above, if block A is specified and block A has two control points, the MV of three control points may be calculated, and as shown in FIG. 49b above, if block A is specified and block A has three control points, the MV of two control points may be calculated.
[0417] Next, the inter prediction unit (126) performs motion compensation for each of the plurality of sub-blocks included in the current block. That is, for each of the plurality of sub-blocks, the inter prediction unit (126) calculates the MV of the sub-block as an affine MV by using two motion vectors v0 and v1 and the above-described equation (1A), or by using three motion vectors v0, v1, and v2 and the above-described equation (1B) (step Sk_2). Then, the inter prediction unit (126) performs motion compensation for the sub-block using the affine MV and the encoded reference picture (step Sk_3). When the processing of steps Sk_2 and Sk_3 is executed for each of the sub-blocks included in the current block, the processing of generating a prediction image using the affine merge mode for the current block is terminated. That is, motion compensation is performed for the current block, and a prediction image of the current block is generated.
[0418] Additionally, in step Sk_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47a to 47c, the MV derivation methods shown in FIGS. 48a and 48b, the MV derivation methods shown in FIGS. 49a and 49b, and other MV derivation methods.
[0419] In addition, the candidate MV list may include candidate MVs of modes that perform predictions on a sub-block basis, other than affine mode.
[0420] In addition, as a candidate MV list, a candidate MV list may be generated that includes, for example, a candidate MV of an affine merge mode having two control points and a candidate MV of an affine merge mode having three control points. Alternatively, a candidate MV list may be generated that includes a candidate MV of an affine merge mode having two control points and a candidate MV of an affine merge mode having three control points, respectively. Alternatively, a candidate MV list may be generated that includes a candidate MV of either an affine merge mode having two control points or an affine merge mode having three control points. The candidate MV may be, for example, an MV of coded complete block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or an MV of a valid block among them.
[0421] In addition, as MV selection information, you may send an index indicating which candidate MV is in the list of candidate MVs.
[0422] [MV Derivation > Affine Mode > Affine Inter Mode]
[0423] Figure 51 is a flowchart showing an example of an affine intermode.
[0424] In the affine inter mode, first, the inter prediction unit (126) derives the predicted MV(v0, v1) or (v0, v1, v2) for each of the two or three control points of the current block (step Sj_1). The control points are the points of the upper left corner, upper right corner, or lower left corner of the current block, as shown in FIG. 46a or FIG. 46b.
[0425] For example, when using the method for deriving MVs shown in FIG. 48a and FIG. 48b, the inter prediction unit (126) derives the predicted MV (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the MV of any one of the encoded completed blocks near each control point of the current block shown in FIG. 48a or FIG. 48b. At this time, the inter prediction unit (126) encodes predicted MV selection information into a stream to identify the selected two or three predicted MVs.
[0426] For example, the inter prediction unit (126) may determine which block's MV from the encoding completed block adjacent to the current block to select as the predicted MV of the control point using cost evaluation, etc., and may describe a flag indicating which predicted MV was selected in the bit stream. That is, the inter prediction unit (126) outputs the predicted MV selection information, such as the flag, as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0427] Next, the inter prediction unit (126) performs motion search while updating the predicted MV selected or derived in step Sj_1 (step Sj_2) respectively (steps Sj_3 and Sj_4). That is, the inter prediction unit (126) sets the MV of each sub-block corresponding to the updated predicted MV as the affine MV and calculates it using the above-described Equation (1A) or Equation (1B) (step Sj_3). Then, the inter prediction unit (126) performs motion compensation for each sub-block using the affine MV and the encoded reference picture (step Sj_4). The processing of steps Sj_3 and Sj_4 is executed for all blocks within the current block whenever the predicted MV is updated in step Sj_2. As a result, the inter prediction unit (126) determines, for example, the predicted MV that yields the smallest cost in the motion search loop as the MV of the control point (step Sj_5). At this time, the inter prediction unit (126) also encodes the difference value between the determined MV and the predicted MV into the stream as the difference MV. That is, the inter prediction unit (126) outputs the difference MV as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0428] Finally, the inter prediction unit (126) generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0429] Additionally, in step Sj_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47a to 47c, the MV derivation methods shown in FIGS. 48a and 48b, the MV derivation methods shown in FIGS. 49a and 49b, and other MV derivation methods.
[0430] In addition, the candidate MV list may include candidate MVs of modes that perform predictions on a sub-block basis, other than affine mode.
[0431] In addition, as a candidate MV list, a candidate MV list may be generated that includes a candidate MV of an affine inter mode having two control points and a candidate MV of an affine inter mode having three control points. Alternatively, a candidate MV list may be generated that includes a candidate MV of an affine inter mode having two control points and a candidate MV of an affine inter mode having three control points, respectively. Alternatively, a candidate MV list may be generated that includes a candidate MV of either an affine inter mode having two control points or an affine inter mode having three control points. The candidate MV may be, for example, an MV of a completed coding block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or an MV of a valid block among them.
[0432] In addition, as prediction MV selection information, you may send an index indicating which candidate MV is in the list of candidate MVs.
[0433] [MV Derivation > Triangle Mode]
[0434] In the example described above, the inter prediction unit (126) generates one rectangular prediction image for the rectangular current block. However, the inter prediction unit (126) may generate a plurality of prediction images with shapes different from the rectangle for the rectangular current block and generate a final rectangular prediction image by combining the plurality of prediction images. The shapes different from the rectangle may be, for example, triangles.
[0435] FIG. 52a is a diagram illustrating the generation of a predicted image of two triangles.
[0436] The inter prediction unit (126) generates a predicted image of a triangle by performing motion compensation on a first partition of a triangle within a current block using the first MV of the first partition. Likewise, the inter prediction unit (126) generates a predicted image of a triangle by performing motion compensation on a second partition of a triangle within a current block using the second MV of the second partition. Then, the inter prediction unit (126) generates a predicted image of a rectangle similar to the current block by combining these predicted images.
[0437] Additionally, as a prediction image of the first partition, a first prediction image of a rectangle corresponding to the current block may be generated using the first MV. Also, as a prediction image of the second partition, a second prediction image of a rectangle corresponding to the current block may be generated using the second MV. A prediction image of the current block may be generated by weighted addition of the first prediction image and the second prediction image. Additionally, the area subject to weighted addition may be only a portion of the area between the boundary of the first partition and the second partition.
[0438] FIG. 52b is a conceptual diagram illustrating an example of a first portion of a first partition overlapping with a second partition, and a first sample set and a second sample set that may be weighted as part of a correction process. The first portion may, for example, be one-fourth of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to the edge of the first partition. Here, N is an integer greater than 0, for example, N may be an integer 2. FIG. 52b illustrates a rectangular partition having a rectangular portion with a width of one-fourth of the width of the first partition. Here, the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples within the first portion. The central example of FIG. 52b illustrates a rectangular partition having a rectangular portion with a height of one-fourth of the height of the first partition. Here, the first sample set includes a sample outside the first part and a sample inside the first part, and the second sample set includes a sample inside the first part. The example on the right side of FIG. 52b shows a triangular partition having a polygonal part with a height corresponding to two samples. Here, the first sample set includes a sample outside the first part and a sample inside the first part, and the second sample set includes a sample inside the first part.
[0439] The first part may be a part of the first partition that overlaps with an adjacent partition. FIG. 52c is a conceptual diagram showing the first part of the first partition that overlaps with a part of the first partition that overlaps with a part of the adjacent partition. For simplicity of explanation, a rectangular partition having a part that overlaps with a spatially adjacent rectangular partition is shown. Partitions having other shapes, such as a triangular partition, may be used, and the overlapping part may overlap with a spatially or temporally adjacent partition.
[0440] In addition, an example is shown in which a prediction image is generated for each of the two partitions using inter-prediction, but a prediction image may also be generated for at least one partition using intra-prediction.
[0441] FIG. 53 is a flowchart showing an example of a triangle mode.
[0442] In triangle mode, first, the inter prediction unit (126) divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter prediction unit (126) may encode partition information, which is information regarding the division into each partition, into a stream as a prediction parameter. That is, the inter prediction unit (126) may output the partition information as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0443] Next, the inter prediction unit (126) first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of encoded completed blocks located around the current block in time or space (step Sx_2). That is, the inter prediction unit (126) creates a list of candidate MVs.
[0444] Then, the inter prediction unit (126) selects the candidate MV of the first partition and the candidate MV of the second partition from among the plurality of candidate MVs obtained in step Sx_2 as the first MV and the second MV, respectively (step Sx_3). At this time, the inter prediction unit (126) may encode MV selection information for identifying the selected candidate MVs into a stream as a prediction parameter. That is, the inter prediction unit (126) may output the MV selection information as a prediction parameter to the entropy encoding unit (110) through the prediction parameter generation unit (130).
[0445] Next, the inter prediction unit (126) generates a first predicted image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Likewise, the inter prediction unit (126) generates a second predicted image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).
[0446] Finally, the inter prediction unit (126) generates a prediction image of the current block by adding the first prediction image and the second prediction image with weights (step Sx_6).
[0447] In addition, in the example shown in FIG. 52a, the first partition and the second partition are each triangular, but they may be trapezoidal or have different shapes. In addition, in the example shown in FIG. 52a, the current block is composed of two partitions, but it may be composed of three or more partitions.
[0448] Additionally, the first partition and the second partition may overlap. That is, the first partition and the second partition may contain the same pixel area. In this case, the predicted image of the current block may be generated using the predicted image in the first partition and the predicted image in the second partition.
[0449] In addition, this example shows an example where a predicted image is generated by inter-prediction with two partitions, but a predicted image may also be generated by intra-prediction for at least one partition.
[0450] In addition, the candidate MV list for selecting the first MV and the candidate MV list for selecting the second MV may be different, or they may be the same candidate MV list.
[0451] Additionally, the partition information may include an index indicating the direction of partitioning at least the current block into multiple partitions. The MV selection information may include an index indicating a selected first MV and an index indicating a selected second MV. A single index may represent multiple pieces of information. For example, a single index representing a combination of part or all of the partition information and part or all of the MV selection information may be encoded.
[0452] [MV Derivation > ATMVP Mode]
[0453] FIG. 54 is a diagram showing an example of an ATMVP mode in which MV is derived in sub-block units.
[0454] ATMVP mode is a mode classified as a merge mode. For example, in ATMVP mode, candidate MVs at the sub-block level are registered in the candidate MV list used in normal merge mode.
[0455] Specifically, in ATMVP mode, first, as shown in FIG. 54, a time MV reference block corresponding to the current block is identified in the encoding completed reference picture designated by the MV (MV0) of the block adjacent to the lower left of the current block. Next, for each sub-block within the current block, the MV used during encoding of the area corresponding to that sub-block within that time MV reference block is identified. The MV identified in this way is included in the candidate MV list as a candidate MV for the sub-block of the current block. When such a candidate MV for each sub-block is selected from the candidate MV list, motion compensation using that candidate MV as the MV of the sub-block is executed for that sub-block. By doing so, a predicted image of each sub-block is generated.
[0456] In addition, in the example shown in FIG. 54, a block adjacent to the lower left of the current block was used as the surrounding MV reference block, but other blocks may be used. Also, the size of the sub-block may be 4×4 pixels, 8×8 pixels, or any other size. The size of the sub-block may be converted into units such as slices, bricks, or pictures.
[0457] [Movement Exploration > DMVR]
[0458] Figure 55 is a diagram showing the relationship between the merge mode and the DMVR.
[0459] The inter prediction unit (126) derives the MV of the current block in merge mode (step Sl_1). Next, the inter prediction unit (126) determines whether to search for the MV, that is, to search for movement (step Sl_2). Here, if the inter prediction unit (126) determines that it does not search for movement (No in step Sl_2), it determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). That is, in this case, the MV of the current block is determined in merge mode.
[0460] Meanwhile, if it is determined that motion search is performed in step Sl_1 (Yes in step Sl_2), the inter prediction unit (126) derives the final MV for the current block by searching the surrounding area of the reference picture represented by the MV derived in step Sl_1 (step Sl_3). That is, in this case, the MV of the current block is determined by the DMVR.
[0461] FIG. 56 is a conceptual diagram illustrating an example of a DMVR for determining MV.
[0462] First, for example in merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), a reference pixel is identified from the first reference picture (L0), which is the encoded picture of the L0 list. Likewise, according to the candidate MV (L1), a reference pixel is identified from the second reference picture (L1), which is the encoded picture of the L1 list. A template is generated by taking the average of these reference pixels.
[0463] Next, using the template, the surrounding areas of the candidate MVs of the first reference picture (L0) and the second reference picture (L1) are searched, respectively, and the MV with the minimum cost is determined as the final MV of the current block. Additionally, the cost may be calculated using, for example, the difference between each pixel value of the template and each pixel value of the search area and the candidate MV value.
[0464] Even if it is not the processing described here, any processing that can derive the final MV by searching the surroundings of the candidate MV may be used.
[0465] FIG. 57 is a conceptual diagram illustrating another example of DMVR for determining MV. Unlike the example of DMVR shown in FIG. 56, the example shown in FIG. 57 calculates the cost without generating a template.
[0466] First, the inter prediction unit (126) searches around the reference blocks included in the reference pictures of the L0 list and L1 list, respectively, based on the initial MV, which is a candidate MV obtained from the candidate MV list. For example, as shown in FIG. 57, the initial MV corresponding to the reference block of the L0 list is InitMV_L0, and the initial MV corresponding to the reference block of the L1 list is InitMV_L1. In motion search, the inter prediction unit (126) first sets a search position for the reference picture of the L0 list. The difference vector representing the set search position, specifically the difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to the search position, is MVd_L0. Then, the inter prediction unit (126) determines the search position in the reference picture of the L1 list. This search position is represented by the difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. Specifically, the inter prediction unit (126) determines the difference vector as MVd_L1 by mirroring MVd_L0. That is, the inter prediction unit (126) sets the position that is symmetrical from the position indicated by the initial MV in the reference picture of the L0 list and the L1 list as the search position. For each search position, the inter prediction unit (126) calculates the total sum of absolute difference values (SAD) of pixel values within the block at that search position as the cost, and finds the search position where the cost is minimized.
[0467] FIG. 58a is a diagram showing an example of motion search in DMVR, and FIG. 58b is a flowchart showing an example of motion search.
[0468] First, the inter prediction unit (126) calculates the cost at the search location (also called the starting point) indicated by the initial MV in Step 1 and at eight search locations surrounding it. Then, the inter prediction unit (126) determines whether the cost of the search locations other than the starting point is the minimum. Here, if the inter prediction unit (126) determines that the cost of the search locations other than the starting point is the minimum, it moves to the search location where the cost is the minimum and performs the processing of Step 2. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit (126) skips the processing of Step 2 and performs the processing of Step 3.
[0469] In Step 2, the inter prediction unit (126) sets the search location moved according to the processing result of Step 1 as a new starting point and performs the same search as in Step 1. Then, the inter prediction unit (126) determines whether the cost of the search location other than the starting point is the minimum. Here, if the cost of the search location other than the starting point is the minimum, the inter prediction unit (126) performs the processing of Step 4. Meanwhile, if the cost of the starting point is the minimum, the inter prediction unit (126) performs the processing of Step 3.
[0470] In Step 4, the inter prediction unit (126) treats the search position of the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as a difference vector.
[0471] In Step 3, the inter prediction unit (126) determines a pixel location with fractional precision where the cost is minimized based on the cost at four points located above, below, left, and right of the starting point of Step 1 or Step 2, and sets that pixel location as the final search location. That pixel location with fractional precision is determined by weighted addition of the vectors of four points located above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)), using the cost at each of the four search locations as a weight. Then, the inter prediction unit (126) determines the difference between the location indicated by the initial MV and the final search location as a difference vector.
[0472] [Movement Compensation > BIO / OBMC / LIC]
[0473] In motion compensation, there is a mode for generating a prediction image and correcting the prediction image. The mode is, for example, BIO, OBMC, and LIC described below.
[0474] Figure 59 is a flowchart showing an example of the generation of a predicted image.
[0475] The inter prediction unit (126) generates a prediction image (step Sm_1) and corrects the prediction image by any one of the modes described above (step Sm_2).
[0476] Figure 60 is a flowchart showing another example of the generation of a predicted image.
[0477] The inter prediction unit (126) derives the MV of the current block (step Sn_1). Next, the inter prediction unit (126) generates a prediction image using the MV (step Sn_2) and determines whether to perform correction processing (step Sn_3). Here, if the inter prediction unit (126) determines that correction processing is to be performed (Yes in step Sn_3), it generates a final prediction image by correcting the prediction image (step Sn_4). In addition, in the LIC described later, luminance and color difference may be corrected in step Sn_4. On the other hand, if the inter prediction unit (126) determines that correction processing is not to be performed (No in step Sn_3), it outputs the prediction image as a final prediction image without correcting it (step Sn_5).
[0478] [Movement Reward > OBMC]
[0479] Inter-predicted images may be generated using not only motion information of the current block obtained by motion search, but also motion information of adjacent blocks. Specifically, inter-predicted images may be generated at the sub-block level within the current block by weighted addition of a prediction image based on motion information obtained by motion search (in the reference picture) and a prediction image based on motion information of adjacent blocks (in the current picture). Such inter-predicted (motion compensation) is sometimes referred to as OBMC (overlapped block motion compensation) or OBMC mode.
[0480] In OBMC mode, information indicating the size of a sub-block for OBMC (e.g., called OBMC block size) may be signaled at the sequence level. Additionally, information indicating whether to apply OBMC mode (e.g., called OBMC flag) may be signaled at the CU level. Furthermore, the signaling levels of these information are not limited to the sequence level and CU level, and may be other levels (e.g., picture level, slice level, brick level, CTU level, or sub-block level).
[0481] The OBMC mode will be explained in more detail. Figures 61 and 62 are a flowchart and a conceptual diagram illustrating an overview of the prediction image correction processing by OBMC.
[0482] First, as shown in FIG. 62, a predicted image (Pred) is obtained by normal motion compensation using the MV assigned to the current block. In FIG. 62, the arrow "MV" indicates a reference picture and shows what the current block of the current picture is referencing to obtain the predicted image.
[0483] Next, the MV (MV_L) already derived for the left adjacent block of completed encoding is applied (reused) to the current block to obtain a predicted image (Pred_L). The MV (MV_L) is represented by an arrow "MV_L" pointing from the current block to a reference picture. Then, the first correction of the predicted image is performed by overlapping the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0484] Likewise, the MV (MV_U) already derived for the upper adjacent block of completed encoding is applied (reused) to the current block to obtain a predicted image (Pred_U). The MV (MV_U) is represented by an arrow "MV_U" pointing from the current block to a reference picture. Then, a second correction of the predicted image is performed by superimposing the predicted image Pred_U onto the predicted image (e.g., Pred and Pred_L) that has undergone a first correction. This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block in which the boundaries with adjacent blocks are blended (smoothed).
[0485] In addition, the example described above is a two-pass correction method using left adjacent and upper adjacent blocks, but the correction method may also be a three-pass or more-pass correction method using right adjacent and / or lower adjacent blocks.
[0486] In addition, the area for overlapping does not have to be the entire pixel area of the block, but only a part of the area near the block boundary.
[0487] In addition, the OBMC prediction image correction process for obtaining a single prediction image Pred by superimposing additional prediction images Pred_L and Pred_U from a single reference picture has been described here. However, when a prediction image is corrected based on multiple reference pictures, the same processing may be applied to each of the multiple reference pictures. In such cases, by performing OBMC image correction based on multiple reference pictures, a corrected prediction image is obtained from each reference picture, and then the obtained multiple corrected prediction images are further superimposed to obtain a final prediction image.
[0488] In addition, in OBMC, the unit of the current block may be a PU unit, or a sub-block unit formed by further dividing the PU.
[0489] As a method for determining whether to apply OBMC, for example, there is a method of using obmc_flag, which is a signal indicating whether to apply OBMC. As a specific example, the encoding device (100) may determine whether the current block belongs to a complex region of motion. If the encoding device (100) belongs to a complex region of motion, it sets the value 1 as obmc_flag to apply OBMC and performs encoding, and if it does not belong to a complex region of motion, it sets the value 0 as obmc_flag to perform encoding of the block without applying OBMC. Meanwhile, the decoding device (200) performs decoding by decoding the obmc_flag described in the stream, switching whether to apply OBMC according to the value.
[0490] [Movement Compensation > BIO]
[0491] Next, we will explain how to derive the MV. First, we will explain the mode for deriving the MV based on a model assuming uniform linear motion. This mode is sometimes called the BIO (bi-directional optical flow) mode. Also, this bi-directional optical flow may be denoted as BDOF instead of BIO.
[0492] FIG. 63 is a diagram for explaining a model assuming uniform linear motion. In FIG. 63, (vx, vy) represents a velocity vector, and τ0 and τ1 represent the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1), respectively. (MVx0, MVy0) represents the MV corresponding to the reference picture Ref0, and (MVx1, MVy1) represents the MV corresponding to the reference picture Ref1.
[0493] At this time, under the assumption of uniform linear motion of the velocity vector (vx, vy), (MVx0, MVy0) and (MVx1, MVy1) are represented as (vxτ0, vyτ0) and (-vxτ1, -vyτ1), respectively, and the following optical flow equation (2) holds.
[0494]
[0495] Here, I(k) represents the luminance value of reference image k (k=0, 1) after motion compensation. This optical flow equation indicates that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is zero. Based on a combination of this optical flow equation and Hermite interpolation, block-unit motion vectors obtained from candidate MV lists, etc., may be corrected at the pixel level.
[0496] In addition, the MV may be derived from the decoder (200) side in a manner different from the derivation of the motion vector based on a model assuming uniform linear motion. For example, the motion vector may be derived in sub-block units based on the MV of multiple adjacent blocks.
[0497] FIG. 64 is a flowchart showing an example of inter prediction based on BIO. FIG. 65 is a diagram showing an example of the functional configuration of an inter prediction unit (126) that performs inter prediction based on BIO.
[0498] As shown in FIG. 65, the inter prediction unit (126) comprises, for example, a memory (126a), an interpolation image derivation unit (126b), a gradient image derivation unit (126c), an optical flow derivation unit (126d), a correction value derivation unit (126e), and a prediction image correction unit (126f). Additionally, the memory (126a) may be a frame memory (122).
[0499] The inter prediction unit (126) derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) that are different from the picture (Cur Pic) containing the current block. Then, the inter prediction unit (126) derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_1). Additionally, motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0500] Next, the interpolation image derivation unit (126b) refers to the memory (126a) and uses the motion vector M0 and reference picture L0 to derive the interpolation image I of the current block. 0 It derives the interpolated image. Also, the interpolated image derivation unit (126b) refers to the memory (126a) and uses the motion vector M1 and reference picture L1 to derive the interpolated image I of the current block. 1 Derive (Step Sy_2). Here, the interpolated image I 0 is an image included in reference picture Ref0, derived for the current block, and interpolated image I 1 is an image included in reference picture Ref1, derived for the current block. Interpolated image I 0 and interpolated image I 1 Each may be the same size as the current block. Or, interpolated image I 0 and interpolated image I 1 Each of these may be an image larger than the current block in order to properly derive the gradient image described later. In addition, the interpolated image (I 0 and I 1 ) may include motion vectors (M0, M1) and reference pictures (L0, L1), and a predicted image derived by applying a motion compensation filter.
[0501] Also, the gradient image extraction part (126c) is the interpolated image I0 and interpolated image I 1 From, the gradient image of the current block (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) derives (step Sy_3). Also, the horizontal gradient image is, (Ix 0 , Ix 1 ) and the vertical gradient image is, (Iy 0 , Iy 1 The gradient image derivation unit (126c) may derive the gradient image by, for example, applying a gradient filter to an interpolated image. The gradient image may represent the amount of spatial change of pixel values along the horizontal or vertical direction.
[0502] Next, the optical flow derivation unit (126d) is an interpolated image (I) in units of a plurality of sub-blocks constituting a current block. 0 , I 1 ) and gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 Using ), the optical flow (vx, vy), which is the velocity vector described above, is derived (step Sy_4). The optical flow is a coefficient that corrects the spatial displacement of a pixel and may be called a local motion estimate, a corrected motion vector, or a corrected weight vector. As an example, the sub-block may be a 4×4 pixel sub-CU. In addition, the derivation of the optical flow may be performed at a different unit, such as a pixel unit, rather than at a sub-block unit.
[0503] Next, the inter prediction unit (126) corrects the predicted image of the current block using optical flows (vx, vy). For example, the correction value derivation unit (126e) derives a correction value for the value of a pixel included in the current block using optical flows (vx, vy) (step Sy_5). Then, the predicted image correction unit (126f) may correct the predicted image of the current block using the correction value (step Sy_6). Additionally, the correction value may be derived on a pixel-by-pixel basis, or may be derived on a multiple pixel-by-pixel basis or on a sub-block basis.
[0504] In addition, the processing flow of BIO is not limited to the processing disclosed in FIG. 64. Only some of the processing disclosed in FIG. 64 may be performed, different processing may be added or replaced, or different processing may be executed in a different order.
[0505] [Movement Compensation > LIC]
[0506] Next, an example of a mode that generates a predicted image (prediction) using LIC (local illumination compensation) will be described.
[0507] FIG. 66a is a diagram illustrating an example of a method for generating a predicted image using luminance correction processing by LIC. FIG. 66b is a flowchart showing an example of a method for generating a predicted image using the LIC.
[0508] First, the inter prediction unit (126) derives an MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).
[0509] Next, the inter prediction unit (126) extracts information indicating how the luminance value has changed in the reference picture and the current picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference area (peripheral reference area) and the encoded upper adjacent reference area (peripheral reference area) in the current picture, and the luminance pixel values at the same location within the reference picture designated as the derived MV. Then, the inter prediction unit (126) calculates a luminance correction parameter using the information indicating how the luminance value has changed (step Sz_3).
[0510] The inter prediction unit (126) generates a prediction image for the current block by performing a luminance correction process that applies the luminance correction parameter to the reference image within the reference picture designated as MV (step Sz_4). That is, a correction based on the luminance correction parameter is performed on the prediction image, which is the reference image within the reference picture designated as MV. In this correction, the luminance may be corrected or the color difference may be corrected. That is, a color difference correction parameter is calculated using information indicating how the color difference has changed, and a color difference correction process is performed.
[0511] In addition, the shape of the surrounding reference area in FIG. 66a is an example, and other shapes may be used.
[0512] In addition, the process of generating a predicted image from one reference picture has been described here, but the same applies when generating a predicted image from multiple reference pictures, and the predicted image may be generated after performing luminance correction processing on the reference image obtained from each reference picture in the same way as described above.
[0513] As a method for determining whether to apply LIC, for example, there is a method of using lic_flag, which is a signal indicating whether to apply LIC. As a specific example, in an encoding device (100), a current block is determined to be in a region where a luminance change occurs, and if it is in a region where a luminance change occurs, a value of 1 is set as lic_flag to apply LIC and perform encoding, and if it is not in a region where a luminance change occurs, a value of 0 is set as lic_flag to perform encoding without applying LIC. Meanwhile, in a decoding device (200), by decoding the lic_flag described in the stream, decoding may be performed by switching whether to apply LIC according to the value.
[0514] As another method for determining whether to apply LIC, for example, there is also a method of determining based on whether LIC has been applied in a surrounding block. As a specific example, when the current block is processed in merge mode, the inter prediction unit (126) determines whether the surrounding completed encoding block selected when deriving the MV in merge mode has been encoded by applying LIC. The inter prediction unit (126) performs encoding by switching whether to apply LIC based on the result. In addition, in this example, the same processing is applied to the processing on the side of the decoder (200).
[0515] LIC (luminance correction processing) was explained using FIGS. 66a and FIGS. 66b, and the details thereof will be explained below.
[0516] First, the inter prediction unit (126) derives an MV for obtaining a reference image corresponding to the current block from a reference picture, which is an encoded finished picture.
[0517] Next, the inter prediction unit (126) calculates a luminance correction parameter by extracting information indicating how the luminance value has changed in the reference picture and the current picture using the luminance pixel value of the encoding completed surrounding reference area of the left adjacent and upper adjacent sides for the current block and the luminance pixel value at an equivalent position within the reference picture designated as MV. For example, the luminance pixel value of a pixel in the surrounding reference area within the current picture is set to p0, and the luminance pixel value of a pixel in the surrounding reference area within the reference picture at an equivalent position to the pixel is set to p1. The inter prediction unit (126) calculates coefficients A and B as luminance correction parameters that optimize A×p1+B=p0 for a plurality of pixels in the surrounding reference area.
[0518] Next, the inter prediction unit (126) generates a prediction image for the current block by performing a luminance correction process on a reference image within a reference picture designated as MV using a luminance correction parameter. For example, the luminance pixel value within the reference image is set to p2, and the luminance pixel value of the prediction image after the luminance correction process is set to p3. The inter prediction unit (126) generates a prediction image after the luminance correction process by calculating A×p2+B=p3 for each pixel within the reference image.
[0519] Additionally, a portion of the peripheral reference area shown in FIG. 66a may be used. For example, an area containing a predetermined number of pixels selected from each of the upper adjacent pixel and the left adjacent pixel may be used as the peripheral reference area. Furthermore, the peripheral reference area is not limited to an area adjacent to the current block, but may be an area not adjacent to the current block. Also, in the example shown in FIG. 66a, the peripheral reference area within the reference picture is an area designated as the MV of the current picture from the peripheral reference area within the current picture, but may be an area designated as a different MV. For example, the other MV may be the MV of the peripheral reference area within the current picture.
[0520] In addition, the operation of the encoding device (100) has been described here, and the operation of the decoding device (200) is the same.
[0521] In addition, LIC may be applied to color difference as well as luminance. In this case, correction parameters may be derived individually for Y, Cb, and Cr, respectively, or a common correction parameter may be used for any one of them.
[0522] In addition, LIC processing may be applied on a sub-block basis. For example, correction parameters may be derived using the surrounding reference area of the current sub-block and the surrounding reference area of the reference sub-block within the reference picture designated as the MV of the current sub-block.
[0523] [Predictive Control Unit]
[0524] The prediction control unit (128) selects either an intra prediction image (an image or signal output from the intra prediction unit (124)) or an inter prediction image (an image or signal output from the inter prediction unit (126)) and outputs the selected prediction image to the subtraction unit (104) and the addition unit (116).
[0525] [Prediction Parameter Generation Unit]
[0526] The prediction parameter generation unit (130) may output information regarding intra prediction, inter prediction, and selection of a prediction image in the prediction control unit (128) as a prediction parameter to the entropy encoding unit (110). The entropy encoding unit (110) may generate a stream based on the prediction parameter input from the prediction parameter generation unit (130) and the quantization coefficient input from the quantization unit (108). The prediction parameter may be used in the decoder (200). The decoder (200) may receive the stream, decode it, and perform processing such as the prediction processing performed in the intra prediction unit (124), the inter prediction unit (126), and the prediction control unit (128). The prediction parameter may include any index, flag, or value based on a selected prediction signal (e.g., MV, prediction type, or prediction mode used in the intra prediction unit (124) or the inter prediction unit (126)), or prediction processing performed in the intra prediction unit (124), the inter prediction unit (126), and the prediction control unit (128), or representing the prediction processing.
[0527] [Decoding device]
[0528] Next, a decoding device (200) capable of decoding a stream output from the above-described encoding device (100) will be described. FIG. 67 is a block diagram showing an example of the functional configuration of the decoding device (200) according to an embodiment. The decoding device (200) is a device that decodes a stream of encoded images in blocks.
[0529] As shown in FIG. 67, the decoder (200) comprises an entropy decoder (202), an inverse quantization unit (204), an inverse transform unit (206), an adder (208), a block memory (210), a loop filter unit (212), a frame memory (214), an intra prediction unit (216), an inter prediction unit (218), a prediction control unit (220), a prediction parameter generation unit (222), and a division decision unit (224). Additionally, the intra prediction unit (216) and the inter prediction unit (218) are each configured as part of a prediction processing unit.
[0530] [Example of implementation of a decoding device]
[0531] FIG. 68 is a block diagram showing an example of implementation of a decoding device (200). The decoding device (200) includes a processor (b1) and a memory (b2). For example, a plurality of components of the decoding device (200) shown in FIG. 67 are implemented by the processor (b1) and memory (b2) shown in FIG. 68.
[0532] The processor (b1) is a circuit that performs information processing and is capable of accessing memory (b2). For example, the processor (b1) is a dedicated or general-purpose electronic circuit that decodes a stream. The processor (b1) may be a processor such as a CPU. Also, the processor (b1) may be an assembly of multiple electronic circuits. Also, for example, the processor (b1) may serve as a component of the decoding device (200) shown in FIG. 67, excluding the component for storing information.
[0533] Memory (b2) is a dedicated or general-purpose memory in which information for the processor (b1) to decode a stream is stored. Memory (b2) may be an electronic circuit or may be connected to the processor (b1). Also, memory (b2) may be included in the processor (b1). Also, memory (b2) may be an assembly of multiple electronic circuits. Also, memory (b2) may be a magnetic disk or an optical disk, etc., or may be expressed as a storage or a recording medium. Also, memory (b2) may be a non-volatile memory or a volatile memory.
[0534] For example, an image may be stored in memory (b2) or a stream may be stored. Also, a program for the processor (b1) to decode the stream may be stored in memory (b2).
[0535] Also, for example, memory (b2) may serve as a component for storing information among the multiple components of the decoding device (200) shown in FIG. 67, etc. Specifically, memory (b2) may serve as a block memory (210) and frame memory (214) shown in FIG. 67. More specifically, a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.) may be stored in memory (b2).
[0536] In addition, in the decoding device (200), not all of the multiple components shown in FIG. 67, etc., are implemented, and not all of the multiple processes described above are performed. Some of the multiple components shown in FIG. 67, etc., may be included in another device, and some of the multiple processes described above may be performed by another device.
[0537] Hereinafter, after explaining the overall processing flow of the decoding device (200), each component included in the decoding device (200) will be described. In addition, regarding each component included in the decoding device (200) that performs the same processing as the component included in the encoding device (100), a detailed description is omitted. For example, the inverse quantization unit (204), inverse transformation unit (206), adder (208), block memory (210), frame memory (214), intra prediction unit (216), inter prediction unit (218), prediction control unit (220), and loop filter unit (212) included in the decoder (200) each perform the same processing as the inverse quantization unit (112), inverse transformation unit (114), adder (116), block memory (118), frame memory (122), intra prediction unit (124), inter prediction unit (126), prediction control unit (128), and loop filter unit (120) included in the encoding device (100).
[0538] [Full flow of decoding processing]
[0539] FIG. 69 is a flowchart showing an example of overall decoding processing by a decoding device (200).
[0540] First, the division determination unit (224) of the decoder (200) determines the division pattern of each of the plurality of fixed-size blocks (128×128 pixels) included in the picture based on parameters input from the entropy decoder (202) (step Sp_1). This division pattern is a division pattern selected by the encoding device (100). Then, the decoder (200) performs the processing of steps Sp_2 to Sp_6 for each of the plurality of blocks constituting the division pattern.
[0541] The entropy decoding unit (202) decodes the encoded quantization coefficients and prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).
[0542] Next, the inverse quantization unit (204) and the inverse transformation unit (206) restore the predicted residual of the current block by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (step Sp_3).
[0543] Next, the prediction processing unit, consisting of an intra prediction unit (216), an inter prediction unit (218), and a prediction control unit (220), generates a prediction image of the current block (step Sp_4).
[0544] Next, the adder (208) reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted image to the predicted residual (step Sp_5).
[0545] And, when this reconstructed image is generated, the loop filter unit (212) performs filtering on the reconstructed image (step Sp_6).
[0546] Then, the decoding device (200) determines whether the decoding of the entire picture is complete (step Sp_7), and if it determines that it is not complete (No in step Sp_7), it repeats the processing from step Sp_1.
[0547] Additionally, the processing of these steps Sp_1 to Sp_7 may be performed sequentially by a decoding device (200), and some of the processing may be performed in parallel or the order may be changed.
[0548] [Division Decision Section]
[0549] FIG. 70 is a diagram showing the relationship between the division decision unit (224) and other components. The division decision unit (224) may perform the following processing as an example.
[0550] The division decision unit (224) collects block information from, for example, a block memory (210) or a frame memory (214) and also obtains parameters from the entropy decoding unit (202). Then, the division decision unit (224) may determine a division pattern of a fixed-size block based on the block information and parameters. Then, the division decision unit (224) may output information representing the determined division pattern to the inverse transformation unit (206), the intra prediction unit (216), and the inter prediction unit (218). The inverse transformation unit (206) may perform an inverse transformation on the transformation coefficients based on the division pattern represented by the information from the division decision unit (224). The intra prediction unit (216) and the inter prediction unit (218) may generate a prediction image based on the division pattern represented by the information from the division decision unit (224).
[0551] [Entropy Decoding Section]
[0552] FIG. 71 is a block diagram showing an example of the functional configuration of the entropy decoding unit (202).
[0553] The entropy decoding unit (202) generates quantization coefficients, prediction parameters, and parameters regarding the partitioning pattern, etc., by entropy decoding the stream. For example, CABAC is used for the entropy decoding. Specifically, the entropy decoding unit (202) is equipped with, for example, a binary arithmetic decoding unit (202a), a context control unit (202b), and a multi-valued unit (202c). The binary arithmetic decoding unit (202a) arithmetic decodes the stream into a binary signal using a context value derived by the context control unit (202b). The context control unit (202b), similar to the context control unit (110b) of the encoding device (100), derives a context value, i.e., the probability of occurrence of a binary signal, based on the characteristics of the syntax element or surrounding conditions. The debinarization unit (202c) performs debinarization by converting the 2-value signal output from the 2-value arithmetic decoding unit (202a) into a debinarized signal representing the aforementioned quantization coefficients, etc. This debinarization is performed according to the debinarization method described above.
[0554] The entropy decoder (202) outputs quantization coefficients in blocks to the inverse quantization unit (204). The entropy decoder (202) may output prediction parameters included in the stream (see FIG. 1) to the intra prediction unit (216), the inter prediction unit (218), and the prediction control unit (220). The intra prediction unit (216), the inter prediction unit (218), and the prediction control unit (220) may perform prediction processing such as the processing performed by the intra prediction unit (124), the inter prediction unit (126), and the prediction control unit (128) on the side of the encoding device (100).
[0555] [Entropy Decoding Section]
[0556] FIG. 72 is a diagram showing the flow of CABAC in the entropy decoding unit (202).
[0557] First, initialization is performed in the CABAC of the entropy decoding unit (202). In this initialization, initialization in the binary arithmetic decoding unit (202a) and the setting of the initial context value are performed. Then, the binary arithmetic decoding unit (202a) and the multi-valuation unit (202c) perform arithmetic decoding and multi-valuation on, for example, the encoded data of the CTU. At this time, the context control unit (202b) updates the context value whenever arithmetic decoding is performed. Then, the context control unit (202b) resets the context value as a post-processing step. This reset context value is used, for example, as the initial value of the context value for the next CTU.
[0558] [Inverse Quantumization Unit]
[0559] The inverse quantization unit (204) inversely quantizes the quantization coefficients of the current block, which are inputs from the entropy decoding unit (202). Specifically, the inverse quantization unit (204) inversely quantizes each quantization coefficient of the current block based on a quantization parameter corresponding to the quantization coefficient. Then, the inverse quantization unit (204) outputs the inversely quantized quantization coefficients (i.e., conversion coefficients) of the current block to the inverse conversion unit (206).
[0560] FIG. 73 is a block diagram showing an example of the functional configuration of the inverse quantization unit (204).
[0561] The inverse quantization unit (204) comprises, for example, a quantization parameter generation unit (204a), a prediction quantization parameter generation unit (204b), a quantization parameter memory unit (204d), and an inverse quantization processing unit (204e).
[0562] FIG. 74 is a flowchart showing an example of inverse quantization by an inverse quantization unit (204).
[0563] The inverse quantization unit (204) may, for example, perform inverse quantization processing for each CU according to the flow shown in FIG. 74. Specifically, the quantization parameter generation unit (204a) determines whether to perform inverse quantization (step Sv_11). Here, if it is determined that inverse quantization is to be performed (Yes in step Sv_11), the quantization parameter generation unit (204a) obtains the difference quantization parameter of the current block from the entropy decoding unit (202) (step Sv_12).
[0564] Next, the predicted quantization parameter generation unit (204b) obtains a quantization parameter of a processing unit different from the current block from the quantization parameter memory unit (204d) (step Sv_13). The predicted quantization parameter generation unit (204b) generates a predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).
[0565] Then, the quantization parameter generation unit (204a) adds the difference quantization parameter of the current block obtained from the entropy decoding unit (202) and the prediction quantization parameter of the current block generated by the prediction quantization parameter generation unit (204b) (step Sv_15). By this addition, the quantization parameter of the current block is generated. Also, the quantization parameter generation unit (204a) stores the quantization parameter of the current block in the quantization parameter memory unit (204d) (step Sv_16).
[0566] Next, the inverse quantization processing unit (204e) inversely quantizes the quantization coefficients of the current block into conversion coefficients using the quantization parameters generated in step Sv_15 (step Sv_17).
[0567] Additionally, the differential quantization parameter may be decoded at the bit sequence level, picture level, slice level, brick level, or CTU level. Also, the initial value of the quantization parameter may be decoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the quantization parameter may be generated using the initial value of the quantization parameter and the differential quantization parameter.
[0568] Additionally, the inverse quantization unit (204) may be equipped with a plurality of inverse quantizers, and may inverse quantize the quantization coefficient using an inverse quantization method selected from a plurality of inverse quantization methods.
[0569] [Inverse Transformation Section]
[0570] The inverse transformation unit (206) restores the predicted residual by inversely transforming the transformation coefficients, which are inputs from the inverse quantization unit (204).
[0571] For example, if the information decoded from the stream indicates that EMT or AMT is applied (e.g., the AMT flag is true), the inverse conversion unit (206) inversely converts the conversion coefficients of the current block based on the information indicating the decoded conversion type.
[0572] Also, for example, if the information decoded from the stream indicates that NSST is applied, the inverse transformation unit (206) applies an inverse transformation to the transformation coefficients.
[0573] FIG. 75 is a flowchart showing an example of processing by the inverse conversion unit (206).
[0574] For example, the inverse transform unit (206) determines whether there is information in the stream indicating that no orthogonal transformation is performed (step St_11). Here, if it is determined that such information does not exist (No in step St_11), the inverse transform unit (206) obtains information indicating a transformation type that is decoded by the entropy decoding unit (202) (step St_12). Next, the inverse transform unit (206) determines the transformation type used for the orthogonal transformation of the encoding device (100) based on the information (step St_13). Then, the inverse transform unit (206) performs an inverse orthogonal transformation using the determined transformation type (step St_14).
[0575] FIG. 76 is a flowchart showing another example of processing by the inverse conversion unit (206).
[0576] For example, the inverse transformation unit (206) determines whether the transformation size is less than or equal to a predetermined value (step Su_11). Here, if it is determined that it is less than or equal to a predetermined value (Yes in step Su_11), the inverse transformation unit (206) obtains information from the entropy decoding unit (202) indicating which of the one or more transformation types included in the first transformation type group was used by the encoding device (100) (step Su_12). In addition, this information is decoded by the entropy decoding unit (202) and output to the inverse transformation unit (206).
[0577] The inverse transform unit (206) determines the transformation type used for orthogonal transformation in the encoding device (100) based on the information (step Su_13). Then, the inverse transform unit (206) performs inverse orthogonal transformation of the transformation coefficients of the current block using the determined transformation type (step Su_14). Meanwhile, if the inverse transform unit (206) determines in step Su_11 that the transformation size is not less than or equal to a predetermined value (No in step Su_11), it performs inverse orthogonal transformation of the transformation coefficients of the current block using a second transformation type group (step Su_15).
[0578] In addition, the inverse orthogonal transformation by the inverse transformation unit (206) may be performed according to the flow shown in FIG. 75 or FIG. 76 for each TU as an example. In addition, the inverse orthogonal transformation may be performed using a predetermined transformation type without decoding information indicating the transformation type used for the orthogonal transformation. In addition, the transformation type is specifically DST7 or DCT8, and in the inverse orthogonal transformation, an inverse transformation basis function corresponding to that transformation type is used.
[0579] [Additional]
[0580] The adder (208) reconstructs the current block by adding the predicted residual, which is an input from the inverse transform unit (206), and the predicted image, which is an input from the prediction control unit (220). That is, a reconstructed image of the current block is generated. Then, the adder (208) outputs the reconstructed image of the current block to the block memory (210) and the loop filter unit (212).
[0581] [Block Memory]
[0582] The block memory (210) is a block referenced in intra-prediction and is a memory unit for storing blocks within the current picture. Specifically, the block memory (210) stores the reconstructed image output from the adder (208).
[0583] [Loop Filter Section]
[0584] The loop filter unit (212) performs a loop filter on the reconstructed image generated by the adder (208) and outputs the reconstructed image with the filter applied to the frame memory (214) and display device, etc.
[0585] When the information indicating the on / off status of the ALF decoded from the stream indicates that the ALF is on, one filter is selected from a plurality of filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0586] FIG. 77 is a block diagram showing an example of the functional configuration of a loop filter section (212). Additionally, the loop filter section (212) has the same configuration as the loop filter section (120) of the encoding device (100).
[0587] The loop filter unit (212) comprises, for example, a deblocking filter processing unit (212a), an SAO processing unit (212b), and an ALF processing unit (212c), as shown in FIG. 77. The deblocking filter processing unit (212a) performs the deblocking filter processing described above on the reconstructed image. The SAO processing unit (212b) performs the SAO processing described above on the reconstructed image after the deblocking filter processing. Additionally, the ALF processing unit (212c) applies the ALF processing described above to the reconstructed image after the SAO processing. Furthermore, the loop filter unit (212) does not need to have all the processing units disclosed in FIG. 77, and may have only some of the processing units. Also, the loop filter unit (212) may be configured to perform each of the processing units described above in a different order from the processing order disclosed in FIG. 77.
[0588] [Frame Memory]
[0589] The frame memory (214) is a memory unit for storing a reference picture used for inter prediction, and is also called a frame buffer. Specifically, the frame memory (214) stores a reconstructed image that has been filtered by the loop filter unit (212).
[0590] [Prediction Unit (Intra Prediction Unit · Inter Prediction Unit · Prediction Control Unit)]
[0591] FIG. 78 is a flowchart illustrating an example of processing performed in the prediction unit of a decoding device (200). Additionally, as an example, the prediction unit is composed of all or part of the components of an intra prediction unit (216), an inter prediction unit (218), and a prediction control unit (220). The prediction processing unit includes, for example, an intra prediction unit (216) and an inter prediction unit (218).
[0592] The prediction unit generates a prediction image of the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. Additionally, the prediction signal may include, for example, an intra prediction signal or an inter prediction signal. Specifically, the prediction unit generates a prediction image of the current block using a reconstruction image that has already been obtained by generating a prediction image for another block, restoring the prediction residual, and adding the prediction images. The prediction unit of the decoder (200) generates a prediction image identical to the prediction image generated by the prediction unit of the encoding unit (100). That is, the methods for generating prediction images used in the prediction units are common or corresponding to each other.
[0593] The reconstructed image may be, for example, an image of a reference picture, or an image of a decoded block (i.e., another block described above) within a current picture that includes the current block. The decoded block within the current picture is, for example, an adjacent block of the current block.
[0594] FIG. 79 is a flowchart showing another example of processing performed in the prediction section of the decoder (200).
[0595] The prediction unit determines a method or mode for generating a prediction image (step Sr_1). For example, this method or mode may be determined based on prediction parameters, for example.
[0596] If the prediction unit determines a first method as a mode for generating a prediction image, it generates a prediction image according to the first method (step Sr_2a). Additionally, if the prediction unit determines a second method as a mode for generating a prediction image, it generates a prediction image according to the second method (step Sr_2b). Additionally, if the prediction unit determines a third method as a mode for generating a prediction image, it generates a prediction image according to the third method (step Sr_2c).
[0597] The first, second, and third methods are different methods for generating a prediction image, and may each be, for example, an inter-prediction method, an intra-prediction method, and a prediction method other than those mentioned above. In these prediction methods, the reconstructed image described above may be used.
[0598] FIGS. 80a and FIGS. 80b are flowcharts showing other examples of processing performed in the prediction section of the decoder (200).
[0599] The prediction unit may perform prediction processing according to the flow shown in FIGS. 80a and FIGS. 80b as an example. In addition, the intra-block copy shown in FIGS. 80a and FIGS. 80b is a mode belonging to inter-prediction, and is a mode in which a block included in the current picture is referenced as a reference image or a reference block. That is, in the intra-block copy, a picture different from the current picture is not referenced. In addition, the PCM mode shown in FIG. 80a is a mode belonging to intra-prediction, and is a mode in which conversion and quantization are not performed.
[0600] [Intra Prediction Department]
[0601] The intra prediction unit (216) generates a predicted image of a current block (i.e., an intra prediction image) by performing intra prediction based on an intra prediction mode decoded from a stream and referencing a block within a current picture stored in a block memory (210). Specifically, the intra prediction unit (216) generates an intra prediction image by performing intra prediction based on pixel values (e.g., luminance values, chrominance values) of a block adjacent to the current block, and outputs the intra prediction image to the prediction control unit (220).
[0602] In addition, when an intra prediction mode that references a luminance block is selected for intra prediction of a color difference block, the intra prediction unit (216) may predict the color difference component of the current block based on the luminance component of the current block.
[0603] Also, when the information decoded from the stream indicates the application of PDPC, the intra prediction unit (216) corrects the pixel value after intra prediction based on the gradient of the reference pixel in the horizontal / vertical direction.
[0604] FIG. 81 is a diagram showing an example of processing by the intra prediction unit (216) of the decoding device (200).
[0605] The intra prediction unit (216) first determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, if it is determined that an MPM flag indicating 1 exists (Yes in step Sw_11), the intra prediction unit (216) obtains information indicating an intra prediction mode selected in the encoding device (100) among the MPMs from the entropy decoder (202) (step Sw_12). In addition, the information is decoded by the entropy decoder (202) and output to the intra prediction unit (216). Next, the intra prediction unit (216) determines the MPM (step Sw_13). The MPM consists of, for example, six intra prediction modes. Then, the intra prediction unit (216) determines an intra prediction mode indicated by the information obtained in step Sw_12 among the plurality of intra prediction modes included in the MPM (step Sw_14).
[0606] Meanwhile, if the intra prediction unit (216) determines that an MPM flag indicating 1 does not exist in the stream in step Sw_11 (No in step Sw_11), it obtains information indicating an intra prediction mode selected in the encoding device (100) (step Sw_15). That is, the intra prediction unit (216) obtains information indicating an intra prediction mode selected in the encoding device (100) from the entropy decoding unit (202) among one or more intra prediction modes not included in the MPM. In addition, that information is decoded by the entropy decoding unit (202) and output to the intra prediction unit (216). Then, the intra prediction unit (216) determines an intra prediction mode indicated by the information obtained in step Sw_15 among one or more intra prediction modes not included in the MPM (step Sw_17).
[0607] The intra prediction unit (216) generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0608] [Inter Prediction Department]
[0609] The inter prediction unit (218) predicts a current block by referring to a reference picture stored in the frame memory (214). The prediction is performed in units of a current block or a sub-block within the current block. Additionally, a sub-block is included in a block and is a unit smaller than a block. The size of the sub-block may be 4×4 pixels, 8×8 pixels, or any other size. The size of the sub-block may be converted into units such as slices, bricks, or pictures.
[0610] For example, the inter prediction unit (218) generates an inter prediction image of a current block or sub-block by performing motion compensation using motion information (e.g., MV) decoded from a stream (e.g., prediction parameter output from the entropy decoding unit (202)), and outputs the inter prediction image to the prediction control unit (220).
[0611] When the information decoded from the stream indicates that the OBMC mode is applied, the inter prediction unit (218) generates an inter prediction image by using not only the motion information of the current block obtained by motion search, but also the motion information of the adjacent block.
[0612] Also, when the information decoded from the stream indicates that the FRUC mode is applied, the inter prediction unit (218) derives motion information by performing motion search according to the pattern matching method (bilateral matching or template matching) decoded from the stream. Then, the inter prediction unit (218) performs motion compensation (prediction) using the derived motion information.
[0613] Additionally, when the BIO mode is applied, the inter prediction unit (218) derives the MV based on a model assuming uniform linear motion. Also, when the information decoded from the stream indicates that the affine mode is applied, the inter prediction unit (218) derives the MV in sub-block units based on the MVs of multiple adjacent blocks.
[0614] [Flow for Deriving MV]
[0615] FIG. 82 is a flowchart showing an example of MV derivation in a decoding device (200).
[0616] The inter prediction unit (218) determines whether to decode motion information (e.g., MV), for example. For example, the inter prediction unit (218) may determine based on a prediction mode included in the stream, or may determine based on other information included in the stream. Here, if the inter prediction unit (218) determines that motion information is to be decoded, it derives the MV of the current block in a mode for decoding motion information. On the other hand, if the inter prediction unit (218) determines that motion information is not to be decoded, it derives the MV in a mode for not decoding motion information.
[0617] Here, the modes for deriving MV include the normal inter mode, normal merge mode, FRUC mode, and affine mode described later. Among these modes, the modes for decoding motion information include the normal inter mode, normal merge mode, and affine mode (specifically, affine inter mode and affine merge mode). In addition, the motion information may include not only MV but also the predicted MV selection information described later. Also, the modes that do not decode motion information include the FRUC mode. The inter prediction unit (218) selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.
[0618] FIG. 83 is a flowchart showing another example of MV derivation in a decoding device (200).
[0619] The inter prediction unit (218), for example, determines whether to decode the differential MV, and the inter prediction unit (218) may determine this based on the prediction mode included in the stream or based on other information included in the stream. Here, if the inter prediction unit (218) determines that the differential MV is to be decoded, it may derive the MV of the current block in the mode for decoding the differential MV. In this case, for example, the differential MV included in the stream is decoded as a prediction parameter.
[0620] Meanwhile, if the inter prediction unit (218) determines that the differential MV is not to be decoded, it derives the MV in a mode where the differential MV is not decoded. In this case, the encoded differential MV is not included in the stream.
[0621] Here, as described above, the modes for deriving MV include the normal inter, normal merge mode, FRUC mode, and affine mode described later. Among these modes, the modes for encoding differential MV include the normal inter mode and affine mode (specifically, affine inter mode). Also, the modes for not encoding differential MV include the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode). The inter prediction unit (218) selects a mode for deriving the MV of the current block from these multiple modes and uses the selected mode to derive the MV of the current block.
[0622] [MV Derivation > Normal Inter Mode]
[0623] For example, if the information decoded from the stream indicates that a normal inter mode is applied, the inter prediction unit (218) derives an MV in a normal merge mode based on the information decoded from the stream, and performs motion compensation (prediction) using the MV.
[0624] FIG. 84 is a flowchart showing an example of inter prediction by normal inter mode in a decoder (200).
[0625] The inter prediction unit (218) of the decoding device (200) performs motion compensation for each block. At this time, the inter prediction unit (218) first obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of decoding completed blocks located around the current block in time or space (step Sg_11). That is, the inter prediction unit (218) creates a list of candidate MVs.
[0626] Next, the inter prediction unit (218) selects each of the N candidate MVs (where N is an integer greater than or equal to 2) among the multiple candidate MVs obtained in step Sg_11 as a predicted motion vector candidate (also called a predicted MV candidate) and extracts them according to a predetermined priority (step Sg_12). In addition, the priority is predetermined for each of the N predicted MV candidates.
[0627] Next, the inter prediction unit (218) decodes prediction MV selection information from the input stream and, using the decoded prediction MV selection information, selects one prediction MV candidate from among the N prediction MV candidates as the prediction MV of the current block (step Sg_13).
[0628] Next, the inter prediction unit (218) decodes the differential MV from the input stream and derives the MV of the current block by adding the differential value, which is the decoded differential MV, and the selected prediction MV (step Sg_14).
[0629] Finally, the inter prediction unit (218) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processing of steps Sg_11 to Sg_15 is executed for each block. For example, if the processing of steps Sg_11 to Sg_15 is executed for each of all blocks included in the slice, the inter prediction using normal inter mode for that slice is terminated. Also, if the processing of steps Sg_11 to Sg_15 is executed for each of all blocks included in the picture, the inter prediction using normal inter mode for that picture is terminated. Additionally, if the processing of steps Sg_11 to Sg_15 is not executed for all blocks included in the slice but is executed for some blocks, the inter prediction using normal inter mode for that slice may be terminated. If the processing of steps Sg_11 to Sg_15 is executed for some blocks included in the picture, the inter prediction using normal inter mode for that picture may be terminated.
[0630] [MV Derivation > Normal Merge Mode]
[0631] For example, if the information decoded from the stream indicates the application of normal merge mode, the inter prediction unit (218) derives an MV in normal merge mode and performs motion compensation (prediction) using the MV.
[0632] FIG. 85 is a flowchart showing an example of inter prediction by normal merge mode in a decoder (200).
[0633] The inter prediction unit (218) first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded completed blocks located around the current block in time or space (step Sh_11). That is, the inter prediction unit (218) creates a list of candidate MVs.
[0634] Next, the inter prediction unit (218) derives the MV of the current block by selecting one candidate MV from among the multiple candidate MVs obtained in step Sh_11 (step Sh_12). Specifically, the inter prediction unit (218) obtains, for example, MV selection information included as a prediction parameter in a stream, and selects the candidate MV identified by the MV selection information as the MV of the current block.
[0635] Finally, the inter prediction unit (218) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The processing of steps Sh_11 to Sh_13 is executed for each block, for example. For example, when the processing of steps Sh_11 to Sh_13 is executed for each of all blocks included in the slice, the inter prediction using normal merge mode for that slice is terminated. Also, when the processing of steps Sh_11 to Sh_13 is executed for each of all blocks included in the picture, the inter prediction using normal merge mode for that picture is terminated. Additionally, when the processing of steps Sh_11 to Sh_13 is not executed for all blocks included in the slice but is executed for some blocks, the inter prediction using normal merge mode for that slice may be terminated. If the processing of steps Sh_11~Sh_13 is executed for some blocks included in the picture, the inter prediction using normal merge mode for that picture may be terminated.
[0636] [MV Derivation > FRUC Mode]
[0637] For example, if the information decoded from the stream indicates the application of FRUC mode, the inter prediction unit (218) derives the MV in FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is not signaled from the encoding device (100) side but is derived from the decoding device (200) side. For example, the decoding device (200) may derive motion information by performing motion search. In this case, the decoding device (200) performs motion search without using the pixel value of the current block.
[0638] FIG. 86 is a flowchart showing an example of inter prediction by FRUC mode in a decoder (200).
[0639] First, the inter prediction unit (218) generates a list (i.e., a candidate MV list, which may be common to the candidate MV list of normal merge mode) by referencing the MV of each decoding completed block that is spatially or temporally adjacent to the current block and representing those MVs as candidate MVs (step Si_11). Next, the inter prediction unit (218) selects the best candidate MV from among the multiple candidate MVs registered in the candidate MV list (step Si_12). For example, the inter prediction unit (218) calculates an evaluation value for each candidate MV included in the candidate MV list and selects one candidate MV as the best candidate MV based on the evaluation value. Then, the inter prediction unit (218) derives an MV for the current block based on the selected best candidate MV (step Si_14). Specifically, for example, the selected best candidate MV is derived as is as the MV for the current block. For example, an MV for the current block may be derived by performing pattern matching on the surrounding area of a location within the reference picture corresponding to the selected best candidate MV. That is, a search using pattern matching and evaluation values in the reference picture is performed on the surrounding area of the best candidate MV, and if there is an MV with a good evaluation value, the best candidate MV may be updated to that MV and used as the final MV for the current block. It is not necessary to update to an MV with a better evaluation value.
[0640] Finally, the inter prediction unit (218) generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). The processing of steps Si_11 to Si_15 is executed for each block, for example. For example, when the processing of steps Si_11 to Si_15 is executed for each of all blocks included in a slice, the inter prediction using FRUC mode for that slice is terminated. Also, when the processing of steps Si_11 to Si_15 is executed for each of all blocks included in a picture, the inter prediction using FRUC mode for that picture is terminated. The processing may also be performed in the same way as the block unit described above at the sub-block level.
[0641] [MV Derivation > Affine Merge Mode]
[0642] For example, if the information decoded from the stream indicates the application of an affine merge mode, the inter prediction unit (218) derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0643] FIG. 87 is a flowchart showing an example of inter-prediction by an affine merge mode in a decoder (200).
[0644] In the affine merge mode, first, the inter prediction unit (218) derives the MV of each control point of the current block (step Sk_11). The control points are the points at the top left corner and the top right corner of the current block as shown in FIG. 46a, or the points at the top left corner, the top right corner, and the bottom left corner of the current block as shown in FIG. 46b.
[0645] For example, when using the method of deriving MV shown in FIG. 47a to FIG. 47c, the inter prediction unit (218) examines these blocks in the order of decoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), as shown in FIG. 47a, and identifies the first valid block decoded in affine mode.
[0646] The inter prediction unit (218) derives the MV of the control point using the first valid block decoded in a specified affine mode. For example, if block A is specified and block A has two control points, as shown in FIG. 47b, the inter prediction unit (218) calculates the motion vector v0 of the left-top corner control point and the motion vector v1 of the right-top corner control point of the current block by projecting the motion vectors v3 and v4 of the left-top corner and right-top corner of the decoded block containing block A onto the current block. By doing so, the MV of each control point is derived.
[0647] In addition, as shown in FIG. 49a, when block A is specified and block A has two control points, the MV of three control points may be calculated, and as shown in FIG. 49b, when block A is specified and block A has three control points, the MV of two control points may be calculated.
[0648] Also, if the stream includes MV selection information as a prediction parameter, the inter prediction unit (218) may use the MV selection information to derive the MV of each control point of the current block.
[0649] Next, the inter prediction unit (218) performs motion compensation for each of the plurality of sub-blocks included in the current block. That is, for each of the plurality of sub-blocks, the inter prediction unit (218) calculates the MV of the sub-block as an affine MV by using two motion vectors v0 and v1 and the above-described equation (1A), or by using three motion vectors v0, v1, and v2 and the above-described equation (1B) (step Sk_12). Then, the inter prediction unit (218) performs motion compensation for the sub-block using the affine MV and the decoded reference picture (step Sk_13). When the processing of steps Sk_12 and Sk_13 is executed for each of the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block is terminated. That is, motion compensation is performed for the current block, and a predicted image of the current block is generated.
[0650] Additionally, in step Sk_11, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be any combination of the MV derivation methods shown in FIGS. 47a to 47c, the MV derivation methods shown in FIGS. 48a and 48b, the MV derivation methods shown in FIGS. 49a and 49b, and other MV derivation methods.
[0651] In addition, the candidate MV list may include candidate MVs of modes that perform predictions on a sub-block basis, other than affine mode.
[0652] In addition, as a candidate MV list, for example, a candidate MV list including a candidate MV of an affine merge mode having 2 control points and a candidate MV of an affine merge mode having 3 control points may be generated. Alternatively, a candidate MV list including a candidate MV of an affine merge mode having 2 control points and a candidate MV list including a candidate MV of an affine merge mode having 3 control points may be generated, respectively. Alternatively, a candidate MV list including a candidate MV of one of the modes having 2 control points and an affine merge mode having 3 control points may be generated.
[0653] [MV Derivation > Affine Intermode]
[0654] For example, if the information decoded from the stream indicates the application of an affine inter mode, the inter prediction unit (218) derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.
[0655] FIG. 88 is a flowchart showing an example of inter prediction by an affine inter mode in a decoder (200).
[0656] In the affine inter mode, first, the inter prediction unit (218) derives the predicted MV(v0, v1) or (v0, v1, v2) for each of the two or three control points of the current block (step Sj_11). The control points are points at the top left corner, top right corner, or bottom left corner of the current block, for example, as shown in FIG. 46a or FIG. 46b.
[0657] The inter prediction unit (218) obtains prediction MV selection information included as a prediction parameter in the stream and derives the prediction MV of each control point of the current block using the MV identified by the prediction MV selection information. For example, when using the method of deriving MV shown in FIG. 48a and FIG. 48b, the inter prediction unit (218) derives the prediction MV (v0, v1) or (v0, v1, v2) of the control point of the current block by selecting the MV of the block identified by the prediction MV selection information among the decoded completed blocks near each control point of the current block shown in FIG. 48a or FIG. 48b.
[0658] Next, the inter-prediction unit (218) obtains, for example, each differential MV included as a prediction parameter in the stream, and adds the predicted MV of each control point of the current block and the differential MV corresponding to the predicted MV (step Sj_12). By doing so, the MV of each control point of the current block is derived.
[0659] Next, the inter prediction unit (218) performs motion compensation for each of the plurality of sub-blocks included in the current block. That is, for each of the plurality of sub-blocks, the inter prediction unit (218) calculates the MV of the sub-block as an affine MV by using two motion vectors v0 and v1 and the above-described equation (1A), or by using three motion vectors v0, v1, and v2 and the above-described equation (1B) (step Sj_13). Then, the inter prediction unit (218) performs motion compensation for the sub-block using the affine MV and the decoded reference picture (step Sj_14). When the processing of steps Sj_13 and Sj_14 is executed for each of the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block is terminated. That is, motion compensation is performed for the current block, and a predicted image of the current block is generated.
[0660] In addition, in step Sj_11, the aforementioned candidate MV list may be generated, just as in step Sk_11.
[0661] [MV Derivation > Triangle Mode]
[0662] For example, if the information decoded from the stream indicates the application of triangle mode, the inter prediction unit (218) derives the MV in triangle mode and performs motion compensation (prediction) using the MV.
[0663] FIG. 89 is a flowchart showing an example of inter-prediction by triangle mode in a decoder (200).
[0664] In triangle mode, first, the inter prediction unit (218) divides the current block into a first partition and a second partition (step Sx_11). At this time, the inter prediction unit (218) may obtain partition information, which is information regarding the division into each partition, from the stream as a prediction parameter. Then, the inter prediction unit (218) may divide the current block into a first partition and a second partition according to the partition information.
[0665] Next, the inter prediction unit (218) first obtains a plurality of candidate MVs for the current block based on information such as MVs of a plurality of decoded completed blocks located around the current block in time or space (step Sx_12). That is, the inter prediction unit (218) creates a list of candidate MVs.
[0666] Then, the inter prediction unit (218) selects the candidate MV of the first partition and the candidate MV of the second partition from among the plurality of candidate MVs obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, the inter prediction unit (218) may obtain MV selection information for identifying the selected candidate MVs from the stream as prediction parameters. Then, the inter prediction unit (218) may select the first MV and the second MV according to the MV selection information.
[0667] Next, the inter prediction unit (218) generates a first predicted image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Likewise, the inter prediction unit (218) generates a second predicted image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).
[0668] Finally, the inter prediction unit (218) generates a prediction image of the current block by adding the first prediction image and the second prediction image with weights (step Sx_16).
[0669] [Movement Exploration > DMVR]
[0670] For example, if the information decoded from the stream indicates the application of DMVR, the inter prediction unit (218) performs motion detection with DMVR.
[0671] FIG. 90 is a flowchart showing an example of motion detection by DMVR in a decoder (200).
[0672] The inter prediction unit (218) first derives the MV of the current block in merge mode (step Sl_11). Next, the inter prediction unit (218) derives the final MV for the current block by exploring the surrounding area of the reference picture represented by the MV derived in step Sl_11 (step Sl_12). That is, the MV of the current block is determined by the DMVR.
[0673] FIG. 91 is a flowchart showing a detailed example of motion detection by DMVR in a decoder (200).
[0674] First, the inter prediction unit (218) calculates the cost at the search location (also called the starting point) indicated by the initial MV and eight search locations surrounding it in Step 1 shown in FIG. 58a. Then, the inter prediction unit (218) determines whether the cost of the search locations other than the starting point is the minimum. Here, if the inter prediction unit (218) determines that the cost of the search locations other than the starting point is the minimum, it moves to the search location where the cost is the minimum and performs the processing of Step 2 shown in FIG. 58a. On the other hand, if the cost of the starting point is the minimum, the inter prediction unit (218) skips the processing of Step 2 shown in FIG. 58a and performs the processing of Step 3.
[0675] In Step 2 shown in FIG. 58a, the inter prediction unit (218) sets the search location moved according to the processing result of Step 1 as a new starting point and performs the same search as the processing of Step 1. Then, the inter prediction unit (218) determines whether the cost of the search location other than the starting point is the minimum. Here, if the cost of the search location other than the starting point is the minimum, the inter prediction unit (218) performs the processing of Step 4. Meanwhile, if the cost of the starting point is the minimum, the inter prediction unit (218) performs the processing of Step 3.
[0676] In Step 4, the inter prediction unit (218) treats the search position of the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as a difference vector.
[0677] In Step 3 shown in FIG. 58a, the inter prediction unit (218) determines a pixel location with fractional precision where the cost is minimized based on the cost at four points located above, below, left, and right of the starting point of Step 1 or Step 2, and sets that pixel location as the final search location. That pixel location with fractional precision is determined by weighted addition of the vectors of four points located above, below, left, and right ((0, 1), (0, -1), (-1, 0), (1, 0)), using the cost at each of the four search locations as a weight. Then, the inter prediction unit (218) determines the difference between the location indicated by the initial MV and the final search location as a difference vector.
[0678] [Movement Compensation > BIO / OBMC / LIC]
[0679] For example, if the information decoded from the stream indicates the application of correction to the predicted image, the inter-prediction unit (218) generates the predicted image and corrects the predicted image according to the mode of correction. The mode is, for example, the aforementioned BIO, OBMC, and LIC.
[0680] FIG. 92 is a flowchart showing an example of the generation of a predicted image in a decoding device (200).
[0681] The inter prediction unit (218) generates a prediction image (step Sm_11) and corrects the prediction image by any one of the modes described above (step Sm_12).
[0682] FIG. 93 is a flowchart showing another example of the generation of a predicted image in a decoding device (200).
[0683] The inter prediction unit (218) derives the MV of the current block (step Sn_11). Next, the inter prediction unit (218) generates a prediction image using the MV (step Sn_12) and determines whether to perform correction processing (step Sn_13). For example, the inter prediction unit (218) obtains a prediction parameter included in the stream and determines whether to perform correction processing based on the prediction parameter. This prediction parameter is, for example, a flag indicating whether to apply each mode described above. Here, if the inter prediction unit (218) determines that correction processing is to be performed (Yes in step Sn_13), it generates a final prediction image by correcting the prediction image (step Sn_14). Additionally, in the LIC, the luminance and color difference of the prediction image may be corrected in step Sn_14. Meanwhile, if the inter prediction unit (218) determines that it will not perform correction processing (No in step Sn_13), it outputs the predicted image as the final predicted image without correcting it (step Sn_15).
[0684] [Movement Reward > OBMC]
[0685] For example, if the information decoded from the stream indicates the application of OBMC, the inter prediction unit (218) generates a prediction image and corrects the prediction image according to OBMC.
[0686] FIG. 94 is a flowchart showing an example of correction of a predicted image by OBMC in a decoder (200). In addition, the flowchart of FIG. 94 shows the flow of correction of a predicted image using the current picture and reference picture shown in FIG. 62.
[0687] First, the inter prediction unit (218) acquires a prediction image (Pred) by normal motion compensation using the MV assigned to the current block as shown in FIG. 62.
[0688] Next, the inter prediction unit (218) obtains a predicted image (Pred_L) by applying (reusing) the MV (MV_L) already derived for the decoding-completed left adjacent block to the current block. Then, the inter prediction unit (218) performs a first correction of the predicted image by overlapping the two predicted images, Pred and Pred_L. This has the effect of mixing the boundaries between adjacent blocks.
[0689] Likewise, the inter prediction unit (218) obtains a prediction image (Pred_U) by applying (reusing) the MV (MV_U) already derived for the adjacent block above the decoding completion to the current block. Then, the inter prediction unit (218) performs a second correction of the prediction image by superimposing the prediction image Pred_U onto the prediction image (e.g., Pred and Pred_L) that has undergone a first correction. This has the effect of blending the boundaries between adjacent blocks. The prediction image obtained by the second correction is the final prediction image of the current block in which the boundaries with the adjacent blocks are blended (smoothed).
[0690] [Movement Compensation > BIO]
[0691] For example, if the information decoded from the stream indicates the application of BIO, the inter-prediction unit (218) generates a prediction image and corrects the prediction image according to the BIO.
[0692] FIG. 95 is a flowchart showing an example of correction of a predicted image by BIO in a decoding device (200).
[0693] The inter prediction unit (218) derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) that are different from the picture (Cur Pic) containing the current block, as shown in FIG. 63. Then, the inter prediction unit (218) derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_11). Additionally, motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0694] Next, the inter prediction unit (218) uses the motion vector M0 and reference picture L0 to interpolate the image I of the current block. 0 It derives the inter prediction unit (218), using motion vector M1 and reference picture L1, the interpolated image I of the current block. 1 Derive (Step Sy_12). Here, the interpolated image I 0 is an image included in reference picture Ref0, derived for the current block, and interpolated image I 1 is an image included in reference picture Ref1, derived for the current block. Interpolated image I 0 and interpolated image I 1 Each may be the same size as the current block. Or, interpolated image I 0 and interpolated image I 1 Each of these may be an image larger than the current block in order to properly derive the gradient image described later. In addition, the interpolated image (I 0 and I 1 ) may include motion vectors (M0, M1) and reference pictures (L0, L1), and a predicted image derived by applying a motion compensation filter.
[0695] Also, the inter prediction unit (218) is an interpolated image I 0 and interpolated image I 1From, the gradient image of the current block (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) derives (step Sy_13). Also, the horizontal gradient image is, (Ix 0 , Ix 1 ) and the vertical gradient image is, (Iy 0 , Iy 1 ) is. The inter prediction unit (218) may derive the gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may represent the amount of spatial change of pixel values according to the horizontal or vertical direction.
[0696] Next, the inter prediction unit (218) is an interpolated image (I) in units of multiple sub-blocks constituting the current block. 0 , I 1 ) and gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 The optical flow (vx, vy), which is the velocity vector described above, is derived using ) (step Sy_14). As an example, the sub-block may be a 4×4 pixel sub-CU.
[0697] Next, the inter prediction unit (218) corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter prediction unit (218) derives a correction value for the value of a pixel included in the current block using the optical flow (vx, vy) (step Sy_15). Then, the inter prediction unit (218) may correct the predicted image of the current block using the correction value (step Sy_16). Additionally, the correction value may be derived on a pixel-by-pixel basis, or on a multiple pixel-by-pixel basis or a sub-block basis.
[0698] In addition, the processing flow of BIO is not limited to the processing disclosed in FIG. 95. Only some of the processing disclosed in FIG. 95 may be performed, different processing may be added or replaced, and different processing order may be executed.
[0699] [Movement Compensation > LIC]
[0700] For example, if the information decoded from the stream indicates the application of LIC, the inter prediction unit (218) generates a prediction image and corrects the prediction image according to the LIC.
[0701] FIG. 96 is a flowchart showing an example of correction of a predicted image by LIC in a decoding device (200).
[0702] First, the inter prediction unit (218) uses MV to obtain a reference image corresponding to the current block from the decoded reference picture (step Sz_11).
[0703] Next, the inter prediction unit (218) extracts information indicating how the luminance value has changed in the reference picture and the current picture for the current block (step Sz_12). This extraction is performed based on the luminance pixel values of the decoding completed left adjacent reference area (peripheral reference area) and the decoding completed upper adjacent reference area (peripheral reference area) in the current picture, as shown in FIG. 66a, and the luminance pixel values at the equivalent positions within the reference picture designated as the derived MV. Then, the inter prediction unit (218) calculates a luminance correction parameter using the information indicating how the luminance value has changed (step Sz_13).
[0704] The inter prediction unit (218) generates a prediction image for the current block by performing a luminance correction process that applies the luminance correction parameter to the reference image within the reference picture designated as MV (step Sz_14). That is, a correction based on the luminance correction parameter is performed on the prediction image, which is the reference image within the reference picture designated as MV. In this correction, the luminance may be corrected or the color difference may be corrected.
[0705] [Predictive Control Unit]
[0706] The prediction control unit (220) selects either an intra prediction image or an inter prediction image and outputs the selected prediction image to the adder (208). Overall, the configuration, function, and processing of the prediction control unit (220), intra prediction unit (216), and inter prediction unit (218) on the decoder (200) side may correspond to the configuration, function, and processing of the prediction control unit (128), intra prediction unit (124), and inter prediction unit (126) on the encoding device (100) side.
[0707] [Explanation of Intra-Block Copy (IBC) Mode]
[0708] In the inter prediction unit (126) (see FIG. 7) of the encoding device (100) in the first embodiment, in addition to the pictures preceding and succeeding in time, the decoded pixels (i.e., pixels that have already been encoded and decoded) in the picture containing the block to be processed (so-called current picture) may be referenced. For example, the inter prediction unit (126) may only reference pixels within the screen. This prediction mode is called Intra Block Copy (IBC). The IBC mode may be used as one of the modes of normal inter prediction, and may be used together with, for example, a merge mode, a normal inter mode, and a skip mode. The IBC mode will be explained in more detail below with reference to the drawings.
[0709] FIG. 97 is a diagram illustrating the IBC mode. FIG. 97 (a) is a diagram showing an example of a region that can be referenced in the IBC mode. FIG. 97 (b) is a diagram showing an example of a block vector (BV).
[0710] In the inter prediction processing of the block to be processed, the inter prediction unit (126) refers to a block that has already been encoded and decoded within a picture different from the picture containing the block to be processed (i.e., the current picture). However, in the prediction processing using the IBC mode, the inter prediction unit (126) refers only to a pixel that has already been encoded and decoded within the same picture as the picture containing the block to be processed, just as in the intra prediction processing. For example, as shown in FIG. 97 (a), the referenceable pixel within the screen is a pixel located above or to the left of the block to be processed. The area containing these pixels is called the "referenceable area."
[0711] Additionally, the encoding device (100) may determine a vector representing the displacement of the position of the block to be processed and the reference block in order to identify the reference block within a frame or picture (i.e., within a screen) that includes the block to be processed. This vector is called a block vector (BV). A method for representing the BV will be described below.
[0712] As shown in FIG. 97(b), the block vector includes, for example, an x component and a y component. The x component (BVx) represents the horizontal displacement between the block to be processed and the reference block within the screen. The y component (BVy) represents the vertical displacement between the block to be processed and the reference block within the screen. For example, if the origin (i.e., the starting point) of the BV is the pixel location of the top-left corner of the block to be processed, the end point of the BV is the pixel location of the top-left corner of the reference block. In this case, the BV becomes a vector connecting the top-left corner of the block to be processed and the top-left corner of the reference block. The encoding device (100) signals the BV in the encoding bit stream so that the reference block selected by the encoding device (100) can be identified when the decoder (200) decodes the encoding bit stream.
[0713] [HMVP Mode]
[0714] Next, regarding the HMVP mode, it will be explained in more detail with reference to FIG. 42. The candidate MV list in the figure (hereinafter also referred to as the motion vector candidate list) may be a candidate MV list for merge mode or a candidate MV list for normal inter mode.
[0715] In merge mode and normal inter mode, one candidate MV (hereinafter also referred to as a prediction candidate or motion vector candidate) is selected from a list of candidate MVs generated by referencing a completed block, and the motion vector (MV) of the block to be processed is determined. For example, when the encoding device (100) generates a prediction image of the block to be processed in normal inter mode, it selects one candidate MV from the list of candidate MVs, takes the difference between the MV of the block to be processed and the candidate MV, and encodes the difference and the index of the candidate MV. Also, for example, when the encoding device (100) generates a prediction image of the block to be processed in merge mode, it selects one merge index from the list of merge candidate MVs, thereby selecting a candidate MV and a reference picture index corresponding to the merge index. The encoding device (100) encodes the selected merge index. In the candidate MV list, there are registered motion vectors such as spatial proximity prediction motion vectors (also called spatial proximity candidate MVs), which are motion vectors of multiple encoded blocks located spatially around the target block, and temporal proximity prediction motion vectors (also called temporal proximity candidate MVs), which are motion vectors of nearby blocks projected from the position of the target block in the encoded reference picture. Among the candidate MVs registered in this candidate MV list, there is a candidate MV of HMVP mode (hereinafter also called HMVP motion vector candidate).
[0716] In HMVP mode, candidate MVs are managed using a FIFO buffer for HMVP (hereinafter also referred to as the HMVP table), separate from the candidate MV lists of merge mode and normal inter mode. In the FIFO buffer, a predetermined number (e.g., 5 in FIG. 42) of prediction candidates (also referred to as candidate MVs) having information regarding the MV of a previously processed block (i.e., a block processed before the target block) (hereinafter referred to as MV information) are stored in order of block processing, starting from the newest. For example, as shown in FIG. 42, whenever the encoding device (100) processes one block, it stores a prediction candidate having the MV information of the newest block (in other words, the block processed immediately before the block) in the FIFO buffer. Additionally, whenever the encoding device (100) finishes processing one block, it may store a prediction candidate having the MV information of the block in the FIFO buffer.
[0717] Additionally, the block vector candidate list for IBC mode is generated using BV instead of MV, and block vector candidates (candidate BV) are stored in the HMVP table instead of candidate MV. Hereinafter, MV and BV are referred to as vectors, candidate MV and candidate BV are referred to as vector candidates, and the motion vector candidate list and block vector candidate list are referred to as vector candidate lists. Candidate MV and MV may be appropriately replaced with BV and BV candidates, and candidate BV and BV may be appropriately replaced with MV and MV candidates. Additionally, BV candidates and MV candidates may be mixed and registered in the motion vector candidate list and block vector candidate list.
[0718] Additionally, when the encoding device (100) stores a prediction candidate having new MV information in the FIFO buffer, it deletes the prediction candidate having MV information of the oldest block (in other words, the first processed block) in the FIFO buffer from the FIFO buffer. By doing so, the encoding device (100) can manage the prediction candidate in the FIFO buffer in the latest state. Also, in the example of FIG. 42, HMVP1 in the FIFO buffer is a prediction candidate having MV information of the newest block, and HMVP5 in the FIFO buffer is a prediction candidate having MV information of the oldest block.
[0719] Next, with reference to FIG. 42, the pruning process in the process of registering candidate MVs of HMVP mode to the candidate MV lists for merge mode and normal inter mode, and the process of updating the FIFO buffer will be explained.
[0720] [Pruning Treatment]
[0721] Pruning processing refers to a process of comparing an MV that is to be registered in a list, etc., with an MV that is already registered in a list, etc., to determine whether these MV values (or MV information) match. In HMVP mode, there are two types of pruning processing: (1) pruning processing for registering a candidate MV from the HMVP table into a list of candidate MVs (this is referred to as the first pruning processing), and (2) pruning processing for registering the MV of the block processed immediately before (i.e., the latest MV) into the HMVP table (i.e., updating the HMVP table) (this is referred to as the second pruning processing).
[0722] Referring to FIG. 42, for example, in the first pruning process of (1), the encoding device (100) determines whether the MV information is different from all candidate MVs already registered in the candidate MV list for each of the plurality of candidate MVs (e.g., HMVP1 to HMVP5) in the FIFO buffer. More specifically, the encoding device (100) determines whether the MV information is different from all prediction candidates (so-called candidate MVs) already registered in the candidate MV list in order, starting from the candidate MV (here, HMVP1) having the MV information of the newest block for the candidate MVs (in this example, HMVP1 to HMVP5) in the FIFO buffer. If the encoding device (100) determines that the MV information of HMVP1 is different from all candidate MVs in the candidate MV list, the encoding device (100) registers HMVP1 in the candidate MV list. And, the encoding device (100) performs the same processing for each of the remaining candidate MVs (here, HMVP2 to HMVP5) in the FIFO buffer. In addition, the HMVP motion vector candidates registered in the candidate MV list from the FIFO buffer may be one or multiple.
[0723] By using the HMVP mode in this way, it becomes possible to register in the candidate MV list not only candidate MVs having MV information of blocks spatially or temporally adjacent to the block to be processed, but also candidate MVs having MV information of blocks processed in the past. As a result, the variation of candidate MVs for merge mode and normal inter mode increases. Accordingly, the encoding device (100) can select a more appropriate candidate MV for the block to be processed, and thus the encoding efficiency is improved.
[0724] Also, for example, in the second pruning process of (2) above, the encoding device (100) compares the MV of the block processed immediately before (let's call this HMVP0) with all candidate MVs already registered in the FIFO buffer (HMVP1 to HMVP5 in the drawing) and determines whether the MV information of HMVP0 is different from the MV information of all candidate MVs. If the encoding device (100) determines that the MV information of HMVP0 is different from the MV information of all candidate MVs in the FIFO buffer, it deletes the MV of the oldest block in the FIFO buffer (in this example, HMVP5) and registers HMVP0 as the MV of the newest block in the FIFO buffer. Meanwhile, if the encoding device (100) determines that the MV information of HMVP0 matches the MV information of any one of the candidate MVs in the FIFO buffer, it does not register HMVP0 in the FIFO buffer and maintains the current FIFO buffer (i.e., the FIFO buffer in which HMVP1 to HMVP5 are registered).
[0725] In addition, the MV information may include not only the value of the MV, but also information such as the referenced picture, the referenced direction, and the number of referenced pictures.
[0726] Here, candidate MV lists for merge mode and normal inter mode have been described, but are not limited thereto. For example, in the prediction processing of IBC mode, the encoding device (100) generates a candidate BV list in which multiple candidate BVs are registered for each block to be processed.
[0727] In addition, the candidate MV list and FIFO buffer shown in FIG. 42 are examples, and may be lists and buffers of different sizes than the city, and may be configured to register candidate MVs in a different order than the city.
[0728] In addition, the encoding device (100) has been described as an example here, and the above processing is common to the encoding device (100) and the decoding device (200).
[0729] [First Mode]
[0730] Next, the encoding device (100) and the decoding device (200) according to the first embodiment will be described.
[0731] The encoding device (100) uses an IBC mode that references the encoding completion area of the picture to which the processing target block belongs in the generation of a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, it generates a vector candidate list by registering an HMVP vector candidate from the HMVP table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, it generates a vector candidate list by performing a first pruning process and registering an HMVP vector candidate from the HMVP table to the vector candidate list, and the HMVP table stores a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, as HMVP vector candidates in a FIFO manner, and encodes the processing target block using the vector candidate list.
[0732] Additionally, the decoding device (200) uses an IBC mode that references the encoding completion area of the picture to which the processing target block belongs in the generation of a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, it generates a vector candidate list by registering an HMVP vector candidate from the HMVP table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, it generates a vector candidate list by performing a first pruning process and registering an HMVP vector candidate from the HMVP table to the vector candidate list, and the HMVP table stores a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, as HMVP vector candidates in a FIFO manner, and decodes the processing target block using the vector candidate list.
[0733] FIG. 98 is a flowchart (1000) showing an example of an operation performed by an encoding device (100) and a decoding device (200) according to a first embodiment.
[0734] The encoding device (100) initiates block-unit loop processing in the prediction processing of the picture to be processed (not shown). First, the encoding device (100) determines the size of the block to be processed in IBC mode (S1001). IBC mode is a prediction mode that refers to the processing completion area of the picture to which the block to be processed belongs in the generation of the prediction image of the block to be processed.
[0735] Next, the encoding device (100) determines whether the size of the block to be processed is less than or equal to a threshold (S1002). For example, the size of the block to be processed may be defined by the number of pixels within the block to be processed. In this case, the threshold is, for example, 16 pixels. Also, for example, the size of the block to be processed may be defined by at least one of the width and height of the block to be processed. In this case, the threshold is, for example, 4×4 pixels.
[0736] Next, when the encoding device (100) determines that the size of the block to be processed is not below the threshold, in other words, that it is larger than the threshold (No in S1003), it performs a first pruning process and registers the HMVP block vector candidate in the block vector candidate list, thereby generating a block vector candidate list (S1004). For example, when the encoding device (100) generates a block vector candidate list for IBC mode, in step S1004, the encoding device (100) performs a first pruning process on each HMVP block vector candidate from the last HMVP block vector candidate entered into the HMVP table to the first HMVP block vector candidate entered, in order from the newest. More specifically, the encoding device (100) compares each HMVP block vector candidate with all block vector candidates registered in the current block vector candidate list table, in order from the last HMVP block vector candidate entered into the HMVP table, and determines whether they are different from all block vector candidates. For example, a vector (MV or BV) used in an adjacent block may be registered in the block vector candidate list. For example, if a block vector candidate matching the last HMVP block vector candidate entered into the HMVP table exists in the block vector candidate list, the encoding device (100) does not add the said HMVP block vector candidate to the block vector candidate list. On the other hand, if a block vector candidate matching the last HMVP mode block vector candidate entered into the HMVP table does not exist in the block vector candidate list, the encoding device (100) adds the said HMVP mode block vector candidate to the block vector candidate list. The same processing is performed for other HMVP block vector candidates stored in the HMVP table. The number of HMVP block vector candidates added to the block vector candidate list may be one, or multiple depending on the number of empty slots in the list.
[0737] Meanwhile, when the encoding device (100) determines that the size of the block to be processed is below a threshold (Yes in S1003), it generates a block vector candidate list by registering the HMVP block vector candidate in the block vector candidate list without performing the first pruning process (S1005). That is, in step S1005, the HMVP block vector candidates stored in the HMVP table are added to the block vector candidate list in order, starting from the newest ones. The number of HMVP block vector candidates added to the block vector candidate list may be one, or multiple depending on the number of empty slots in the list. In this way, the encoding device (100) can reduce the number of processing cycles required to generate the block vector candidate list by skipping the pruning process when the size of the block to be processed is below a threshold. By doing so, the encoding device (100) improves the encoding efficiency. Accordingly, the encoding device (100) does not need to share one block vector candidate list between the block to be processed and the block adjacent to the block to be processed in order to reduce the number of processing cycles for generating a block vector candidate list. For example, in merge mode, the encoding device (100) does not need to share a merge list between the block to be processed and the block adjacent to the block to be processed.
[0738] Referring again to FIG. 98, the encoding device (100) encodes the block to be processed using the block vector candidate list generated in steps S1004 and S1005 (S1006).
[0739] The encoding device (100) repeats the processing of steps S1001 to S1006 for all blocks in the picture to be processed, and then terminates the block-unit loop processing (not shown).
[0740] In addition, the above process is common to both the encoding device (100) and the decoding device (200), although the encoding device (100) is described as an example here.
[0741] In addition, here, the encoding device (100) determines whether to skip the pruning process based on whether the size of the block to be processed is below a threshold when generating a vector candidate list in IBC mode, and the same process may be performed when generating a vector candidate list in merge mode and normal inter mode. For example, when generating a vector candidate list (merge candidate MV list) in merge mode, the encoding device (100) determines whether to skip the pruning process based on whether the size of the block to be processed is below a threshold. In the merge candidate MV list, information of the candidate MV and the reference picture (reference picture index) is combined and registered. Therefore, even if the MV values are the same, if the reference picture information is different, it is determined in the pruning process that the vector candidates do not match.
[0742] In addition, an example has been described here in which the encoding device (100) performs prediction processing using an IBC mode. The encoding device (100) may perform prediction processing using the IBC mode when it decides to use an IBC mode among a plurality of prediction modes, and may perform prediction processing using the prediction mode when it decides to use a prediction mode different from the IBC mode among a plurality of prediction modes. For example, when the encoding device (100) uses a prediction mode different from the IBC mode, it may perform a first pruning process and generate a vector candidate list by registering HMVP vector candidates from the HMVP table into a vector candidate list, and may encode the block to be processed using the vector candidate list.
[0743] [Technical advantages of the first embodiment]
[0744] The first aspect of the present disclosure introduces a constraint based on the size of the block to be processed into the pruning process of HMVP vector candidates in the generation of a vector candidate list. Accordingly, the encoding device (100) and the decoding device (200) can skip the pruning process of comparing the HMVP vector candidates stored in the HMVP table with the vector candidates already registered in the vector candidate list when the size of the block to be processed is below a threshold during the generation of the vector candidate list. Therefore, when the size of the block to be processed is below a threshold, the number of cycles for performing the pruning process during the generation of the vector candidate list is reduced. Accordingly, the encoding device (100) and the decoding device (200) according to the first aspect reduce the throughput, thereby improving the encoding efficiency of the encoding device (100) and the processing efficiency of the decoding device (200).
[0745] In addition, in the first embodiment, an example was described in which the above size constraint is introduced in the pruning process when registering HMVP block vector candidates from the HMVP table to the block vector candidate list for the generation of the block vector candidate list for IBC mode, and the above size constraint may also be introduced in the generation of the candidate MV list for merge mode and normal inter mode.
[0746] [Second Mode]
[0747] Hereinafter, an encoding device (100), a decoding device (200), an encoding method, and a decoding method according to the second embodiment will be described. In the first embodiment, the process of generating a vector candidate list was described, and in the second embodiment, the process of updating an HMVP table will be described.
[0748] The encoding device (100) according to the second embodiment also updates the HMVP table using a second vector candidate having information regarding the second vector, and in updating the HMVP table, determines whether the size of the block to be processed is below a threshold, and if the size of the block to be processed is below the threshold, updates the HMVP table without performing the second pruning pr...
Claims
Claim 1 The circuit comprises a circuit and a memory connected to the circuit, and in operation, the circuit uses an Intra Block Copy (IBC) mode that references a processing completion area of a picture to which the processing target block belongs in generating a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including multiple vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from the HMVP (History-based Motion Vector Predictor) table into the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from the HMVP table into the vector candidate list, wherein the HMVP table stores multiple first vector candidates, each having information regarding a first vector used in a processing completion block, as HMVP vector candidates in a FIFO manner, and using the vector candidate list An encoding device that encodes the above-mentioned processing target block. Claim 2 A decoding device comprising a circuit and a memory connected to the circuit, wherein, in operation, the circuit uses an IBC mode that references a processing completion area of a picture to which the processing target block belongs in generating a predicted image of the processing target block, determines whether the size of the processing target block, which is a unit for generating a vector candidate list including a plurality of vector candidates, is below a threshold, and if the size of the processing target block is below the threshold, generates the vector candidate list by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the processing target block is greater than the threshold, generates the vector candidate list by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and decodes the processing target block using the vector candidate list. Claim 3 A encoding method for generating a predicted image of a block to be processed, wherein an IBC mode is used to reference a processing completion area of a picture to which the block to be processed belongs, and the size of the block to be processed, which is a unit for generating a vector candidate list including multiple vector candidates, is determined to be below a threshold; if the size of the block to be processed is below the threshold, the vector candidate list is generated by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the block to be processed is greater than the threshold, the vector candidate list is generated by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and the block to be processed is encoded using the vector candidate list. Claim 4 A decoding method for generating a predicted image of a block to be processed, wherein an IBC mode is used to reference a processing completion area of a picture to which the block to be processed belongs, and the size of the block to be processed, which is a unit for generating a vector candidate list including multiple vector candidates, is determined to be below a threshold; if the size of the block to be processed is below the threshold, the vector candidate list is generated by registering an HMVP vector candidate from an HMVP table to the vector candidate list without performing a first pruning process, and if the size of the block to be processed is greater than the threshold, the vector candidate list is generated by performing the first pruning process and registering an HMVP vector candidate from an HMVP table to the vector candidate list, wherein a plurality of first vector candidates, each having information regarding a first vector used in a processing completion block, are stored in the HMVP table in a FIFO manner as HMVP vector candidates, and the block to be processed is decoded using the vector candidate list. Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 delete Claim 9 delete Claim 10 delete Claim 11 delete Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 delete Claim 16 delete Claim 17 delete Claim 18 delete Claim 19 delete Claim 20 delete Claim 21 delete Claim 22 delete