Systems and methods for video coding
By introducing cross-component adaptive loop filtering and adaptive loop filtering processes in video encoding, encoding efficiency and image quality are optimized, and the problem of limited improvement in encoding efficiency and image quality in the prior art is solved, thereby realizing the reduction of resource utilization and circuit scale.
Patent Information
- Application Number
- CN202510364128.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-08
- Filing Date
- 2020-08-07
- Publication Date
- 2025-07-11
AI Technical Summary
When existing video encoding technologies process the increasing amount of digital video data, the encoding efficiency and image quality improvement are limited, and the circuit scale is large, so further optimization is needed.
Cross-component adaptive loop filtering (CCALF) and adaptive loop filtering (ALF) processes are used to generate coefficient values by copying the reconstruction samples within the virtual boundary and adding them to improve coding efficiency and image quality, and optimize the encoding process with technologies such as block segmentation, intra- and inter-frame prediction, transformation and quantization.
Improve encoding efficiency, enhance image quality, and reduce the utilization rate and circuit scale of encoding/decoding processing resources, and improve encoding/decoding speed.
Smart Images

Figure CN120302049A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the application date of August 7, 2020, titled "Systems and Methods for Video Coding", and the application number of 202080054368.X. Technical Field
[0002] The present disclosure relates to video coding, and in particular to video coding and decoding systems, components and methods in video coding and decoding, such as for performing the CCALF (Cross Component Adaptive Loop Filtering) process. Background Art
[0003] With the progress of video coding technology, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec), there is still a continuous need to improve and optimize video coding technology to handle the increasing amount of digital video data in various applications. The present disclosure relates to further progress, improvement, and optimization in video coding, particularly in the CCALF (Cross Component Adaptive Loop Filtering) process. Summary of the Invention
[0004] According to one aspect, an encoder is provided that includes circuitry and a memory coupled to the circuitry. In response to a first reconstructed image sample being outside a virtual boundary, the circuitry copies a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample. The circuitry generates a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component. The circuitry generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of the chrominance component. The circuitry generates a third coefficient value by adding the first coefficient value and the second coefficient value, and encodes a third reconstructed image sample of the chrominance component using the third coefficient value.
[0005] According to another aspect, the first reconstructed image sample and the second reconstructed image sample are located adjacent to each other.
[0006] According to another aspect, the circuitry sets the first coefficient value to zero in operation in response to the first coefficient value being less than 64.
[0007] According to another aspect, there is provided an encoder including: a block splitter that divides a first image into a plurality of blocks in operation; an intra predictor that predicts a block included in the first image using a reference block included in the first image in operation; an inter predictor that predicts a block included in the first image using a reference block included in a second image different from the first image in operation; a loop filter that filters a block included in the first image in operation; a transformer that transforms a prediction error between an original signal and a prediction signal generated by the intra predictor or the inter predictor to generate transform coefficients in operation; a quantizer that quantizes the transform coefficients to generate quantized coefficients in operation; and an entropy encoder that variably encodes the quantized coefficients to generate an encoded bitstream including the encoded quantized coefficients and control information. The loop filter performs the following operations:
[0008] In response to a first reconstructed image sample being outside a virtual boundary, replicate a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample;
[0009] Generate a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component;
[0010] Generate a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image sample of the chrominance component;
[0011] Generate a third coefficient value by adding the first coefficient value and the second coefficient value; and
[0012] Encode the third reconstructed image sample of the chrominance component using the third coefficient value.
[0013] According to another aspect, there is provided a decoder including circuitry and a memory coupled to the circuitry. In response to a first reconstructed image sample being outside a virtual boundary, the circuitry replicates a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample. The circuitry generates a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component. The circuitry generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image sample of the chrominance component. The circuitry generates a third coefficient value by adding the first coefficient value and the second coefficient value, and decodes the third reconstructed image sample of the chrominance component using the third coefficient value.
[0014] According to another aspect, there is provided a decoding apparatus including: a decoder that decodes an encoded bitstream in operation to output quantized coefficients; an inverse quantizer that inverse quantizes the quantized coefficients in operation to output transform coefficients; an inverse transformer that inverse transforms the transform coefficients in operation to output a prediction error; an intra predictor that predicts a block included in the first image using a reference block included in the first image in operation; an inter predictor that predicts a block included in the first image using a reference block included in a second image different from the first image in operation; a loop filter that filters a block included in the first image in operation; and an output terminal that outputs a picture including the first image in operation. The loop filter performs the following operations:
[0015] In response to a first reconstructed image sample being outside a virtual boundary, replicate a reconstructed sample that is inside and adjacent to the virtual boundary to generate the first reconstructed image sample;
[0016] Generate a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to the first reconstructed image sample of the luminance component;
[0017] Generate a second coefficient value by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image sample of the chrominance component;
[0018] Generate a third coefficient value by adding the first coefficient value and the second coefficient value; and
[0019] Decode a third reconstructed image sample of the chrominance component using the third coefficient value.
[0020] According to another aspect, there is provided an encoding method including:
[0021] In response to a first reconstructed image sample being outside a virtual boundary, replicate a reconstructed sample that is inside and adjacent to the virtual boundary to generate the first reconstructed image sample;
[0022] Generate a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to the first reconstructed image sample of the luminance component;
[0023] Generate a second coefficient value by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image sample of the chrominance component;
[0024] Generate a third coefficient value by adding the first coefficient value and the second coefficient value; and
[0025] Encode a third reconstructed image sample of the chrominance component using the third coefficient value.
[0026] According to another aspect, a decoding method is provided, which includes:
[0027] In response to a first reconstructed image sample being outside a virtual boundary, replicating a reconstructed sample that is inside and adjacent to the virtual boundary to generate the first reconstructed image sample;
[0028] Generating a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filter) process to the first reconstructed image sample of a luminance component;
[0029] Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to a second reconstructed image sample of a chrominance component;
[0030] Generating a third coefficient value by adding the first coefficient value and the second coefficient value; and
[0031] Decoding a third reconstructed image sample of the chrominance component using the third coefficient value.
[0032] In video coding technology, new methods need to be proposed to improve coding efficiency, enhance image quality, and reduce circuit scale. Some implementations of the embodiments of the present disclosure, including the constituent elements of the embodiments of the present disclosure considered alone or in various combinations, can promote one or more of the following: improvement of coding efficiency, enhancement of image quality, reduction of utilization of processing resources associated with encoding / decoding, reduction of circuit scale, improvement of encoding / decoding processing speed, etc.
[0033] In addition, some implementations of the embodiments of the present disclosure, including the constituent elements of the embodiments of the present disclosure considered alone or in various combinations, can promote an appropriate selection of one or more elements in encoding and decoding, such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations. Note that the present disclosure includes a disclosure of configurations and methods that can provide advantages other than the above advantages. Examples of such configurations and methods include configurations or methods for improving coding efficiency while reducing the use of processing resources.
[0034] According to the specification and the drawings, additional benefits and advantages of the disclosed embodiments will become apparent. The benefits and / or advantages can be obtained individually through various embodiments and features of the specification and the drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.
[0035] It should be noted that a general or specific embodiment can be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Figure 1 It is a schematic diagram showing an example of the functional configuration of a transmission system according to an embodiment.
[0037] Figure 2 Figure 2 It is a conceptual diagram showing an example of the hierarchical structure of data in a stream.
[0038] Figure 3 Figure 3 It is a conceptual diagram showing an example of a slice configuration.
[0039] Figure 4 Figure 4 It is a conceptual diagram showing an example of a tile configuration.
[0040] Figure 5 Figure 5 It is a conceptual diagram showing an example of the encoding structure in scalable coding.
[0041] Figure 6 Figure 6 It is a conceptual diagram showing an example of the encoding structure in scalable coding.
[0042] Figure 7 Figure 7 It is a block diagram showing an example of the functional configuration of an encoder according to an embodiment.
[0043] Figure 8 Figure 8 It is a functional block diagram showing an example of the installation of an encoder.
[0044] Figure 9 Figure 9 It is a flowchart showing an example of the overall encoding process performed by an encoder.
[0045] Figure 10 Figure 10 It is a conceptual diagram showing an example of block segmentation.
[0046] Figure 11 Figure 11 It is a block diagram showing an example of the functional configuration of a splitter according to an embodiment.
[0047] Figure 12 Figure 12 It is a conceptual diagram for showing an example of a splitting pattern.
[0048] Figure 13A Figure 13A It is a conceptual diagram showing an example of the syntax tree of a splitting pattern.
[0049] Figure 13B Figure 13B is a conceptual diagram showing another example of a syntactic tree for indicating a segmentation pattern.
[0050] Figure 14 Figure 14 is a diagram indicating example transform basis functions for various transform types.
[0051] Figure 15 Figure 15 is a conceptual diagram for showing an example of a spatial-variant transform (SVT).
[0052] Figure 16 Figure 16 is a flowchart showing an example of a process performed by a transducer.
[0053] Figure 17 Figure 17 is a flowchart showing another example of a process performed by a transducer.
[0054] Figure 18 Figure 18 is a block diagram showing an example of the functional configuration of a quantizer according to an embodiment.
[0055] Figure 19 Figure 19 is a flowchart showing an example of a quantization process performed by a quantizer.
[0056] Figure 20 Figure 20 is a block diagram showing an example of the functional configuration of an entropy encoder according to an embodiment.
[0057] Figure 21 Figure 21 is a conceptual diagram for illustrating an example process of context-based adaptive binary arithmetic coding (CABAC) in an entropy encoder.
[0058] Figure 22 Figure 22 is a block diagram showing an example of the functional configuration of a loop filter according to an embodiment.
[0059] Figure 23A Figure 23A is a conceptual diagram for showing an example of a filter shape used in an adaptive loop filter (ALF).
[0060] Figure 23B Figure 23B is a conceptual diagram for showing another example of a filter shape used in an ALF.
[0061] Figure 23C Figure 23C It is a conceptual diagram for showing another example of the filter shape used in the ALF.
[0062] Figure 23D Figure 23D It is a conceptual diagram for showing an example process of the cross-component ALF (CC-ALF).
[0063] Figure 23E Figure 23E It is a conceptual diagram for showing an example of the filter shape used in the CC-ALF.
[0064] Figure 23F Figure 23F It is a conceptual diagram for showing an example process of the joint chroma CCALF (JC-CCALF).
[0065] Figure 23G Figure 23G It is a table showing example weight index candidates that can be adopted in the JC-CCALF.
[0066] Figure 24 Figure 24 It is a block diagram showing an example of the specific configuration of a loop filter used as a deblocking filter (DBF).
[0067] Figure 25 Figure 25 It is a conceptual diagram for showing an example of a deblocking filter with symmetric filtering characteristics with respect to the block boundary.
[0068] Figure 26 Figure 26 It is a conceptual diagram for showing the block boundary where the deblocking filtering process is performed.
[0069] Figure 27 Figure 27 It is a conceptual diagram for showing an example of the boundary strength (Bs) value.
[0070] Figure 28 Figure 28 It is a flowchart showing an example of the process performed by the predictor of the encoder.
[0071] Figure 29 Figure 29 It is a flowchart showing another example of the process performed by the predictor of the encoder.
[0072] Figure 30 Figure 30 It is a flowchart showing another example of the process performed by the predictor of the encoder.
[0073] Figure 31 Figure 31 It is a conceptual diagram for showing the sixty-seven intra prediction modes used in intra prediction in an embodiment.
[0074] Figure 32 Figure 32 It is a flowchart showing an example of the process executed by the intra predictor.
[0075] Figure 33 Figure 33 It is a conceptual diagram for showing an example of a reference picture.
[0076] Figure 34 Figure 34 It is a conceptual diagram for showing an example of a reference picture list.
[0077] Figure 35 Figure 35 It is a flowchart showing an example of the basic process flow of inter prediction.
[0078] Figure 36 Figure 36 It is a flowchart showing an example of the derivation process of a motion vector.
[0079] Figure 37 Figure 37 It is a flowchart showing another example of the derivation process of a motion vector.
[0080] Figure 38A Figure 38A It is a conceptual diagram for showing an example representation of the mode for MV derivation.
[0081] Figure 38B Figure 38B It is a conceptual diagram for showing an example representation of the mode for MV derivation.
[0082] Figure 39 Figure 39 It is a flowchart showing an example of the inter prediction process in a conventional inter mode.
[0083] Figure 40 Figure 40 It is a flowchart showing an example of the inter prediction process in a conventional merge mode.
[0084] Figure 41 Figure 41 It is a conceptual diagram for showing an example of the motion vector derivation process in the merge mode.
[0085] Figure 42 Figure 42 It is a conceptual diagram for showing an example of the MV derivation process for the current picture through the HMVP merge mode.
[0086] Figure 43 Figure 43 is a flowchart showing an example of the frame rate up-conversion (FRUC) process.
[0087] Figure 44 Figure 44 is a conceptual diagram showing an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.
[0088] Figure 45 Figure 45 is a conceptual diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture.
[0089] Figure 46A Figure 46A is a conceptual diagram showing an example of deriving a motion vector for each sub-block based on motion vectors of multiple adjacent blocks.
[0090] Figure 46B Figure 46B is a conceptual diagram showing an example of deriving a motion vector for each sub-block in an affine mode using three control points.
[0091] Figure 47A Figure 47A is a conceptual diagram showing an example of MV derivation at control points in an affine mode.
[0092] Figure 47B Figure 47B is a conceptual diagram showing an example of MV derivation at control points in an affine mode.
[0093] Figure 47C Figure 47C is a conceptual diagram showing an example of MV derivation at control points in an affine mode.
[0094] Figure 48A Figure 48A is a conceptual diagram showing an affine mode using two control points.
[0095] Figure 48B Figure 48B is a conceptual diagram showing an affine mode using three control points.
[0096] Figure 49A Figure 49A is a conceptual diagram showing an example of a method for MV derivation at control points when the number of control points for an encoded block and the number of control points for the current block are different from each other.
[0097] Figure 49B Figure 49B It is a conceptual diagram showing another example of a method for MV derivation at control points when the number of control points for an encoding block and the number of control points for the current block are different from each other.
[0098] Figure 50 Figure 50 It is a flowchart showing an example of the process in the affine merge mode.
[0099] Figure 51 Figure 51 It is a flowchart showing an example of the process in the affine inter prediction mode.
[0100] Figure 52A Figure 52A It is a conceptual diagram for showing the generation of two triangular prediction images.
[0101] Figure 52B Figure 52B It is a conceptual diagram showing an example of the first part of the first partition overlapping with the second partition and the first sample set and the second sample set that can be weighted as part of the correction process.
[0102] Figure 52C Figure 52C It is a conceptual diagram for showing the first part of the first partition, which is the part of the first partition overlapping with a part of the adjacent partition.
[0103] Figure 53 Figure 53 It is a flowchart showing an example of the process in the triangular mode.
[0104] Figure 54 Figure 54 It is a conceptual diagram showing an example of the Advanced Temporal Motion Vector Prediction (ATMVP) mode in which the MV is derived in units of sub-blocks.
[0105] Figure 55 Figure 55 It is a flowchart showing the relationship between the merge mode and the Dynamic Motion Vector Refresh (DMVR).
[0106] Figure 56 Figure 56 It is a conceptual diagram showing an example of the DMVR.
[0107] Figure 57 Figure 57 It is a conceptual diagram showing another example of the DMVR for determining the MV.
[0108] Figure 58A Figure 58A It is a conceptual diagram for showing an example of motion estimation in DMVR.
[0109] Figure 58B Figure 58B It is a flowchart showing an example of the motion estimation process in DMVR.
[0110] Figure 59 Figure 59 It is a flowchart showing an example of the generation process of a predicted image.
[0111] Figure 60 Figure 60 It is a flowchart showing another example of the generation process of a predicted image.
[0112] Figure 61 Figure 61 It is a flowchart showing an example of the correction process of a predicted image by overlapping block motion compensation (OBMC).
[0113] Figure 62 Figure 62 It is a conceptual diagram for showing an example of the predicted image correction process by OBMC.
[0114] Figure 63 Figure 63 It is a conceptual diagram for showing a model assuming uniform linear motion.
[0115] Figure 64 Figure 64 It is a flowchart showing an example of the inter-frame prediction process according to BIO.
[0116] Figure 65 Figure 65 It is a functional block diagram showing an example of the functional configuration of an inter-frame predictor that can perform inter-frame prediction according to BIO.
[0117] Figure 66A Figure 66A It is a conceptual diagram for showing an example of the process of a predicted image generation method using the luminance correction process performed by LIC.
[0118] Figure 66B Figure 66B It is a flowchart showing an example of the process of a predicted image generation method using LIC.
[0119] Figure 67 Figure 67 It is a block diagram showing the functional configuration of a decoder according to an embodiment.
[0120] Figure 68 Figure 68 It is a functional block diagram showing an installation example of a decoder.
[0121] Figure 69 Figure 69 It is a flowchart showing an example of the overall decoding process executed by the decoder.
[0122] Figure 70 Figure 70 It is a conceptual diagram for showing the relationship between the segmentation determiner and other constituent elements.
[0123] Figure 71 Figure 71 It is a block diagram showing an example of the functional configuration of an entropy decoder.
[0124] Figure 72 Figure 72 It is a conceptual diagram for showing an example flow of the CABAC process in the entropy decoder.
[0125] Figure 73 Figure 73 It is a block diagram showing an example of the functional configuration of an inverse quantizer.
[0126] Figure 74 Figure 74 It is a flowchart showing an example of the inverse quantization process executed by the inverse quantizer.
[0127] Figure 75 Figure 75 It is a flowchart showing an example of the process executed by the inverse transformer.
[0128] Figure 76 Figure 76 It is a flowchart showing another example of the process executed by the inverse transformer.
[0129] Figure 77 Figure 77 It is a block diagram showing an example of the functional configuration of a loop filter.
[0130] Figure 78 Figure 78 It is a flowchart showing an example of the process executed by the predictor of the decoder.
[0131] Figure 79 Figure 79 It is a flowchart showing another example of the process executed by the predictor of the decoder.
[0132] Figure 80A Figure 80A It is a flowchart showing another example of the process executed by the predictor of the decoder.
[0133] Figure 80B Figure 80B It is a flowchart showing another example of the process performed by the predictor of the decoder.
[0134] Figure 80C Figure 80C It is a flowchart showing another example of the process performed by the predictor of the decoder.
[0135] Figure 81 Figure 81 It is a diagram showing an example of the process performed by the intra predictor of the decoder.
[0136] Figure 82 Figure 82 It is a flowchart showing an example of the MV derivation process in the decoder.
[0137] Figure 83 Figure 83 It is a flowchart showing another example of the MV derivation process in the decoder.
[0138] Figure 84 Figure 84 It is a flowchart showing an example of the process of inter prediction through the conventional inter mode in the decoder.
[0139] Figure 85 Figure 85 It is a flowchart showing an example of the process of inter prediction through the conventional merge mode in the decoder.
[0140] Figure 86 Figure 86 It is a flowchart showing an example of the process of inter prediction through the FRUC mode in the decoder.
[0141] Figure 87 Figure 87 It is a flowchart showing an example of the process of inter prediction through the affine merge mode in the decoder.
[0142] Figure 88 Figure 88 It is a flowchart showing an example of the process of inter prediction through the affine inter mode in the decoder.
[0143] Figure 89 Figure 89 It is a flowchart showing an example of the process of inter prediction through the triangle mode in the decoder.
[0144] Figure 90 Figure 90 It is a flowchart showing an example of the process of motion estimation through DMVR in the decoder.
[0145] Figure 91 Figure 91 It is a flowchart showing an example process of motion estimation by DMVR in a decoder.
[0146] Figure 92 Figure 92 It is a flowchart showing an example of the process of generating a prediction image in a decoder.
[0147] Figure 93 Figure 93 It is a flowchart showing another example of the process of generating a prediction image in a decoder.
[0148] Figure 94 Figure 94 It is a flowchart showing an example of the process of correcting a prediction image by OBMC in a decoder.
[0149] Figure 95 Figure 95 It is a flowchart showing an example of the process of correcting a prediction image by BIO in a decoder.
[0150] Figure 96 Figure 96 It is a flowchart showing an example of the process of correcting a prediction image by LIC in a decoder.
[0151] Figure 97 Figure 97 It is a flowchart of a sample process flow for decoding an image by applying the CCALF (Cross-Component Adaptive Loop Filter) process according to the first aspect.
[0152] Figure 98 Figure 98 It is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.
[0153] Figure 99 Figure 99 It is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.
[0154] Figure 100 Figure 100 It is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.
[0155] Figure 101 Figure 101 It is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.
[0156] Figure 102 Figure 102 It is a flowchart of a sample process flow for decoding an image by applying the CCALF process according to the second aspect.
[0157] Figure 103 Figure 103 Illustrates the sample positions of the cropping parameters to be parsed from, for example, the VPS, APS, SPS, PPS, slice header, CTU, or TU of a bitstream.
[0158] Figure 104 Figure 104 Illustrates an example of the cropping parameters.
[0159] Figure 105 Figure 105 Is a flowchart of a sample process flow for decoding an image using the CCALF process with filter coefficients according to a third aspect.
[0160] Figure 106 Figure 106 Is a conceptual diagram indicating an example of the positions of the filter coefficients to be used in the CCALF process.
[0161] Figure 107 Figure 107 Is a conceptual diagram indicating an example of the positions of the filter coefficients to be used in the CCALF process.
[0162] Figure 108 Figure 108 Is a conceptual diagram indicating an example of the positions of the filter coefficients to be used in the CCALF process.
[0163] Figure 109 Figure 109 Is a conceptual diagram indicating an example of the positions of the filter coefficients to be used in the CCALF process.
[0164] Figure 110 Figure 110 Is a conceptual diagram indicating an example of the positions of the filter coefficients to be used in the CCALF process.
[0165] Figure 111 Figure 111 Is a conceptual diagram indicating a further example of the positions of the filter coefficients to be used in the CCALF process.
[0166] Figure 112 Figure 112 Is a conceptual diagram indicating a further example of the positions of the filter coefficients to be used in the CCALF process.
[0167] Figure 113 Figure 113 Is a block diagram showing the functional configuration of the CCALF process performed by an encoder and a decoder according to an embodiment.
[0168] Figure 114 Figure 114 It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using a filter selected from multiple filters according to the fourth aspect.
[0169] Figure 115 Figure 115 It illustrates an example of a process flow for selecting a filter.
[0170] Figure 116 Figure 116 It illustrates an example of a filter.
[0171] Figure 117 Figure 117 It illustrates an example of a filter.
[0172] Figure 118 Figure 118 It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the fifth aspect.
[0173] Figure 119[[END It illustrates an example of the number of coefficients to be parsed from a bitstream.
[0174] It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the sixth aspect.
[0175] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.
[0176] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.
[0177] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.
[0178] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component of a current sample by calculating the weighted average of adjacent samples, where the positions of the adjacent samples are adaptively determined as the chrominance type.
[0179] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component of the current sample by calculating the weighted average of adjacent samples, where the positions of the adjacent samples are adaptively determined according to the chromaticity type.
[0180] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component by applying a bit shift to the output value of the weighted calculation.
[0181] It is a conceptual diagram showing an example of generating the CCALF value of the luminance component by applying a bit shift to the output value of the weighted calculation.
[0182] It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the seventh aspect.
[0183] It shows the sample positions of one or more parameters to be parsed from the bitstream.
[0184] It shows the sample process of retrieving one or more parameters.
[0185] Figure 131 Figure 131 It shows the sample value of the second parameter.
[0186] Figure 132 Figure 132 It shows an example of parsing the second parameter using arithmetic coding.
[0187] Figure 133 Figure 133 It is a conceptual diagram of a variant of the present embodiment applied to rectangular partitions and non-rectangular partitions (such as triangular partitions).
[0188] Figure 134 Figure 134 It is a flowchart of an example process flow for decoding an image by applying the CCALF process using parameters according to the eighth aspect.
[0189] Figure 135 Figure 135 It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the eighth aspect.
[0190] Figure 136 Figure 136 It shows the example positions of chromaticity sample types 0 to 5.
[0191] Figure 137 Figure 137 It is a conceptual diagram showing the symmetric filling of a sample.
[0192] Figure 138 Figure 138 It is a conceptual diagram showing the symmetric filling of a sample.
[0193] Figure 139 Figure 139 It is a conceptual diagram showing the symmetric filling of a sample.
[0194] Figure 140 Figure 140 It is a conceptual diagram showing the asymmetric filling of a sample.
[0195] Figure 141 Figure 141 It is a conceptual diagram showing the asymmetric filling of a sample.
[0196] Figure 142 Figure 142 It is a conceptual diagram showing the asymmetric filling of a sample.
[0197] Figure 143 Figure 143 It is a conceptual diagram showing the asymmetric filling of a sample.
[0198] Figure 144 Figure 144 It is a conceptual diagram showing the further asymmetric filling of a sample.
[0199] Figure 145 Figure 145 It is a conceptual diagram showing the further asymmetric filling of a sample.
[0200] Figure 146 Figure 146 It is a conceptual diagram showing the further asymmetric filling of a sample.
[0201] Figure 147 Figure 147 It is a conceptual diagram showing the further asymmetric filling of a sample.
[0202] Figure 148 Figure 148 It is a conceptual diagram showing the further symmetric filling of a sample.
[0203] Figure 149 Figure 149 It is a conceptual diagram showing the further symmetric filling of a sample.
[0204] Figure 150 Figure 150 It is a conceptual diagram showing the further symmetric filling of a sample.
[0205] Figure 151 Figure 151 is a conceptual diagram showing further sample asymmetric padding.
[0206] Figure 152 Figure 152 is a conceptual diagram showing further sample asymmetric padding.
[0207] Figure 153 Figure 153 is a conceptual diagram showing further sample asymmetric padding.
[0208] Figure 154 Figure 154 is a conceptual diagram showing further sample asymmetric padding.
[0209] Figure 155 Figure 155 illustrates a further example of padding with horizontal and vertical virtual boundaries.
[0210] Figure 156 Figure 156 is a block diagram showing the functional configuration of an encoder and a decoder according to an example, where symmetric padding is used for the virtual boundary positions of ALF and symmetric or asymmetric padding is used for the virtual boundary positions of CC-ALF.
[0211] Figure 157 Figure 157 is a block diagram showing the functional configuration of an encoder and a decoder according to another example, where symmetric padding is used for the virtual boundary positions of ALF and unilateral padding is used for the virtual boundary positions of CC-ALF.
[0212] Figure 158 Figure 158 is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.
[0213] Figure 159 Figure 159 is a conceptual diagram showing an example of unilateral padding with horizontal and vertical virtual boundaries.
[0214] Figure 160 Figure 160 is a diagram showing an example of the overall configuration of a content providing system for implementing a content distribution service.
[0215] Figure 161 Figure 161 is a conceptual diagram for showing an example of a display screen of a web page.
[0216] Figure 162 Figure 162 is a conceptual diagram for showing an example of a display screen of a web page.
[0217] Figure 163 Figure 163 is a block diagram showing an example of a smart phone.
[0218] Figure 164 Figure 164 is a block diagram showing an example of the functional configuration of a smart phone. DETAILED DESCRIPTION
[0219] In the drawings, unless otherwise indicated by the context, the same reference numerals denote similar elements. The sizes and relative positions of the elements in the drawings are not necessarily drawn to scale.
[0220] Hereinafter, embodiments will be described with reference to the drawings. Note that each of the embodiments described below shows a general or specific example. The numerical values, shapes, materials, components, arrangements and connections of the components, steps, relationships and orders of the steps, etc. indicated in the following embodiments are only examples and are not intended to limit the scope of the claims.
[0221] Embodiments of an encoder and a decoder will be described below. The embodiments are examples of the encoder and the decoder, and the processes and / or configurations presented in the description of the aspects of the present disclosure can be applied to the encoder and the decoder. The processes and / or configurations can also be implemented in encoders and decoders different from the encoder and decoder according to the embodiments. For example, regarding the processes and / or configurations applied to the embodiments, any of the following can be implemented:
[0222] (1) Any component of the encoder or decoder according to the embodiments presented in the description of the aspects of the present disclosure can be replaced with or combined with another component presented anywhere in the description of the aspects of the present disclosure.
[0223] (2) In the encoder or decoder according to the embodiments, arbitrary changes can be made to the functions or processes performed by one or more components of the encoder or decoder, such as addition, replacement, removal, etc. of the functions or processes. For example, any function or process can be replaced with or combined with another function or process presented anywhere in the description of the aspects of the present disclosure.
[0224] (3) In the method implemented by the encoder or decoder according to the embodiments, arbitrary changes can be made, such as addition, replacement, and removal of one or more processes included in the method. For example, any process in the method can be replaced with or combined with another process presented anywhere in the description of the aspects of the present disclosure.
[0225] (4) One or more components included in an encoder or decoder according to an embodiment can be combined with components presented anywhere in the description of aspects of the present disclosure, can be combined with components including one or more functions presented anywhere in the description of aspects of the present disclosure, and can be combined with components implementing one or more processes implemented by components presented in the description of aspects of the present disclosure.
[0226] (5) A component including one or more functions of an encoder or decoder according to an embodiment, or a component implementing one or more processes of an encoder or decoder according to an embodiment, can be combined with or replaced by components presented anywhere in the description of aspects of the present disclosure, can be combined with or replaced by components including one or more functions presented anywhere in the description of aspects of the present disclosure, or can be combined with or replaced by components implementing one or more processes presented anywhere in the description of aspects of the present disclosure.
[0227] (6) In a method implemented by an encoder or decoder according to an embodiment, any process included in the method can be replaced by or combined with a process presented anywhere in the description of aspects of the present disclosure or with any corresponding or equivalent process.
[0228] (7) One or more processes included in a method implemented by the encoder or decoder according to an embodiment can be combined with processes presented anywhere in the description of aspects of the present disclosure.
[0229] (8) The implementation manners of processes and / or configurations presented in the description of aspects of the present disclosure are not limited to an encoder or decoder according to an embodiment. For example, the processes and / or configurations can be implemented in a device for purposes different from those of a motion image encoder or a motion image decoder disclosed in the embodiment.
[0230] (Term Definitions)
[0231] The corresponding terms can be defined as indicated by the following examples.
[0232] An image is a data unit configured with a set of pixels, is a picture, or includes blocks smaller than pixels. In addition to video, an image also includes a still image.
[0233] A picture is an image processing unit configured with a set of pixels, and can also be referred to as a frame or a field. For example, a picture can take the form of an array of luminance samples in a monochrome format or an array of luminance samples and two corresponding arrays of chrominance samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0234] A block is a processing unit which is a set of a determined number of pixels. A block can have any number of different shapes. For example, a block can have a rectangle of M×N (M columns × N rows) pixels, a square of M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, shards, bricks, CTUs, super blocks, basic segmentation units, VPDUs, processing segmentation units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and sub-blocks. A block can be in the form of an M×N sample array or an M×N transform coefficient array. For example, a block can be a square or rectangular pixel region including a luminance matrix and two chrominance matrices.
[0235] A pixel or sample is the smallest point of an image. A pixel or sample includes pixels at integer positions, as well as pixels at sub-pixel positions, e.g., generated based on pixels at integer positions.
[0236] A pixel value or sample value is a characteristic value of a pixel. A pixel value or sample value can include one or more of a luminance value, a chrominance value, an RGB gray level, a depth value, a binary value of zero or 1, etc.
[0237] Chroma or chrominance is the intensity of a color, usually denoted by the symbols Cb and Cr, which specify the value of an array of samples or the value of a single sample representing one of two color difference signals related to the primary colors.
[0238] Luma or luminance is the brightness of an image, usually denoted by the symbol or subscript Y or L, which specify the value of an array of samples or the value of a single sample representing a monochromatic signal related to the primary colors.
[0239] A flag includes one or more bits indicating a value of, for example, a parameter or an index. A flag can be a binary flag, which indicates the binary value of the flag, and it can also indicate a non-binary value of a parameter.
[0240] A signal conveys information which is symbolized or encoded into the signal. Signals include discrete digital signals and continuous analog signals.
[0241] A stream or bitstream is a digital data string of a digital data stream. A stream or bitstream can be a single stream or can be configured with multiple streams having multiple hierarchical layers. A stream or bitstream can be transmitted in a serial communication manner using a single transmission path, or can be transmitted in a packet communication manner using multiple transmission paths.
[0242] The difference refers to various mathematical differences, such as the simple difference (x - y), the absolute value of the difference (|x - y|), the difference of squares (x^2 - y^2), the square root of the difference (√(x – y)), the weighted difference (ax - by: a and b are constants), the offset difference (x - y + a: a is the offset), etc. In the case of scalars, the simple difference is sufficient, and the difference calculation is included.
[0243] The sum refers to various mathematical sums, such as the simple sum (x + y), the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), the weighted difference (ax + by: a and b are constants), the offset sum (x + y + a: a is the offset), etc. In the case of scalars, the simple sum is sufficient, and the sum calculation is included.
[0244] A frame is a combination of a top field and a bottom field, where sampling rows 0, 2, 4,... are from the top field, and sampling rows 1, 3, 5,... are from the bottom field.
[0245] A slice is an integer number of coding tree units in all subsequent dependent slices (if any) that are included in an independent slice segment and before the next independent slice segment (if any) within the same access unit.
[0246] A tile is a rectangular region of coding tree blocks within a specific tile column and a specific tile row in a picture. A tile can be a rectangular region of a frame that is designed to be independently decodable and encodable, although loop filtering across tile edges can still be applied.
[0247] A coding tree unit (CTU) can be a coding tree block of the luma samples of a picture with three sample arrays, or two corresponding coding tree blocks of the chroma samples. Alternatively, a CTU can be a coding tree block of the samples of a monochrome picture and a picture encoded using three separate color planes and a syntax structure for encoding samples. A superblock can be a square block of 64×64 pixels consisting of 1 or 2 mode information blocks, or recursively divided into four 32×32 blocks, which can themselves be further divided.
[0248] (System configuration)
[0249] First, a transmission system according to an embodiment will be described. Figure 1 is a schematic diagram showing an example of the configuration of a transmission system 400 according to an embodiment.
[0250] The transmission system 400 is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. As shown, the transmission system 400 includes an encoder 100, a network 300, and a decoder 200 as Figure 1 shown.
[0251] The image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image and outputs the stream to the network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by encoding.
[0252] It should be noted that the image before being encoded by the encoder 100 is also referred to as the original image, original signal, or original sample. The image can be a video or a still image. The image is a general concept of sequences, pictures, and blocks, and thus, unless otherwise specified, is not limited to a spatial region with a specific size and a temporal region with a specific size. The image is an array of pixels or pixel values, and the signal representing the image or pixel values is also referred to as a sample. The stream can be referred to as a bitstream, encoded bitstream, compressed bitstream, or encoded signal. In addition, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, image encoding method, or video encoding method.
[0253] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 200 can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a two-way communication network and can be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasts and satellite broadcasts. Alternatively, the network 300 can be replaced by a recording medium such as a digital versatile disc (DVD) and a Blu-ray Disc (BD), on which the stream is recorded.
[0254] The decoder 200 generates a decoded image as an uncompressed image by, for example, decoding the stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method adopted by the encoder 100.
[0255] It should be noted that the decoder 200 can also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 can also be referred to as a decoding method, image decoding method, or video decoding method.
[0256] (Data Structure)
[0257] Figure 2 is a conceptual diagram showing an example of the hierarchical structure of data in the stream. For convenience, the transmission system 400 of Figure 1 will be described Figure 2 . The stream includes, for example, a video sequence. As shown in (a) of Figure 2 , the video sequence includes one or more video parameter sets (VPSs), one or more sequence parameter sets (SPSs), one or more picture parameter sets (PPSs), supplementary enhancement information (SEI), and a plurality of pictures.
[0258] In a video having multiple layers, the VPS may include coding parameters shared between some of the multiple layers, as well as coding parameters related to some of the multiple layers included in the video or to a single layer.
[0259] The SPS includes parameters for a sequence, i.e., coding parameters that the decoder 200 refers to for decoding the sequence. For example, the coding parameters may indicate the width or height of a picture. It should be noted that there may be multiple SPSs.
[0260] The PPS includes parameters for a picture, i.e., coding parameters that the decoder 200 refers to for decoding each picture in the sequence. For example, the coding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. It should be noted that there may be multiple PPSs. Each of the SPS and PPS may be abbreviated as a parameter set.
[0261] As Figure 2 shown in (b) of [], a picture may include a picture header and one or more slices. The picture header includes coding parameters that the decoder 200 refers to for decoding the one or more slices.
[0262] As Figure 2 shown in (c) of [], a slice includes a slice header and one or more tiles. The slice header includes coding parameters that the decoder 200 refers to for decoding the one or more tiles.
[0263] As Figure 2 shown in (d) of [], a tile includes one or more coding tree units (CTUs).
[0264] It should be noted that a picture may not include any slices and may include slice groups instead of slices. In this case, a slice group includes at least one slice. Additionally, a tile may include slices.
[0265] The CTU is also referred to as a superblock or a basis splitting unit. As Figure 2 shown in (e) of [], the CTU includes a CTU header and at least one coding unit (CU). As shown in the figure, the CTU includes four coding units CU(10), CU(11), CU(12), and CU(13). The CTU header includes coding parameters that the decoder 200 refers to for decoding at least one CU.
[0266] The CU can be divided into multiple smaller CUs. As shown in the figure, CU (10) is not divided into smaller coding units; CU (11) is divided into four smaller coding units CU (110), CU (111), CU (112), and CU (113); CU (12) is not divided into smaller coding units; and CU (13) is divided into seven smaller coding units CU (1310), CU (1311), CU (1312), CU (1313), CU (132), CU (133), and CU (134). As Figure 2 As shown in (f) of Figure 2 , the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual to be described later. Although the CU is basically the same as the prediction unit (PU) and the transform unit (TU), it should be noted that, for example, the sub-block transform (SBT) to be described later may include multiple TUs smaller than the CU. In addition, the CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. The VPDU is, for example, a fixed unit that can be processed in one stage when performing pipeline processing in hardware.
[0267] It should be noted that the stream may not include Figure 2 all the hierarchical layers shown in Figure 2 . The order of the hierarchical layers can be exchanged, or any hierarchical layer can be replaced by another hierarchical layer. Here, the picture that is the target of the process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as the current picture. When the process is an encoding process, the current picture represents the current picture to be encoded, and when the process is a decoding process, the current picture represents the current picture to be decoded. Similarly, for example, the CU or CU block that is the target of the process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as the current block. When the process is an encoding process, the current block represents the current block to be encoded, and when the process is a decoding process, the current block represents the current block to be decoded.
[0268] (Picture Structure: Slice / Tile)
[0269] The picture can be configured with one or more slice units or one or more tile units to facilitate parallel encoding / decoding of the picture.
[0270] A slice is the basic coding unit included in the picture. The picture can include, for example, one or more slices. In addition, a slice includes one or more coding tree units (CTUs).
[0271] Figure 3 is a conceptual diagram for showing an example of slice configuration. For example, in Figure 3In it, the picture includes 11×8 CTUs and is divided into four slices (slice 1 to slice 4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is the shape obtained by horizontally dividing the picture. The boundary of each slice does not need to coincide with the image edge and can coincide with any boundary between CTUs in the image. The processing order (encoding order or decoding order) of CTUs in the slice is, for example, the raster scan order. The slice includes a slice header and encoded data. The features of the slice can be written into the slice header. The features can include the CTU address of the top CTU in the slice, the slice type, etc.
[0272] A tile is a unit of a rectangular area included in the picture. The tiles of the picture can be assigned numbers called TileId in the raster scan order.
[0273] Figure 4 is a conceptual diagram for showing an example of the tile configuration. For example, in Figure 4 In it, the picture includes 11×8 CTUs and is divided into four tiles (tile 1 to tile 4) of rectangular areas. When using tiles, the processing order of CTUs may be different from the processing order in the case of not using tiles. When not using tiles, multiple CTUs in the picture are usually processed in the raster scan order. When using multiple tiles, at least one CTU in each of the multiple tiles is processed in the raster scan order. For example, as Figure 4 shown, the processing order of the CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then continues from the left end of the second column of tile 1 to the right end of the second column of tile 1.
[0274] It should be noted that one tile can include one or more slices, and one slice can include one or more tiles.
[0275] It should be noted that the picture can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more tiles. The picture can be configured with one of a tile set, a tile group, and a tile. For example, assume that the order of scanning multiple tiles for each tile set in the raster scan order is the basic encoding order of the tiles. Assume that a set of one or more tiles that are consecutive in the basic encoding order in each tile set is a tile group. Such a picture can be configured by a splitter 102 (see Figure 7 ) described later.
[0276] (Scalable Encoding)
[0277] Figure 5 and Figure 6is a conceptual diagram showing an example of a scalable stream structure and will be described with reference to Figure 1 for convenience.
[0278] As Figure 5 shown, the encoder 100 can generate a temporally / spatially scalable stream by dividing each of a plurality of pictures into any of a plurality of layers and encoding the pictures in the layers. For example, the encoder 100 encodes the pictures for each layer, thereby achieving scalability in the case where the enhancement layer exists above the base layer. This encoding of each picture is also referred to as scalable encoding. In this way, the decoder 200 can switch the image quality of the image displayed by decoding the stream. In other words, the decoder 200 can determine which layer to decode based on internal factors such as the processing power of the decoder 200 and external factors such as the communication bandwidth state. As a result, the decoder 200 can decode the content while freely switching between low resolution and high resolution. For example, a user of the stream watches a video of the stream halfway through on a smartphone on the way home and continues to watch the video on a device (e.g., a TV connected to the Internet) at home. It should be noted that each of the above smartphone and device includes a decoder 200 with the same or different performance. In this case, when the device decodes the layer to a higher layer in the stream, the user can watch a high-quality video at home. In this way, the encoder 100 does not need to generate multiple streams with different image qualities of the same content, and thus can reduce the processing load.
[0279] In addition, the enhancement layer can include meta-information based on statistical information about the image. The decoder 200 can generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the metadata. Super-resolution imaging can include, for example, an increase in the SN ratio at the same resolution, an increase in resolution, etc. The metadata can include, for example, information for identifying linear or non-linear filter coefficients used in the super-resolution process, or information for identifying parameter values in filtering processes, machine learning, or least squares methods (used in the super-resolution process), etc.
[0280] In an embodiment, a configuration can be provided in which a picture is divided into, for example, slices according to the meaning of an object in the picture. In this case, the decoder 200 can decode only a partial region in the picture by selecting the slices to be decoded. In addition, the attributes of the object (person, car, ball, etc.) and the position of the object in the picture (coordinates in the same image) can be stored as metadata. In this case, the decoder 200 can identify the position of the desired object based on the metadata and determine the slices including the object. For example, as Figure 6As shown, a data storage structure different from the image data can be used to store metadata, such as the SEI (Supplemental Enhancement Information) message in HEVC. This metadata indicates, for example, the location, size, or color of the main object.
[0281] The metadata can be stored in units of multiple pictures (e.g., a stream, a sequence, or a random access unit). In this way, the decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the picture unit information, can identify the pictures in which the object (person) appears and determine the position of the object in the pictures.
[0282] (Encoder)
[0283] An encoder according to an embodiment will be described. Figure 7 FIG. is a block diagram showing a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes video in units of blocks.
[0284] As Figure 7 shown, the encoder 100 is a device that encodes an image in units of blocks, and includes a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra predictor 124, an inter predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra predictor 124 and the inter predictor 126 are part of the prediction controller.
[0285] The encoder 100 is implemented as, for example, a general - purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor acts as the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra predictor 124, the inter predictor 126, and the prediction controller 128. Alternatively, the encoder 100 can be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra predictor 124, the inter predictor 126, and the prediction controller 128.
[0286] (Installation example of the encoder)
[0287] Figure 8 FIG. is a functional block diagram showing an installation example of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example,Figure 7 A plurality of components of the encoder 100 shown are mounted on Figure 8 the processor a1 and the memory a2 shown.
[0288] The processor a1 is a circuit that performs information processing and is coupled to the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding images. The processor a1 can be a processor such as a CPU. In addition, the processor a1 can be an aggregate of multiple electronic circuits. Additionally, for example, the processor a1 can assume Figure 7 the role of two or more components among a plurality of components such as the encoder 100 shown.
[0289] The memory a2 is a dedicated or general-purpose memory for storing information used by the processor a1 to encode images. The memory a2 can be an electronic circuit and can be connected to the processor a1. In addition, the memory a2 can be included in the processor a1. In addition, the memory a2 can be an aggregate of multiple electronic circuits. Additionally, the memory a2 can be a magnetic disk, an optical disk, etc., or can be represented as a storage device, a recording medium, etc. In addition, the memory a2 can be a non-volatile memory or a volatile memory.
[0290] For example, the memory a2 can store the image to be encoded or the bitstream corresponding to the encoded image. In addition, the memory a2 can store a program for causing the processor a1 to encode images.
[0291] In addition, for example, the memory a2 can act as Figure 7 the role of two or more components for storing information among a plurality of components such as the encoder 100 shown. For example, the memory a2 can act as Figure 7 the role of the block memory 118 and the frame memory 122 shown. More specifically, the memory a2 can store reconstructed blocks, reconstructed pictures, etc.
[0292] It should be noted that in the encoder 100, not all of the plurality of components etc. shown may be implemented, and not all of the processes described here may be executed. Figure 7 A part of the components etc. shown can be included in another device, or a part of the processes described here can be executed by another device. Figure 7 Shown components etc. can be included in another device, or a part of the processes described here can be executed by another device.
[0293] Hereinafter, the overall flow of the process executed by the encoder 100 will be described, and then each component included in the encoder 100 will be described.
[0294] (Overall Flow of Encoding Process)
[0295] Figure 9FIG. 0 is a flowchart showing an example of the overall encoding process performed by the encoder 100, and will be described with reference to Figure 7 for convenience.
[0296] First, the splitter 102 of the encoder 100 splits each picture included in the input image into a plurality of blocks having a fixed size (e.g., 128×128 pixels) (step Sa_1). The splitter 102 then selects a splitting pattern for the fixed-size blocks (also referred to as block shapes) (step Sa_2). In other words, the splitter 102 further splits the fixed-size blocks into a plurality of blocks forming the selected splitting pattern. The encoder 100 performs steps Sa_3 to Sa_9 for each of the plurality of blocks, for that block (i.e., the current block to be encoded).
[0297] The prediction controller 128 and the prediction executor (which includes the intra predictor 124 and the inter predictor 126) generate a predicted image of the current block (step Sa-3). The predicted image may also be referred to as a prediction signal, a prediction block, or a prediction sample.
[0298] Next, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Sa_4). The prediction residual may also be referred to as a prediction error.
[0299] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients may sometimes be referred to as a coefficient block.
[0300] Next, the entropy encoder 110 encodes the plurality of quantized coefficients and the prediction parameters related to the generation of the predicted image (specifically, entropy encoding) to generate a stream (step Sa_6). The stream may sometimes be referred to as an encoded bitstream or a compressed bitstream.
[0301] Next, the inverse quantizer 112 performs inverse quantization of the plurality of quantized coefficients, and the inverse transformer 114 performs inverse transformation of the result to recover the prediction residual (step Sa_7).
[0302] Next, the adder 116 adds the predicted image and the recovered prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image may also be referred to as a reconstructed block or a decoded image block.
[0303] When the reconstructed image is generated, the loop filter 120 performs filtering of the reconstructed image as needed (step Sa_9).
[0304] The encoder 100 then determines whether the encoding of the entire picture has been completed (step Sa_10). When it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeatedly performed on the next block of the image.
[0305] Although in the above example the encoder 100 selects a splitting mode for a block of a fixed size and encodes each block according to the splitting mode, it should be noted that each block can be encoded according to the corresponding splitting mode among multiple splitting modes. In this case, the encoder 100 can evaluate the cost of each of the multiple splitting modes, and for example, can select the stream that can be obtained by encoding according to the splitting mode that generates the minimum cost as the output stream.
[0306] As shown in the figure, the processes in steps Sa_1 to Sa_10 are sequentially executed by the encoder 100. Alternatively, two or more processes can be executed in parallel, the processes can be reordered, and so on.
[0307] The encoding process adopted by the encoder 100 is a hybrid encoding using predictive encoding and transform encoding. In addition, the predictive encoding is executed by an encoding loop configured with a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, a block memory 118, a frame memory 122, an intra predictor 124, an inter predictor 126, and a prediction controller 128. In other words, the prediction executor configured with the intra predictor 124 and the inter predictor 126 is part of the encoding loop.
[0308] (Splitter)
[0309] The splitter 102 divides each picture included in the original image into a plurality of blocks and outputs each block to the subtractor 104. For example, the splitter 102 first divides the picture into blocks of a fixed size (e.g., 128×128 pixels). Other fixed block sizes can be adopted. The blocks of a fixed size are also referred to as coding tree units (CTUs). The splitter 102 then divides each fixed-size block into variable-size blocks (e.g., 64×64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a splitting mode. The variable-size blocks can also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in the picture can be processed in units of CUs, PUs, or TUs.
[0310] Figure 10 is a conceptual diagram for showing an example of block splitting according to an embodiment. In Figure 10In this figure, solid lines indicate the block boundaries of blocks divided by quadtree block division, and dashed lines indicate the block boundaries of blocks divided by binary tree block division.
[0311] Here, block 10 is a square block with 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into four square 64×64 pixel blocks (quadtree block division).
[0312] The upper-left 64×64 pixel block is further vertically divided into two rectangular 32×64 pixel blocks, and the left 32×64 pixel block is further vertically divided into two rectangular 16×64 pixel blocks (binary tree block division). As a result, the upper-left 64×64 pixel block is divided into two 16×64 pixel blocks 11 and 12 and a 32×64 pixel block 13.
[0313] The upper-right 64×64 pixel block is horizontally divided into two rectangular 64×32 pixel blocks 14 and 15 (binary tree block division).
[0314] The lower-left 64×64 pixel block is first divided into four square 32×32 pixel blocks (quadtree block division). The upper-left and lower-right blocks among the four square 32×32 pixel blocks are further divided. The upper-left square 32×32 pixel block is vertically divided into two rectangular 16×32 pixel blocks, and the right 16×32 pixel block is further horizontally divided into two 16×16 pixel blocks (binary tree block division). The lower-right 32×32 pixel block is horizontally divided into two 32×16 pixel blocks (binary tree block division). The upper-right square 32×32 pixel block is horizontally divided into two rectangular 32×16 pixel blocks (binary tree block division). As a result, the lower-left square 64×64 pixel block is divided into rectangular 16×32 pixel block 16, two square 16×16 pixel blocks 17 and 18, two square 32×32 pixel blocks 19 and 20, and two rectangular 32×16 pixel blocks 21 and 22.
[0315] The lower-right 64×64 pixel block 23 is not divided.
[0316] As described above, in Figure 10 , based on recursive quadtree and binary tree block division, block 10 is divided into 13 variable-size blocks 11 to 23. This type of division is also called quadtree plus binary tree (QTBT) division.
[0317] It should be noted that in Figure 10 , a block is divided into four or two blocks (quadtree or binary tree block division), but the division is not limited to these examples. For example, a block can be divided into three blocks (ternary block division). The division including this ternary block division is also called multi-type tree (MBT) division.
[0318] Figure 11 is a block diagram showing an example of the functional configuration of the splitter 102 according to one embodiment. As Figure 11 shown, the splitter 102 may include a block splitting determiner 102a. As an example, the block splitting determiner 102a may perform the following process.
[0319] For example, the block splitting determiner 102a may obtain or retrieve block information from the block memory 118 and / or the frame memory 122, and determine a splitting pattern (e.g., the above-mentioned splitting pattern) based on the block information. The splitter 102 splits the original image according to the splitting pattern and outputs at least one block obtained by the splitting to the subtractor 104.
[0320] In addition, for example, the block splitting determiner 102a outputs one or more parameters indicating the determined splitting pattern (e.g., the above-mentioned splitting pattern) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 may transform the prediction residual based on one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 may generate a prediction image based on one or more parameters. In addition, the entropy encoder 110 may perform entropy coding on one or more parameters.
[0321] The parameters related to the splitting pattern may be written in the stream. As an example, it is as follows.
[0322] Figure 12 is a conceptual diagram for showing an example of the splitting pattern. Examples of the splitting pattern include: splitting into four regions (QT), where one block is split into two regions both horizontally and vertically; splitting into three regions (HT or VT), where one block is split in the same direction in a 1:2:1 ratio; splitting into two regions (HB or VB), where one block is split in the same direction in a 1:1 ratio; and no splitting (NS).
[0323] It should be noted that the splitting pattern does not have a block splitting direction in the cases of splitting into four regions and no splitting, and the splitting pattern has splitting direction information in the cases of splitting into two regions or three regions.
[0324] Figure 13A is a conceptual diagram for showing an example of the syntax tree of the splitting pattern.
[0325] Figure 13B is a conceptual diagram for showing another example of the syntax tree of the splitting pattern.
[0326] Figure 13A and Figure 13BA conceptual diagram showing an example of a syntax tree for a segmentation pattern. In Figure 13A the example, first, there is information indicating whether to perform segmentation (S: segmentation flag), and next, there is information indicating whether to perform segmentation into four regions (QT: QT flag). Next, there is information indicating which of the three-region and two-region segmentations to perform (TT: TT flag or BT: BT flag), and then there is information indicating the division direction (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of at least one block obtained by segmenting according to such a segmentation pattern can be further repeatedly segmented in a similar process. In other words, as an example, whether to perform segmentation, whether to perform segmentation into four regions, which of the horizontal and vertical directions is the direction in which the segmentation method is to be performed, which of the three-region and two-region segmentations to perform can be determined recursively, and the determination result can be encoded in the stream according to Figure 13A the encoding order disclosed in the syntax tree shown.
[0327] In addition, although the information items indicating S, QT, TT, and Ver are arranged in the listed order in the Figure 13A syntax tree shown, the information items indicating S, QT, Ver, and BT can also be arranged in the listed order. In other words, in Figure 13B the example, first, there is information indicating whether to perform segmentation (S: segmentation flag), and next, there is information indicating whether to perform segmentation into four regions (QT: QT flag). Next, there is information indicating the division direction (Ver: vertical flag, or Hor: horizontal flag), and next, there is information indicating which of the two-region and three-region segmentations to perform (BT: BT flag or TT: TT flag).
[0328] It should be noted that the above segmentation pattern is an example, and a segmentation pattern other than the described one can be used, or a part of the described segmentation pattern can be used.
[0329] (Subtractor)
[0330] The subtractor 104 subtracts the predicted image (predicted samples input from the prediction controller 128 indicated below) from the original image in units of blocks. The original image is input from the segmenter 102 and segmented by the segmenter 102. In other words, the subtractor 104 calculates the prediction residual (also referred to as error) of the current block. The subtractor 104 then outputs the calculated prediction residual to the transformer 106.
[0331] The original image can be an image for which a signal (e.g., a luminance signal and two chrominance signals) representing each picture included in a video has been input to the encoder 100. The signal representing the image may also be referred to as a sample.
[0332] (Transformer)
[0333] The transformer 106 transforms the prediction residual in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to the quantizer 108. More specifically, the transformer 106 applies, for example, a defined discrete cosine transform (DCT) or discrete sine transform (DST) to the prediction residual in the spatial domain. The defined DCT or DST may be predefined.
[0334] It should be noted that the transformer 106 can adaptively select a transform type from multiple transform types and transform the prediction residual into transform coefficients by using a transform basis function corresponding to the selected transform type. Such a transform is also referred to as an explicit multi-core transform (EMT) or an adaptive multi-core transform (AMT). The transform basis function may also be referred to as a basis.
[0335] The transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transform types can be represented as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 is a chart of example transform basis functions indicating example transform types. In Figure 14 where N represents the number of input pixels. For example, the selection of a transform type from multiple transform types can depend on the prediction type (one of intra prediction and inter prediction) and can depend on the intra prediction mode.
[0336] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0337] In addition, the transformer 106 may perform an inverse transform on the transform coefficients, which are the results of the transform. This inverse transform is also referred to as an Adaptive Secondary Transform (AST) or a Non-Separable Secondary Transform (NSST). For example, the transformer 106 performs the inverse transform on a per-sub-block basis (e.g., a 4×4 pixel sub-block) included in a transform coefficient block corresponding to an intra prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are typically signaled at the CU level. It should be noted that signaling of such information does not necessarily need to be performed at the CU level and may also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0338] The transformer 106 may employ separable transforms and non-separable transforms. A separable transform is a method in which the transform is performed multiple times by separately performing the transform in each of multiple directions according to the dimensions of the input. A non-separable transform is a method of performing a collective transform, in which two or more dimensions in a multi-dimensional input are collectively regarded as a single dimension.
[0339] In one example of a non-separable transform, when the input is a 4×4 pixel block, the 4×4 pixel block is considered as a single array containing 16 elements, and the transform applies a 16×16 transform matrix to this array.
[0340] In another example of a non-separable transform, an input block of 4×4 pixels is considered as a single array containing 16 elements, and then a transform that performs a given rotation on the array multiple times (a given transform of a hypercube) may be performed.
[0341] In the transform in the transformer 106, the type of transform of the transform basis function to be transformed into the frequency domain may be switched according to the region in the CU. Examples include a spatially varying transform (SVT).
[0342] Figure 15 is a conceptual diagram for showing an example of SVT.
[0343] In SVT, as Figure 15As shown, the CU is divided horizontally or vertically into two equal regions, and only one of the regions is transformed into the frequency domain. The transform basis type can be set for each region. For example, DST7 and DST8 are used. For example, in the two regions obtained by vertically dividing the CU into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, in the two regions, DST7 can be used for the region at position 1. Similarly, in the two regions obtained by horizontally dividing the CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Although in Figure 15 the example shown, one of the two regions in the CU is transformed while the other region is not transformed, each of the two regions can be transformed. Additionally, the splitting method can include not only splitting into two regions but also splitting into four regions. Furthermore, the splitting method can be more flexible. For example, the information indicating the splitting method can be encoded and signaled in the same way as the CU splitting. It should be noted that SVT can also be referred to as sub-block transform (SBT).
[0344] The AMT and EMT described above can be referred to as MTS (multiple transform selection). When applying MTS, transform types such as DST7, DCT8, etc. can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. There is another process called IMTS (implicit MTS) as a process for selecting the transform type to be used for an orthogonal transform to be performed without encoded index information. When applying IMTS, for example, when the CU has a rectangular shape, the orthogonal transform of the rectangular shape can be performed using DST7 (for the short side) and DST2 (for the long side). Additionally, for example, when the CU has a square shape, the orthogonal transform of the rectangular shape can be performed by using DCT2 when MTS is valid in the sequence and DST7 when MTS is invalid in the sequence. DCT2 and DST7 are just examples. Other transform types can be used, and the combination of the used transform types can also be changed to a different transform type combination. IMTS can be used only for intra-prediction blocks or for both intra-prediction blocks and inter-prediction blocks.
[0345] The three processes of MTS, SBT, and IMTS have been described above as selection processes for selectively switching the transform type used for orthogonal transformation. However, all three selection processes may be employed, or only some of the selection processes may be selectively employed. For example, it may be identified whether to employ one or more selection processes based on flag information in a header such as SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transformation of the CU is performed. It should be noted that the selection process for selectively switching the transform type may be a selection process different from the above three selection processes, or each of the three selection processes may be replaced by another process. Generally, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transformation of the entire CU and encoding information indicating the transform type used in the transformation. Function [2] is a function for performing the orthogonal transformation of the entire CU and determining the transform type based on a determined rule without encoding information indicating the transform type. Function [3] is a function for performing the orthogonal transformation of a partial region of the CU and encoding information indicating the transform type used in the transformation. Function [4] is a function for performing the orthogonal transformation of a partial region of the CU and determining the transform type based on a determined rule without encoding information indicating the transform type used in the transformation. The determined rule may be predetermined.
[0346] It should be noted that it may be determined for each processing unit whether to apply MTS, IMTS, and / or SBT. For example, it may be determined for each sequence, picture, block, slice, CTU, or CU whether to apply MTS, IMTS, and / or SBT.
[0347] It should be noted that the tool for selectively switching the transform type in the present invention may be described as a method, a selection process, or a process for selectively selecting a basis used in a transform process. Additionally, the tool for selectively switching the transform type may be described as a mode for adaptively selecting the transform type.
[0348] Figure 16 is a flowchart showing an example of the process performed by the transformer 106 and will be described for convenience with reference to Figure 7 for description.
[0349] For example, the transformer 106 determines whether to perform an orthogonal transform (step St_1). Here, when it is determined to perform an orthogonal transform (yes in step St_1), the transformer 106 selects a transform type for the orthogonal transform from among multiple transform types (step St_2). Next, the transformer 106 performs the orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transform type to the entropy encoder 110 to allow the entropy encoder 110 to encode this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transform (no in step St_1), the transformer 106 outputs information indicating that no orthogonal transform is performed to allow the entropy encoder 110 to encode this information (step St_5). It should be noted that whether to perform an orthogonal transform in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transform can also be performed using a defined transform type without encoding information indicating the transform type used in the orthogonal transform. The defined transform type can be predefined.
[0350] Figure 17 is a flowchart showing an example of the process performed by the transformer 106, and will be described for convenience with reference to Figure 7 For illustration. It should be noted that Figure 17 the example shown in Figure 16 is an example of an orthogonal transform in the case where the transform type used in the orthogonal transform is selectively switched (as in the case of the example shown in
[0351] As an example, the first transform type group may include DCT2, DST7, and DCT8. As another example, the second transform type group may include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group may partially overlap with each other, or may be completely different from each other.
[0352] The transformer 106 determines whether the transform size is less than or equal to a determined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the determined value (yes in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform types included in the first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one of the transform types included in the first transform type group to the entropy encoder 110 so as to allow the entropy encoder 110 to encode this information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to the predetermined value (no in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The determined value may be a threshold and may be a predetermined value.
[0353] In step Su_3, the information indicating the transform type used in the orthogonal transform may be information indicating a combination of the transform type to be vertically applied to the current block and the transform type to be horizontally applied to the current block. The first type group may include only one transform type, and the information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and the information indicating the transform type used for the orthogonal transform among one or more of the transform types included in the second transform type group may be encoded.
[0354] Alternatively, the transform type may be indicated based on the transform size without encoding the information indicating the transform type. It should be noted that such determination is not limited to the determination of whether the transform size is less than or equal to the determined value, and other processes for determining the transform type used in the orthogonal transform based on the transform size are also possible.
[0355] (Quantizer)
[0356] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in a determined scan order and quantizes the scanned transform coefficients based on the quantization parameter (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients of the current block (hereinafter also referred to as quantized coefficients) to the entropy encoder 110 and the inverse quantizer 112. The determined scan order may be predetermined.
[0357] The determined scan order is the order for quantizing / inverse quantizing the transform coefficients. For example, the determined scan order may be defined as ascending order of frequencies (from low frequency to high frequency) or descending order of frequencies (from high frequency to low frequency).
[0358] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, when the value of the quantization parameter increases, the quantization step also increases. In other words, when the value of the quantization parameter increases, the error of the quantized coefficient (quantization error) increases.
[0359] In addition, quantization matrices can be used for quantization. For example, multiple quantization matrices can be used corresponding to the frequency transform size (e.g., 4×4, 8×8), prediction mode (e.g., intra prediction, inter prediction), and pixel component (e.g., luminance, chrominance pixel components). It should be noted that quantization means digitizing the sampled values at determined intervals corresponding to determined levels. In the technical field, quantization can be referred to using other expressions, such as rounding and scaling, and rounding and scaling can be adopted. The determined intervals and determined levels can be predetermined.
[0360] The method of using a quantization matrix can include: the method of using the quantization matrix directly set on the encoder 100 side, and the method of using the quantization matrix (default matrix) set as the default. On the encoder 100 side, a quantization matrix suitable for the image characteristics can be set by directly setting the quantization matrix. However, this case may have the disadvantage of increasing the coding amount for encoding the quantization matrix. It should be noted that instead of directly using the default quantization matrix or the encoded quantization matrix, a quantization matrix for quantifying the current block can be generated based on the default quantization matrix or the encoded quantization matrix.
[0361] There is a method for quantifying high-frequency coefficients and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be regarded as equivalent to the method of using a quantization matrix (flat matrix) whose coefficients have the same value.
[0362] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level. The quantization matrix can be specified using, for example, the sequence parameter set (SPS) or the picture parameter set (PPS). The SPS includes the parameters for the sequence, and the PPS includes the parameters for the picture. Each of the SPS and PPS can be abbreviated as a parameter set.
[0363] When using a quantization matrix, the quantizer 108 scales the quantization width, which can be calculated based on, for example, the quantization parameter, using the value of the quantization matrix for each transform coefficient. The quantization process performed without using a quantization matrix can be a process of quantifying the transform coefficients according to the quantization width calculated based on, for example, the quantization parameter. It should be noted that in the quantization process performed without using any quantization matrix, the quantization width can be multiplied by a determined value common to all the transform coefficients in the block. The determined value can be predetermined.
[0364] Figure 18It is a block diagram showing an example of the functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a differential quantization parameter generator 108a, a prediction quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization executor 108e.
[0365] Figure 19 It is a flowchart showing an example of the quantization process performed by the quantizer 108, and for convenience, reference will be made to Figure 7 and 18 for description.
[0366] As an example, the quantizer 108 can perform quantization on each CU based on the Figure 19 flowchart shown. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generator 108c generates quantization parameters for the current block (step Sv_2), and stores the quantization parameters in the quantization parameter storage device 108d (step Sv_3).
[0367] Next, the quantization executor 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). The prediction quantization parameter generator 108b then obtains the quantization parameters of a processing unit different from the current block from the quantization parameter storage device 108d (step Sv_5). The prediction quantization parameter generator 108b generates prediction quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generator 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generator 108c and the prediction quantization parameters of the current block generated by the prediction quantization parameter generator 108b (step Sv_7). The differential quantization parameters can be generated by calculating the difference. The differential quantization parameter generator 108a outputs the differential quantization parameters to the entropy encoder 110 to allow the entropy encoder 110 to encode the differential quantization parameters (step Sv_8).
[0368] It should be noted that the differential quantization parameters can be encoded at, for example, the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At initialization, the initial values of the quantization parameters and the differential quantization parameters can be used to generate the quantization parameters.
[0369] It should be noted that the quantizer 108 can include multiple quantizers, and dependent quantization can be applied, where the transform coefficients are quantized using a quantization method selected from multiple quantization methods.
[0370] (Entropy Encoder)
[0371] Figure 20 is a block diagram showing an example of the functional configuration of the entropy encoder 110 according to an embodiment, and will be described for convenience with reference to Figure 7 The entropy encoder 110 generates a bitstream by performing entropy encoding on the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy encoding. More specifically, the entropy encoder 110 shown in the figure includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, in which a multi-level signal such as a quantized coefficient and a prediction parameter is transformed into a binary signal. Examples of binarization methods include truncated Rice binarization, exponential Golomb code, and fixed-length binarization. The context controller 110b derives a context value according to the characteristics of the syntax element or the surrounding state (i.e., the occurrence probability of the binary signal). Examples of methods for deriving the context value include bypassing, referring to syntax elements, referring to the upper and left adjacent blocks, referring to hierarchical information, etc. The binary arithmetic encoder 110c performs arithmetic coding on the binary signal using the derived context.
[0372] Figure 21 is a conceptual diagram for illustrating an example process of the CABAC process in the entropy encoder 110. First, initialization is performed in the entropy encoder 110 with CABAC. In the initialization, initialization in the binary arithmetic encoder 110c and setting of the initial context value are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic coding of a plurality of quantized coefficients in the CTU. Each time arithmetic coding is performed, the context controller 110b can update the context value. The context controller 110b can then save the context value as post-processing. For example, the saved context value can be used to initialize the context value of the next CTU.
[0373] (Inverse Quantizer)
[0374] The inverse quantizer 112 performs inverse quantization on the quantized coefficients input from the quantizer 108. More specifically, the inverse quantizer 112 performs inverse quantization on the quantized coefficients of the current block in a determined scan order. The inverse quantizer 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transformator 114. The determined scan order can be predetermined.
[0375] (Inverse Transformator)
[0376] The inverse transformer 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantizer 112. More specifically, the inverse transformer 114 restores the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the restored prediction residual to the adder 116.
[0377] It should be noted that since information is usually lost in quantization, the restored prediction residual does not match the prediction residual calculated by the subtractor 104. In other words, the restored prediction residual usually includes quantization errors.
[0378] (Adder)
[0379] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transformer 114 and the predicted image input from the prediction controller 128. Subsequently, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block may also be referred to as a local decoded block.
[0380] (Block Memory)
[0381] The block memory 118 is a storage device for storing blocks in the current picture used for, for example, intra prediction. More specifically, the block memory 118 stores the reconstructed image output from the adder 116.
[0382] (Frame Memory)
[0383] The frame memory 122 is a storage device for storing reference pictures used in inter prediction, for example, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.
[0384] (Loop Filter)
[0385] The loop filter 120 applies a loop filter to the reconstructed image output by the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter used in the encoding loop (in-loop filter). Examples of loop filters include, for example, an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, etc.
[0386] Figure 22 is a block diagram showing an example of the functional configuration of the loop filter 120 according to an embodiment. For example, as Figure 22As shown, the loop filter 120 includes a deblocking filter executor 120a, an SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs a deblocking filtering process on the reconstructed image. The SAO executor 120b performs an SAO process on the reconstructed image after the deblocking filtering process. The ALF executor 120c performs an ALF process on the reconstructed image after the SAO process. The ALF and the deblocking filter will be described in detail later. The SAO process is a process for improving image quality by reducing ringing (a phenomenon in which pixel values are distorted like waves around edges) and correcting pixel value deviations. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that in some embodiments, the loop filter 120 may not include Figure 22 all the constituent elements disclosed in Figure 22 and may include some constituent elements and may include additional elements. In addition, the loop filter 120 may be configured to perform the above processes in a processing order different from that disclosed in
[0387] (Loop filter>Adaptive loop filter)
[0388] In the ALF, a least-squares error filter for removing compression artifacts is applied. For example, for each 2×2 pixel sub-block in the current block, one filter selected from multiple filters is applied based on the local gradient direction and activity.
[0389] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of multiple classes (e.g., fifteen or twenty-five classes). The classification of the sub-block can be based on, for example, gradient directionality and activity. In the example, the class index C (e.g., C = 5D + A) is calculated or determined based on the gradient directionality D (e.g., 0 to 2 or 0 to 4) and the gradient activity A (e.g., 0 to 4). Then, based on the classification index C, each sub-block is classified into one of the multiple classes.
[0390] For example, the gradient directionality D is calculated by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, for example, the gradient activity A is calculated by adding the gradients in multiple directions and quantifying the addition result.
[0391] The filter to be used for each sub-block can be determined from multiple filters based on such classification results.
[0392] The filter shape to be used in the ALF is, for example, a circularly symmetric filter shape. Figures 23A to 23C is a conceptual diagram for showing an example of the filter shape used in the ALF.Figure 23A The figure illustrates a 5×5 diamond filter, Figure 23B the figure illustrates a 7×7 diamond filter, and Figure 23C the figure illustrates a 9×9 diamond filter. Information indicating the filter shape is typically signaled at the picture level. It should be noted that the signaling of such information indicating the filter shape does not necessarily need to be performed at the picture level and can be performed at another level (e.g., at the sequence level, slice level, tile level, CTU level, or CU level).
[0393] For example, the enabling or disabling of the ALF can be determined at the picture level or CU level. For example, a decision on whether to apply the ALF to the luminance can be made at the CU level, and a decision on whether to apply the ALF to the chrominance can be made at the picture level. Information indicating the enabling or disabling of the ALF is typically signaled at the picture level or CU level. It should be noted that the signaling of the information indicating the enabling or disabling of the ALF does not necessarily need to be performed at the picture level or CU level and can be performed at another level (e.g., at the sequence level, slice level, tile level, or CTU level).
[0394] In addition, as described above, one filter is selected from multiple filters, and the ALF process for the sub-block is performed. The set of coefficients for each of the multiple filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that the signaling of the set of coefficients does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0395] (Loop Filter > Cross-Component Adaptive Loop Filter)
[0396] Figure 23D is a conceptual diagram for showing an example process of the cross-component ALF (CC-ALF). Figure 23E is a conceptual diagram for showing an example of the filter shape used in the CC-ALF, such as Figure 23D the CC-ALF of Figure 23D and Figure 23E the example CC-ALF of operates by applying a linear diamond filter to the luminance channels of each chrominance component. For example, the filter coefficients can be transmitted in the APS, scaled by a factor of 2^10, and rounded for fixed-point representation. For example, in Figure 23D Y samples (the first component) are used for the CCALF of Cb and the CCALF of Cr (components different from the first component).
[0397] The application of the filter can be controlled on variable block sizes and signaled by a flag encoded with the context received for each sample block. The block size, together with the CC-ALF enable flag, can be received at the slice level for each chrominance component. CC-ALF can support various block sizes, e.g., 16×16 pixels, 32×32 pixels, 64×64 pixels, 128×128 pixels (in chrominance samples).
[0398] (Loop filter > Joint chrominance cross-component adaptive loop filter)
[0399] An example of joint chrominance - CCALF is shown in Figure 23F and Figure 23G . Figure 23F is a conceptual diagram for showing an example process of joint chrominance CCALF. Figure 23G is a table showing example weight index candidates. As shown, one CCALF filter is used to generate a CCALF filtered output as a chrominance refinement signal for one color component, while applying a weighted version of the same chrominance refinement signal to another color component. In this way, the complexity of the existing CCALF is reduced by approximately half. The weight value can be encoded as a sign flag and a weight index. The weight index (denoted as weight_index) can be encoded into 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight, which is a non-zero magnitude. For example, the magnitude of JcCcWeight can be determined as follows:
[0400] If weight_index is less than or equal to 4, then JcCcWeight equals weight_index >> 2;
[0401] Otherwise, JcCcWeight equals 4 / (weight_index – 4).
[0402] The block-level on / off control of ALF filtering for Cb and Cr can be separate. This is the same as in CCALF, and two separate groups of block-level on / off control flags can be encoded. Different from CCALF, the Cb, Cr on / off control block sizes here are the same, so only one block size variable can be encoded.
[0403] (Loop filter > Deblocking filter)
[0404] During the deblocking filtering process, the loop filter 120 performs a filtering process on the block boundaries in the reconstructed image to reduce the distortion occurring at the block boundaries.
[0405] Figure 24 shows the loop filter 120 acting as a deblocking filter (see Figure 7 and Figure 22)Block diagram of an example of the specific configuration of the deblocking filter actuator 120a.
[0406] The deblocking filter actuator 120a includes: a boundary determiner 1201; a filter determiner 1203; a filtering actuator 1205; a process determiner 1208; a filter characteristic determiner 1207; and switches 1202, 1204, and 1206.
[0407] The boundary determiner 1201 determines whether the pixel to be deblocked filtered (i.e., the current pixel) exists around the block boundary. The boundary determiner 1201 then outputs the determination result to the switch 1202 and the process determiner 1208.
[0408] In the case where the boundary determiner 1201 determines that the current pixel exists around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1204. In the opposite case (where the boundary determiner 1201 determines that the current pixel does not exist around the block boundary), the switch 1202 outputs the unfiltered image to the switch 1206. Note that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.
[0409] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0410] In the case where the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the filtering actuator 1205. In the opposite case (where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel), the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the switch 1206.
[0411] When the unfiltered image is obtained through the switches 1202 and 1204, the filtering actuator 1205 performs deblocking filtering on the current pixel with the filtering characteristics determined by the filter characteristic determiner 1207. The filtering actuator 1205 then outputs the filtered pixel to the switch 1206.
[0412] Under the control of the process determiner 1208, the switch 1206 selectively outputs one of the pixels that have not been deblocked filtered and the pixels that have been deblocked filtered by the filtering actuator 1205.
[0413] The processing determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary and when the filter determiner 1203 has determined that deblocking filtering of the current pixel is to be performed, the processing determiner 1208 causes the switch 1207 to output the pixel for which deblocking filtering has been performed. In addition, except for the above cases, the processing determiner 1208 causes the switch 1206 to output the pixel for which deblocking filtering has not been performed. By repeating the output of the pixel in this way, the filtered image is output from the switch 1206. It should be noted that Figure 24 The configuration shown in
[0414] Figure 25 is a conceptual diagram for showing an example of a deblocking filter having a symmetric filtering characteristic with respect to the block boundary.
[0415] During the deblocking filtering process, pixel values and quantization parameters can be used to select one of two deblocking filters (i.e., a strong filter and a weak filter) having different characteristics. In the case of the strong filter, when pixels p0 to p2 and pixels q0 to q2 exist across the block boundary, as Figure 25 shown, by performing calculations according to the following expressions, for example, the pixel values of the corresponding pixels q0 to q2 are changed to pixel values q'0 to q'2.
[0416] q'0 = (p1 + 2 × p0 + 2 × q0 + 2 × q1 + q2 + 4) / 8
[0417] q'1 = (p0 + q0 + q1 + q2 + 2) / 4
[0418] q'2 = (p0 + q0 + q1 + 3 × q2 + 2 × q3 + 4) / 8
[0419] It should be noted that in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the corresponding pixels p0 to p2 and pixels q0 to q2. In addition, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of the pixel q2 with respect to the block boundary. In addition, on the right side of each expression, the coefficient multiplied by the corresponding pixel value of the pixel to be used for deblocking filtering is the filter coefficient.
[0420] In addition, in the deblocking filtering, clipping can be performed so that the change in the calculated pixel value does not exceed a threshold value. For example, during the clipping process, the pixel value calculated according to the above expression can be clipped to a value obtained according to "calculated pixel value ± 2 × threshold value" (using the threshold value determined based on the quantization parameter). In this way, over-smoothing can be prevented.
[0421] Figure 26 is a conceptual diagram for showing a block boundary on which a deblocking filtering process is performed. Figure 27 is a conceptual diagram for showing an example of a boundary strength (Bs) value.
[0422] The block boundary on which the deblocking filtering process is performed is, for example, a boundary between a CU, a Pu, or a TU having 8×8 pixels, as Figure 26 shown. The deblocking filtering process can be performed, for example, in units of four rows or four columns. First, as Figure 27 shown for block P and block Q ( Figure 26 shown), a boundary strength (Bs) value is determined.
[0423] According to the Figure 27 Bs value in, it can be determined whether to perform a deblocking filtering process on a block boundary belonging to the same image with different strengths. When the Bs value is 2, a deblocking filtering process for a chrominance signal is performed. When the Bs value is 1 or greater and a determined condition is satisfied, a deblocking filtering process for a luminance signal is performed. The determined condition can be predetermined. Note that the conditions for determining the Bs value are not limited to those Figure 27 shown in, and the Bs value can be determined based on another parameter.
[0424] (Predictor (intra predictor, inter predictor, prediction controller))
[0425] Figure 28 is a flowchart showing an example of a process performed by the predictor of the encoder 100. It should be noted that the predictor includes all or part of the following constituent elements: intra predictor 124; inter predictor 126; and prediction controller 128. The prediction executor includes, for example, intra predictor 124 and inter predictor 126.
[0426] The predictor generates a predicted image of a current block (step Sb_1). This predicted image can also be referred to as a prediction signal or a predicted block. It should be noted that the prediction signal is, for example, an intra prediction image (image prediction signal) or an inter prediction image (inter prediction signal). The predictor uses a reconstructed image that has been obtained through the generation of a prediction image, the generation of a prediction residual, the generation of quantized coefficients, the recovery of the prediction residual, and the addition to the prediction image by another block, to generate a predicted image of the current block.
[0427] The reconstructed image can be, for example, an image in a reference picture, or an image of an encoded block (i.e., the above-mentioned other block) in the current picture, where the current picture is a picture including the current block. The encoded block in the current picture is, for example, an adjacent block of the current block.
[0428] Figure 29 is a flowchart showing another example of a process performed by the predictor of the encoder 100.
[0429] The predictor generates a prediction image using a first method (step Sc_1a), generates a prediction image using a second method (step Sc_1b), and generates a prediction image using a third method (step Sc_1c). The first method, the second method, and the third method may be mutually different methods for generating a prediction image. Each of the first to third methods may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.
[0430] Next, the prediction processor evaluates the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates a cost C for the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the prediction images by comparing the costs C of the prediction images. It should be noted that the cost C can be calculated, for example, according to the expression of the R-D optimization model, such as C = D + λ × R. In this expression, D represents the compression artifact of the prediction image and is expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. In addition, R represents the bit rate of the stream. In addition, λ represents, for example, a multiplier according to the Lagrange method multiplier.
[0431] Then, the predictor selects one of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects the method or mode for obtaining the final prediction image. For example, the predictor selects the prediction image with the minimum cost C based on the cost C calculated for the prediction image. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 may be performed based on the parameters used in the encoding process. The encoder 100 may transform the information for identifying the selected prediction image, method, or mode into a stream. This information may be, for example, a flag or the like. In this way, the decoder 200 can generate a prediction image based on this information according to the method or mode selected by the encoder 100. It should be noted that in Figure 29 the example shown, after generating the prediction image using the corresponding method, the predictor selects any prediction image. However, the predictor may select a method or mode based on the parameters used in the above encoding process before generating the prediction image, and may generate a prediction image according to the selected method or mode.
[0432] For example, the first method and the second method may be intra-frame prediction and inter-frame prediction, respectively, and the predictor may select the final prediction image of the current block from the prediction images generated according to the prediction method.
[0433] Figure 30 is a flowchart showing another example of the process performed by the predictor of the encoder 100.
[0434] First, the predictor generates a predicted image using intra prediction (step Sd_1a) and generates a predicted image using inter prediction (step Sd_1b). It should be noted that the predicted image generated by intra prediction is also referred to as an intra-predicted image, and the predicted image generated by inter prediction is also referred to as an inter-predicted image.
[0435] Next, the predictor evaluates each of the intra-predicted image and the inter-predicted image (step Sd_2). The above cost C can be used in the evaluation. The predictor can then select, from the intra-predicted image and the inter-predicted image, the predicted image for which the minimum cost C has been calculated as the final predicted image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the predicted image for the current block is selected.
[0436] The prediction processor then selects, from the intra-predicted image and the inter-predicted image, the predicted image for which the minimum cost C has been calculated as the final predicted image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the predicted image for the current block is selected.
[0437] (Intra Predictor)
[0438] The intra predictor 124 generates a prediction signal (i.e., an intra-predicted image) by referring to one or more blocks in the current picture and stored in the block memory 118 to perform intra prediction of the current block (also referred to as prediction within a frame). More specifically, by referring to the pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block i, the intra predictor 124 generates an intra-predicted image and then outputs the intra-predicted image to the prediction controller 128.
[0439] For example, the intra predictor 124 performs intra prediction by using one of a plurality of defined intra prediction modes. The intra prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes. The defined modes can be predefined.
[0440] One or more non-directional prediction modes include, for example, the planar prediction mode and the DC prediction mode defined in the H.265 / High Efficiency Video Coding (HEVC) standard.
[0441] The plurality of directional prediction modes include, for example, thirty-three directional prediction modes defined in the H.265 / HEVC standard. It should be noted that in addition to the thirty-three directional prediction modes, the plurality of directional prediction modes can also include thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31It is a conceptual diagram for showing a total of sixty-seven intra prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra prediction. Solid arrows indicate thirty-three directions defined in the H.265 / HEVC standard, and dashed arrows indicate an additional thirty-two directions ( Figure 31 The two non-directional prediction modes are not shown in
[0442] In various processing examples, a luminance block can be referred to in the intra prediction of a chrominance block. In other words, the chrominance component of the current block can be predicted based on the luminance component of the current block. This intra prediction is also called cross-component linear model (CCLM) prediction. An intra prediction mode of a chrominance block that references such a luminance block (also called, for example, a CCLM mode) can be added as one of the intra prediction modes of the chrominance block.
[0443] The intra predictor 124 can correct the pixel values of the intra prediction based on the horizontal / vertical reference pixel gradients. The intra prediction accompanied by such correction is also called position-dependent intra prediction combination (PDPC). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0444] Figure 32 It is a flowchart showing an example of the process executed by the intra predictor 124.
[0445] The intra predictor 124 selects one intra prediction mode from a plurality of intra prediction modes (step Sw_1). The intra predictor 124 then generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra prediction modes. For example, two of the six intra prediction modes can be the planar mode and the DC prediction mode, and the other four modes can be directional prediction modes. The intra predictor 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0446] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the intra prediction mode selected among these MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as prediction parameters by the entropy encoder 110.
[0447] When it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flag. The intra predictor 124 then generates information indicating the intra prediction mode selected from among at least one intra prediction mode not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded by the entropy encoder 110 as prediction parameters. The information indicating the intra prediction mode indicates any one of, for example, 0 to 60.
[0448] (Inter-frame predictor)
[0449] The inter-frame predictor 126 performs inter-frame prediction of the current block by referring to one or more blocks in a reference picture (also referred to as inter-frame prediction), and generates a predicted picture (inter-frame predicted picture), where the reference picture is different from the current picture and is stored in the frame memory 122. Inter-frame prediction is performed in units of the current block or a current sub-block in the current block (e.g., a 4×4 block). The sub-block is included in the block and is a unit smaller than the block. The size of the sub-block can be in the form of a slice, a brick, a picture, etc.
[0450] For example, the inter-frame predictor 126 performs motion estimation in the reference picture of the current block or current sub-block and finds the reference block or reference sub-block that best matches the current block or current sub-block. The inter-frame predictor 126 then obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or reference sub-block to the current block or sub-block. The inter-frame predictor 126 generates an inter-frame predicted picture of the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter-frame predictor 126 outputs the generated inter-frame predicted picture to the prediction controller 128.
[0451] The motion information used in motion compensation can be signaled as an inter-frame prediction signal in various forms. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a motion vector predictor can be signaled.
[0452] (Reference picture list)
[0453] Figure 33 is a conceptual diagram for showing an example of a reference picture. Figure 34 is a conceptual diagram for showing an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. It should be noted that in Figure 33In this figure, each rectangle represents a picture, each arrow represents a picture reference relationship, the horizontal axis represents time, I, P, and B in the rectangle represent intra-predicted pictures, single-predicted pictures, and bi-predicted pictures respectively, and the numbers in the rectangle represent the decoding order. As Figure 33 shown, the decoding order of the pictures is I0, P1, B2, B3, B4, and the display order of the pictures is I0, B3, B2, B4, P1. As Figure 34 shown, the reference picture list is a list representing reference picture candidates. For example, a picture (or slice) may include at least one reference picture list. For example, one reference picture list is used when the current picture is a single-predicted picture, and two reference picture lists are used when the current picture is a bi-predicted picture. In Figure 33 and Figure 34 the example of, picture B3 as the current picture currPic has two reference picture lists, namely the L0 list and the L1 list. When the current picture currPic is picture B3, the reference picture candidates for the current picture currPic are I0, P1, B2, and the reference picture lists (i.e., the L0 list and the L1 list) indicate these pictures. The inter-frame predictor 126 or the prediction controller 128 specifies which picture in each reference picture list is to be actually referenced in the form of a reference picture index refidxLx. In Figure 34 this figure, reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.
[0454] Such reference picture lists can be generated for each unit such as a sequence, a picture, a slice, a block, a CTU, or a CU. Additionally, among the reference pictures indicated in the reference picture list, the reference picture indices indicating the reference pictures to be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, block level, CTU level, or CU level. Furthermore, a common reference picture list can be used in multiple inter-frame prediction modes.
[0455] (Basic Process of Inter-Frame Prediction)
[0456] Figure 35 is a flowchart showing an example basic processing flow of inter-frame prediction processing.
[0457] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Then, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Se_4).
[0458] Here, in the generation of a predicted image, the inter-frame predictor 126 generates a predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Further, in the determination of the MV, the inter-frame predictor 126 determines the MV by selecting a motion vector candidate (MV candidate) (step Se_1) and deriving the MV (step Se_2). The selection of the MV candidate is performed, for example, by the inter-frame predictor 126 generating an MV candidate list and selecting at least one MV candidate from the MV candidate list. It should be noted that MVs derived in the past may be added to the MV candidate list. Alternatively, in the derivation of the MV, the inter-frame predictor 126 may also select at least one MV candidate from at least one MV candidate and determine the selected at least one MV candidate as the MV of the current block. Alternatively, the inter-frame predictor 126 may determine the MV of the current block by performing an estimation in the reference picture area specified by each of the selected at least one MV candidates. It should be noted that the estimation in the reference picture area may be referred to as motion estimation.
[0459] In addition, although steps Se_1 to Se_3 are performed by the inter-frame predictor 126 in the above example, processes such as step Se_1 and step Se_2 may be performed by another component included in the encoder 100, for example.
[0460] It should be noted that an MV candidate list may be generated for each process in the inter-frame prediction mode, or a common MV candidate list may be used in a plurality of inter-frame prediction modes. The processes in steps Se_3 and Se_4 respectively correspond to Figure 9 steps Sa_3 and Sa_4 shown in Figure 30 The process in step Se_3 corresponds to the process in step Sd_1b in
[0461] (Motion Vector Derivation Flowchart)
[0462] Figure 36 is a flowchart showing an example of the process of deriving a motion vector.
[0463] The inter-frame predictor 126 may derive the MV of the current block in a mode for encoding motion information (e.g., MV). In this case, for example, the motion information may be encoded as a prediction parameter and signaled. In other words, the encoded motion information is included in the stream.
[0464] Alternatively, the inter-frame predictor 126 may derive the MV in a mode in which the motion information is not encoded. In this case, the motion information is not included in the stream.
[0465] Here, the MV derivation mode may include a conventional inter mode, a conventional merge mode, a FRUC mode, an affine mode, etc., which will be described later. The modes for encoding motion information in the modes include a conventional inter mode, a conventional merge mode, an affine mode (specifically, an affine inter mode and an affine merge mode), etc. It should be noted that the motion information may include not only the MV but also the motion vector predictor selection information which will be described later. The modes that do not encode motion information include a FRUC mode, etc. The inter predictor 126 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.
[0466] Figure 37 is a flowchart showing another example of the derivation of a motion vector.
[0467] The inter predictor 126 may derive the MV of the current block in a mode that encodes the MV difference. In this case, for example, the MV difference may be encoded as a prediction parameter and signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. It should be noted that the MV predictor is a motion vector predictor.
[0468] Alternatively, the inter predictor 126 may derive the MV in a mode that does not encode the MV difference. In this case, the encoded MV difference is not included in the stream.
[0469] Here, as described above, the MV derivation mode includes a conventional inter mode, a conventional merge mode, a FRUC mode, an affine mode, etc., which will be described later. The modes that encode the MV difference in the modes include a conventional inter mode, an affine mode (specifically, an affine inter mode), etc. The modes that do not encode the MV difference include a FRUC mode, a conventional merge mode, an affine mode (specifically, an affine merge mode), etc. The inter predictor 126 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.
[0470] (Motion Vector Derivation Mode)
[0471] Figure 38A and Figure 38B is a conceptual diagram for showing an example classification of the modes for MV derivation. For example, as Figure 38A shown, according to whether motion information is encoded and whether the MV difference is encoded, the MV derivation mode is roughly classified into three modes. The three modes are an inter mode, a merge mode, and a frame rate up-conversion (FRUC) mode. The inter mode is a mode that performs motion estimation and encodes motion information and the MV difference. For example, as Figure 38BAs shown, the inter-frame mode includes the affine inter-frame mode and the normal inter-frame mode. The merge mode is a mode in which motion estimation is not performed, an MV is selected from coded neighboring blocks, and the MV of the current block is derived using this MV. The merge mode is a mode that basically encodes motion information without encoding the MV difference. For example, as Figure 38B shown, the merge mode includes the normal merge mode (also referred to as the regular merge mode or the normal merge mode), the merge with motion vector difference (MMVD) mode, the combined inter-frame merge / intra prediction (CIIP) mode, the triangular mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the modes included in the merge mode, the MV difference is exceptionally encoded. It should be noted that the affine merge mode and the affine inter-frame mode are modes included in the affine mode. The affine mode is a mode used to derive the MV of each of a plurality of sub-blocks included in the current block as the MV of the current block under the assumption of an affine transformation. The FRUC mode is a mode that is used to derive the MV of the current block by performing estimation between coding regions, and neither encodes motion information nor encodes any MV difference. It should be noted that the corresponding modes will be described in more detail later.
[0472] It should be noted that Figure 38A and Figure 38B the classification of the modes shown in
[0473] (MV derivation > normal inter-frame mode)
[0474] The normal inter-frame mode is an inter-frame prediction mode that is used to derive the MV of the current block from a reference picture region specified by an MV candidate based on a block similar to the image of the current block. In this normal inter-frame mode, the MV difference is encoded.
[0475] Figure 39 is a flowchart showing an example of the inter-frame prediction process in the normal inter-frame mode.
[0476] First, the inter-frame predictor 126 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0477] Next, the inter-frame predictor 126 extracts N (an integer of 2 or greater) MV candidates as motion vector predictor candidates (also referred to as MV predictor candidates) from the multiple MV candidates obtained in step Sg_1 according to the determined priority order (step Sg_2). It should be noted that the priority order can be predetermined for each of the N MV candidates.
[0478] Next, the inter-frame predictor 126 selects one motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor for the current block (also referred to as the MV predictor) (step Sg_3). At this time, the inter-frame predictor 126 encodes the motion vector predictor selection information for identifying the selected motion vector predictor in the stream. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0479] Next, the inter-frame predictor 126 derives the MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as the MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130. It should be noted that the encoded reference picture is a picture including multiple blocks that have been reconstructed after being encoded.
[0480] Finally, by performing motion compensation on the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image of the current block (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a slice, the inter-frame prediction of the slice using the conventional inter-frame mode ends. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a picture, the inter-frame prediction of the picture using the conventional inter-frame mode ends. It should be noted that in steps Sg_1 to Sg_5, not all blocks included in the slice can undergo these processes, and when some blocks undergo the processes, the inter-frame prediction of the slice using the conventional inter-frame mode can end. This also applies to the processes in steps Sg_1 to Sg_5. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the conventional inter-frame mode can end.
[0481] It should be noted that the predicted image is an inter-frame prediction signal as described above. In addition, information indicating the inter-frame prediction mode (the conventional inter-frame mode in the above example) for generating the predicted image is encoded as a prediction parameter in the encoded signal.
[0482] Note that the MV candidate list can also be used as a list used in another mode. In addition, the processes related to the MV candidate list can be applied to the processes related to the list for use in another mode. The processes related to the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering the MV candidates, or deleting the MV candidates.
[0483] (MV Derivation > Regular Merge Mode)
[0484] The regular merge mode is an inter prediction mode for deriving an MV by selecting an MV candidate from the MV candidate list as the MV of the current block. Note that the regular merge mode is a type of merge mode and can be abbreviated as the merge mode. In this embodiment, the regular merge mode and the merge mode are distinguished, and the merge mode is used in a broader sense.
[0485] Figure 40 It is a flowchart showing an example of inter prediction in the regular merge mode.
[0486] First, the inter predictor 126 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sh_1). In other words, the inter predictor 126 generates an MV candidate list.
[0487] Next, the inter predictor 126 selects one MV candidate from the plurality of MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter predictor 126 encodes MV selection information for identifying the selected MV candidate in the stream. In other words, the inter predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0488] Finally, by performing motion compensation for the current block using the derived MV and the coded reference picture, the inter predictor 126 generates a predicted image of the current block (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a slice, the inter prediction of the slice using the regular merge mode ends. In addition, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a picture, the inter prediction of the picture using the regular merge mode ends. Note that not all blocks included in the slice can undergo the processes in steps Sh_1 to Sh_3, and when some blocks undergo the processes, the inter prediction of the slice using the regular merge mode can end. This also applies to the processes in steps Sh_1 to Sh_3. When the processes are performed for some blocks in a picture, the inter prediction of the picture using the regular merge mode can be completed.
[0489] In addition, information included in the encoded signal and representing an inter-frame prediction mode (in the above example, the normal merge mode) used to generate a predicted image is encoded, for example, as prediction parameters in the stream.
[0490] Figure 41 is a conceptual diagram showing an example of the motion vector derivation process of the current picture by the normal merge mode.
[0491] First, the inter-frame predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of a plurality of encoded blocks located spatially around the current block; temporally adjacent MV candidates, which are MVs of surrounding blocks onto which the position of the current block in the encoded reference picture is projected; combined MV candidates, which are MVs generated by combining the MV values of spatially adjacent MV predictors and the MV values of temporally adjacent MV predictors; and zero MV candidates, which are MVs having a zero value.
[0492] Next, the inter-frame predictor 126 selects one MV candidate from among the plurality of MV candidates registered in the MV candidate list and determines the MV candidate as the MV of the current block.
[0493] In addition, the entropy encoder 110 writes and encodes in the stream a merge_idx, which is a signal indicating which MV candidate has been selected.
[0494] It should be noted that the MV candidates registered in the Figure 41 MV candidate list described are examples. The number of MV candidates may be different from the number of MV candidates in the figure, and the MV candidate list may be configured in such a way that some types of MV candidates in the figure may not be included, or one or more types of MV candidates other than the types of MV candidates in the figure may be included.
[0495] By using the MV of the current block derived by the normal merge mode to perform dynamic motion vector refresh (DMVR) to be described later, the final MV can be determined. It should be noted that in the normal merge mode, the motion information is encoded and no MV difference is encoded. In the MMVD mode, one MV candidate is selected from the MV candidate list, just as in the case of the normal merge mode, and the MV difference is encoded. As Figure 38B shown, MMVD can be classified as a merge mode together with the normal merge mode. It should be noted that the MV difference in the MMVD mode does not always need to be the same as the MV difference used for the inter-frame mode. For example, the MV difference derivation in the MMVD mode may be a process that requires less processing amount than the MV difference derivation in the inter-frame mode.
[0496] In addition, a combined inter-frame merging / intra-frame prediction (CIIP) mode can be performed. This mode is used to overlap the prediction image generated in inter-frame prediction and the prediction image generated in intra-frame prediction to generate the prediction image of the current block.
[0497] It should be noted that the MV candidate list can be referred to as the candidate list. Additionally, merge_idx is the MV selection information.
[0498] (MV derivation > HMVP mode)
[0499] Figure 42 is a conceptual diagram showing an example of the MV derivation process for the current picture using the HMVP merge mode.
[0500] In the regular merge mode, the MV of a CU, for example, serving as the current block, is determined by selecting one MV candidate from the MV list generated from reference coded blocks (e.g., CUs). Here, another MV candidate can be registered in the MV candidate list. The mode of registering such another MV candidate is called the HMVP mode.
[0501] In the HMVP mode, a first-in, first-out (FIFO) server for HMVP is used to manage MV candidates, separate from the MV candidate list of the regular merge mode.
[0502] In the FIFO buffer, motion information such as the MV of a block processed in the past is stored latest first. In managing the FIFO buffer, each time a block is processed, the MV of the latest block (i.e., the CU processed immediately before) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the CU processed earliest) is deleted from the FIFO buffer. In Figure 42 the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the oldest MV.
[0503] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is an MV different from all the MV candidates already registered in the MV candidate list of the regular merge mode starting from HMVP1. When it is determined that the MV is different from all the MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for the regular merge mode as an MV candidate. At this time, one or more MV candidates in the FIFO buffer can be registered (added to the MV candidate list).
[0504] By using the HMVP mode in this way, not only can the MVs of blocks adjacent to the current block in space or time be added, but also the MVs of blocks processed in the past can be added. As a result, the variation of the MV candidates in the regular merge mode is expanded, which increases the possibility of improving the coding efficiency.
[0505] Note that the MV can be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer can include not only the MV value, but also reference picture information, reference direction, number of pictures, etc. Additionally, the block can be, for example, a CU.
[0506] Note that Figure 42 the MV candidate list and the FIFO buffer shown in are examples. The size of the MV candidate list and the FIFO buffer can be different from that in Figure 42 or can be configured to register MV candidates in an order different from that in Figure 42 Moreover, the processes described herein can be common between the encoder 100 and the decoder 200.
[0507] Note that the HMVP mode can be applied to modes other than the regular merge mode. For example, motion information such as the MV of a block that was previously processed in the affine mode can also be stored most recently and can be used as an MV candidate, which can promote better efficiency. The mode obtained by applying the HMVP mode to the affine mode can be referred to as the history affine mode.
[0508] (MV derivation > FRUC mode)
[0509] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, on the decoder side, motion estimation is performed without using any pixel values in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values in the current block include frame rate up-conversion (FRUC) mode, pattern matching motion vector derivation (PMMVD) mode, etc.
[0510] Figure 43 An example of the FRUC process in flowchart form is shown in. First, each of the coded blocks that are spatially or temporally adjacent to the current block is indicated by referring to the MV as a list of MV candidates (this list can be the MV candidate list and can also be used as the MV candidate list for the regular merge mode) (step Si_1).
[0511] Next, the best MV candidate is selected from among the plurality of MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation value of the corresponding MV candidate included in the MV candidate list is calculated, and one MV candidate is selected based on the evaluation value. Based on the selected motion vector candidate, the motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (the best MV candidate) is directly derived as the motion vector of the current block. Additionally, for example, pattern matching in the surrounding area of the position in the reference picture corresponding to the selected motion vector candidate can be used to derive the motion vector of the current block. In other words, pattern matching and evaluation value estimation can be performed in the surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces a better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, the update of the motion vector that produces a better evaluation value may not be performed.
[0512] Finally, the inter-frame predictor 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a picture, the inter-frame prediction of the picture using the FRUC mode ends. Note that not all blocks included in a slice undergo the processes in steps Si_1 to Si_5, and when some blocks undergo the processes, the inter-frame prediction of the slice using the FRUC mode can end. When the processes in steps Si_1 to Si_5 are performed for some blocks included in a picture in a similar manner, the inter-frame prediction of the picture using the FRUC mode can end.
[0513] Similar processes can be performed on a sub-block basis.
[0514] The evaluation value can be calculated according to various methods. For example, a comparison is made between the reconstructed image in the region in the reference picture corresponding to the motion vector and the reconstructed image in a determined region (this region can be, for example, a region in another reference picture or a region in an adjacent block of the current picture, as described below). The determined region can be predetermined.
[0515] The difference between the pixel values of the two reconstructed images can be used for the evaluation value of the motion vector. Note that information other than the difference can be used to calculate the evaluation value.
[0516] Next, an example of pattern matching will be described in detail. First, one MV candidate included in the MV candidate list (e.g., the merge list) is selected as the estimated starting point through pattern matching. For example, as the pattern matching, the first pattern matching or the second pattern matching can be used. The first pattern matching and the second pattern matching can be respectively referred to as bilateral matching and template matching.
[0517] (MV Derivation>FRUC>Bilateral Matching)
[0518] In the first pattern matching, pattern matching is performed between two blocks that are located along the motion trajectory of the current block and are included in two different reference pictures. Therefore, in the first pattern matching, the region in the other reference picture along the motion trajectory of the current block is used as the determined region for calculating the evaluation value of the above candidate. The determined region can be predetermined.
[0519] Figure 44 is a conceptual diagram showing an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As Figure 44 shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the pair of the best matches among the pairs in two blocks that are included in two different reference pictures (Ref0, Ref1) and are located along the motion trajectory of the current block (Cur block). More specifically, for the current block, the difference between the reconstructed image at the specified position in the first coded reference picture (Ref0) specified by the MV candidate and the reconstructed image at the specified position in the second coded reference picture (Ref1) specified by the symmetric MV obtained by scaling the MV candidate by the display time interval is derived, and the obtained difference value is used to calculate the evaluation value. The MV candidate that produces the best evaluation value and may produce good results can be selected from multiple MV candidates as the final MV.
[0520] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) specifying the two reference blocks are proportional to the time distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the time distances from the current picture to the corresponding two reference pictures are equal to each other, mirror-symmetric bidirectional motion vectors are derived in the first pattern matching.
[0521] (MV Derivation>FRUC>Template Matching)
[0522] In the second mode matching (template matching), pattern matching is performed between a block in a reference picture and a template in the current picture (the template is a block adjacent to the current block in the current picture (the adjacent block is, for example, an upper and / or left adjacent block)). Therefore, in the second mode matching, the adjacent block of the current block in the current picture is used as a determined region for calculating the evaluation value of the above-mentioned MV candidate.
[0523] Figure 45 is a conceptual diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As Figure 45 shown, in the second mode matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference picture (Ref0) that best matches the adjacent block of the current block in the current picture (Cur Pic). More specifically, the difference between the reconstructed image in the coded region adjacent to the left and above or left or above and the reconstructed image in the corresponding region in the coded reference picture (Ref0) specified by the MV candidate is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.
[0524] This information indicating whether to apply the FRUC mode (for example, referred to as the FRUC flag) can be signaled at the CU level. In addition, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the applicable pattern matching method (for example, the first mode matching or the second mode matching) can be signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (for example, the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0525] (MV Derivation > Affine Mode)
[0526] The affine mode is a mode that generates an MV using an affine transformation. For example, the MV can be derived for each sub-block based on the motion vectors of multiple adjacent blocks. This mode is also referred to as the affine motion compensation prediction mode.
[0527] Figure 46A is a conceptual diagram showing an example of MV derivation for each sub-block based on the motion vectors of multiple adjacent blocks. In Figure 46A it, the current block includes, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the upper left control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (vx , v y )。
[0528] [Mathematical Expression 1]
[0529]
[0530] Here, x and y represent the horizontal and vertical positions of the sub - block respectively, and w represents a determined weighting coefficient. The determined weighting coefficient can be pre - determined.
[0531] This information indicating the affine mode (e.g., called an affine flag) can be signaled at the CU level. Note that the signaling of the information indicating the affine mode does not necessarily need to be performed at the CU level and can be performed at another level (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub - block level).
[0532] In addition, the affine mode can include several modes for different methods of deriving the motion vectors at the upper - left and upper - right control points. For example, the affine mode includes two modes: the affine inter - frame mode (also called the affine regular inter - frame mode) and the affine merge mode.
[0533] (MV Derivation > Affine Mode)
[0534] Figure 46B is a conceptual diagram showing an example of MV derivation in units of sub - blocks in the affine mode using three control points. In Figure 46B , the current block includes, for example, sixteen 4×4 blocks. Here, the motion vector V0 at the upper - left control point in the current block is derived based on the motion vectors of adjacent blocks. Here, the motion vector V1 at the upper - right control point in the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V2 at the lower - left control point of the current block is derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v x , v y ) of the corresponding sub - blocks in the current block can be derived.
[0535] [Mathematical Expression 2]
[0536]
[0537] Here, x and y represent the horizontal and vertical positions of the sub - block respectively, and w and h can be weighting coefficients, which can be pre - determined weighting coefficients. In an embodiment, w can represent the width of the current block, and h can represent the height of the current block.
[0538] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0539] In addition, this affine mode using three control points may include different methods for deriving motion vectors at the upper left, upper right and lower left control points. For example, as in the case of the affine mode using two control points, the affine mode using three control points may include two modes, the affine inter-frame mode and the affine merge mode.
[0540] Note that in the affine mode, the size of each sub-block included in the current block may not be limited to 4×4 pixels, and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.
[0541] (MV Derivation > Affine Mode > Control Points)
[0542] Figure 47A , Figure 47B and Figure 47C is a conceptual diagram for illustrating an example of MV derivation at a control point in an affine mode.
[0543] like Figure 47A As shown, in the affine mode, for example, based on a plurality of motion vectors corresponding to blocks encoded according to the affine mode among the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, a motion vector predictor at a corresponding control point of the current block is calculated. More specifically, the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) are checked in the order listed, and the first valid block encoded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on a plurality of motion vectors corresponding to the identified blocks.
[0544] For example, Figure 47B As shown, when block A adjacent to the left side of the current block has been encoded according to the affine mode using two control points, motion vectors v3 and v4 projected at the upper left corner position and the upper right corner position of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated.
[0545] For example, Figure 47CAs shown, when the block A adjacent to the left side of the current block has been encoded according to the affine mode using three control points, the motion vectors v3, v4, and v5 projected at the upper left corner position, upper right corner position, and lower left corner position of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block are calculated.
[0546] Figures 47A to 47C The MV derivation method shown above can be used in the MV derivation at each control point of the current block in Figure 50 the step Sk_1 shown above, or can be used for the MV predictor derivation at each control point of the current block in Figure 51 the step Sj_1 shown later.
[0547] Figure 48A and Figure 48B are conceptual diagrams for showing examples of MV derivation at control points in the affine mode.
[0548] Figure 48A is a conceptual diagram for showing an exemplary affine mode using two control points.
[0549] In the affine mode, as Figure 48A shown, the MV selected from the MVs at the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block.
[0550] Figure 48B is a conceptual diagram for showing an exemplary affine mode using three control points.
[0551] In the affine mode, as Figure 48B shown, the MV selected from the MVs at the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block. Additionally, the MV selected from the MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 at the lower left corner control point of the current block.
[0552] Note that Figure 48A and Figure 48B the MV derivation method shown above can be used in the MV derivation at each control point of the current block in Figure 50 shown in Sk_1 described later, or can be used forFigure 51 Derivation of the MV predictor at each control point of the current block in step Sj_1 shown in
[0553] Here, when affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level, the number of control points of the coded block and the number of control points of the current block can be different from each other.
[0554] Figure 49A and Figure 49B are conceptual diagrams showing examples of methods for MV derivation at control points when the number of control points of the coded block and the number of control points of the current block are different from each other.
[0555] For example, as Figure 49A shown, the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block has been coded according to an affine mode using two control points. In this case, the motion vectors v3 and v4 projected at the upper left position and upper right position in the coded block including block A are derived. Then, the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3 and v4. In addition, the motion vector v2 at the lower left control point is calculated based on the derived motion vectors v0 and v1.
[0556] For example, as Figure 49B shown, the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left side of the current block has been coded according to an affine mode using three control points. In this case, the motion vectors v3, v4, and v5 projected at the upper left position, upper right position, and lower left position in the coded block including block A are derived. Then, the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0557] Note that Figure 49A and Figure 49B the MV derivation methods shown can be used for MV derivation at each control point of the current block in step Sk_1 shown later in Figure 50 or can be used for MV predictor derivation at each control point of the current block in step Sj_1 shown later in Figure 51 shown.
[0558] (MV Derivation > Affine Mode > Affine Merge Mode)
[0559] Figure 50 is a flowchart showing an example of the process in the affine merge mode.
[0560] In the affine merge mode as shown in the figure, first, the inter-frame predictor 126 derives the MVs at the corresponding control points of the current block (step Sk_1). As Figure 46A shown, the control points are the upper left corner point and the upper right corner point of the current block, or as Figure 46B shown, the upper left corner point, the upper right corner point, and the lower left corner point of the current block. The inter-frame predictor 126 can encode MV selection information for identifying two or three derived MVs in the stream.
[0561] For example, when using the Figures 47A to 47C MV derivation method shown, as Figure 47A shown, the inter-frame predictor 126 checks the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) in the listed order and identifies the first valid block encoded according to the affine mode.
[0562] The inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine mode to derive the MVs at the control points. For example, when block A is identified and block A has two control points, as Figure 47B shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left corner control point of the current block and the motion vector v1 at the upper right corner control point of the current block according to the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block including block A. For example, the inter-frame predictor 126 projects the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block onto the current block to calculate the motion vector v0 at the upper left corner control point of the current block and the motion vector v1 at the upper right corner control point of the current block.
[0563] Alternatively, when block A is identified and block A has three control points, as Figure 47C shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block according to the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the encoded block including block A. For example, the inter-frame predictor 126 projects the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block to calculate the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block.
[0564] Note that as described above in Figure 49A shown, when block A is identified and block A has two control points, the MVs at three control points can be calculated, and as described above inFigure 49B As shown, when block A is recognized and block A has three control points, the MVs at two control points can be calculated.
[0565] Next, the inter-frame predictor 126 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the MV of each of the plurality of sub-blocks as an affine MV (step Sk_2) using, for example, two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation of the sub-blocks (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all the sub-blocks included in the current block, the process of generating the predicted image using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate the predicted image of the current block.
[0566] Note that the above MV candidate list can be generated in step Sk_1. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example, Figures 47A to 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in and any combination of other MV derivation methods.
[0567] Note that, in addition to the affine mode, the MV candidate list can include MV candidates in a mode that performs prediction on a sub-block basis.
[0568] Note that, for example, an MV candidate list (which includes MV candidates in the affine merge mode using two control points and the affine merge mode using three control points) can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points can be generated. The MV candidate can be, for example, the MV for the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or the MV of the valid block in the block.
[0569] Note that the index indicating one of the MVs in the MV candidate list can be transmitted as the MV selection information.
[0570] (MV Derivation > Affine Mode > Affine Inter - Frame Mode)
[0571] Figure 51 It is a flowchart showing an example of the process in the affine inter - frame mode.
[0572] In the affine inter - frame mode, first, the inter - frame predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the upper - left corner point, the upper - right corner point, and the lower - right corner point of the current block, as Figure 46A or Figure 46B shown.
[0573] For example, when using the MV derivation method shown in Figure 48A and Figure 48B , the inter - frame predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block by selecting the MV of any block in the coded blocks near the corresponding control points of the current block shown in Figure 48A or Figure 48B . At this time, the inter - frame predictor 126 encodes the MV predictor selection information for identifying the selected two or three MV predictors in the bitstream.
[0574] For example, the inter - frame predictor 126 can determine, using cost evaluation etc., the block from which the MV is selected as the MV predictor at the control point from the coded blocks adjacent to the current block, and can write a flag indicating which MV predictor has been selected in the bitstream. In other words, the inter - frame predictor 126 outputs the MV predictor selection information such as a flag as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0575] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4), while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 calculates the MV of each sub-block corresponding to the updated MV predictor as an affine MV using the above expression (1A) or expression (1B) (step Sj_3). The inter-frame predictor 126 then performs motion compensation for the sub-blocks using these affine MVs and the encoded reference pictures (step Sj_4). When updating the MV predictor in step Sj_2, the processes in steps Sj_3 and Sj_4 are performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor as an MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0576] Finally, the inter-frame predictor 126 generates a predicted image of the current block by performing motion compensation for the current block using the determined MV and the encoded reference pictures (step Sj_6).
[0577] Note that the above MV candidate list can be generated in step Sj_1. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example, Figures 47A to 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in
[0578] Note that in addition to the affine mode, the MV candidate list can include MV candidates in a mode that performs prediction in units of sub-blocks.
[0579] Note that, for example, an MV candidate list including MV candidates in an affine inter-frame mode using two control points and an MV candidate list including MV candidates in an affine inter-frame mode using three control points can be generated as an MV candidate list. Alternatively, an MV candidate list including MV candidates in an affine inter-frame mode using two control points and an MV candidate list including MV candidates in an affine inter-frame mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of an affine inter-frame mode using two control points and an affine inter-frame mode using three control points can be generated. The MV candidate can be, for example, an MV for block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left) to be encoded, or an MV of a valid block in the block.
[0580] Note that an index indicating one of the MV candidates in the MV candidate list can be transmitted as MV predictor selection information.
[0581] (MV derivation > Triangle mode)
[0582] In the above example, the inter-frame predictor 126 generates a rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 can generate a plurality of prediction images, each having a shape different from the rectangle of the current rectangular block, and a plurality of prediction images can be combined to generate a final rectangular prediction image. The shape different from the rectangle can be, for example, a triangle.
[0583] Figure 52A is a conceptual diagram for showing the generation of two triangular prediction images.
[0584] The inter-frame predictor 126 generates a triangular prediction image by performing motion compensation on a first partition having a triangular shape in the current block using a first MV of the first partition to generate a triangular prediction image. Similarly, the inter-frame predictor 126 generates a triangular prediction image by performing motion compensation on a second partition having a triangular shape in the current block using a second MV of the second partition to generate a triangular prediction image. Then, the inter-frame predictor 126 generates a prediction image having a rectangular shape identical to the rectangular shape of the current block by combining these prediction images.
[0585] Note that a first prediction image having a rectangular shape corresponding to the current block can be generated using the first MV as a prediction image for the first partition. In addition, a second prediction image having a rectangular shape corresponding to the current block can be generated using the second MV as a prediction image for the second partition. A prediction image of the current block can be generated by performing weighted addition of the first prediction image and the second prediction image. Note that the part where the weighted addition is performed can be a partial region across the boundary between the first partition and the second partition.
[0586] Figure 52BA conceptual diagram for showing a first part of a first partition that overlaps a second partition and examples of first and second sample sets that can be weighted as part of a correction process. The first part can be, for example, a quarter of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition, where N is an integer greater than zero. For example, N can be the integer 2. As shown in the figure, Figure 52B The left example of Figure 52B shows a rectangular partition with a rectangular part, the width of which is a quarter of the width of the first partition, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The central example of Figure 52B shows a rectangular partition with a rectangular part, the height of which is a quarter of the height of the first partition, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Figure 52B The right example of Figure 52B shows a triangular partition with a polygonal part, the height of which corresponds to two samples, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.
[0587] The first part can be the part of the first partition that overlaps an adjacent partition. [[ID= A conceptual diagram for showing a first part of a first partition, where the first part is the part of the first partition that overlaps a part of an adjacent partition. For ease of illustration, a rectangular partition with an overlapping part with a spatially adjacent rectangular partition is shown. Partitions with other shapes, such as triangular partitions, can be used, and the overlapping part can overlap with a spatially or temporally adjacent partition.
[0588] In addition, although examples of generating a prediction image for each of two partitions using inter-frame prediction are given, an intra-frame prediction can be used to generate a prediction image for at least one partition.
[0589] A flowchart showing an example of a process in a triangular mode.
[0590] In the triangular mode, first, the inter-frame predictor 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode partition information (which is information related to the divided partitions) in the stream as prediction parameters. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0591] First, the inter-frame predictor 126 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sx_2). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0592] The inter-frame predictor 126 then respectively selects an MV candidate for the first partition and an MV candidate for the second partition as the first MV and the second MV from the plurality of MV candidates obtained in step Sx_1 (step Sx_3). At this time, the inter-frame predictor 126 encodes, in the stream, MV selection information for identifying the selected MV candidates as prediction parameters. In other words, the inter-frame predictor 126 outputs, as prediction parameters, the MV selection information to the entropy encoder 110 through the prediction parameter generator 130.
[0593] Next, the inter-frame predictor 126 generates a first prediction image by performing motion compensation using the selected first MV and the coded reference picture (step Sx_4). Similarly, the inter-frame predictor 126 generates a second prediction image by performing motion compensation using the selected second MV and the coded reference picture (step Sx_5).
[0594] Finally, the inter-frame predictor 126 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_6).
[0595] Note that although the first partition and the second partition are triangles in the illustrated example, the first partition and the second partition may be trapezoids, or other shapes different from each other. Further, although the current block includes two partitions in the and illustrated example, the current block may include three or more partitions.
[0596] In addition, the first partition and the second partition may overlap each other. In other words, the first partition and the second partition may include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition may be used to generate the prediction image for the current block.
[0597] In addition, although an example of generating a prediction image for each of two partitions using inter-frame prediction has been shown, an intra-frame prediction may be used to generate a prediction image for at least one partition.
[0598] Note that the list of MV candidates for selecting the first MV and the list of MV candidates for selecting the second MV may be different from each other, or the list of MV candidates for selecting the first MV may also be used as the list of MV candidates for selecting the second MV.
[0599] Note that the partition information may include an index indicating a splitting direction in which at least the current block is split into a plurality of partitions. The MV selection information may include an index indicating the selected first MV and an index indicating the selected second MV. One index may indicate multiple pieces of information. For example, one index that commonly indicates part or all of the partition information and part or all of the MV selection information may be encoded.
[0600] (MV Derivation>ATMVP Mode)
[0601] is a conceptual diagram showing an example of an advanced temporal motion vector prediction (ATMVP) mode for deriving an MV in units of sub-blocks.
[0602] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, MV candidates for each sub-block are registered in an MV candidate list for a regular merge mode.
[0603] More specifically, in the ATMVP mode, first, as shown, a temporal MV reference block associated with the current block is identified in an encoded reference picture specified by the MV (MV0) of an adjacent block located at the lower left position relative to the current block. Next, in each sub-block in the current block, an MV used to encode a region corresponding to the sub-block in the temporal MV reference block is identified. The MV identified in this way is included in the MV candidate list as an MV candidate for the sub-block in the current block. When an MV candidate for each sub-block is selected from the MV candidate list, the sub-block undergoes motion compensation, where the MV candidate is used as the MV of the sub-block. In this way, a predicted image for each sub-block is generated.
[0604] Although in the example shown, the block located at the lower left position relative to the current block is used as a surrounding MV reference block, it should be noted that another block may be used. Additionally, the size of the sub-block may be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub-block may be switched for units such as slices, tiles, pictures, etc.
[0605] (MV Derivation>DMVR)
[0606] is a flowchart showing the relationship between the merge mode and decoder motion vector refinement DMVR.
[0607] The inter - frame predictor 126 derives the motion vector of the current block according to the merge mode (step S1_1). Next, the inter - frame predictor 126 determines whether to perform the estimation of the motion vector, that is, motion estimation (step S1_2). Here, when it is determined not to perform motion estimation (No in step S1_2), the inter - frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merge mode.
[0608] When it is determined to perform motion estimation in step S1_1 (Yes in step S1_2), the inter - frame predictor 126 derives the final motion vector of the current block by estimating the surrounding area of the reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to DMVR.
[0609] is a conceptual diagram showing an example of the DMVR process for determining the MV.
[0610] First, for example, in the merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference picture (L0) which is an encoded picture in the L0 list according to the MV candidate (L0). Similarly, reference pixels are identified from the second reference picture (L1) which is an encoded picture in the L1 list according to the MV candidate (L1). A template is generated by calculating the average value of these reference pixels.
[0611] Next, each of the surrounding areas of the MV candidates of the first reference picture (L0) and the second reference picture (L1) is estimated using the template, and the MV that generates the minimum cost is determined as the final MV. Note that the cost can be calculated, for example, using the difference between each pixel value in the template and the corresponding pixel value in the estimated area, the value of the MV candidate, etc.
[0612] It is not always necessary to perform exactly the same process described here. Other processes for deriving the final MV through estimation in the surrounding area of the MV candidate can be used.
[0613] is a conceptual diagram showing another example of DMVR for determining the MV. Different from the example of DMVR shown in in the example shown in
[0614] First, the inter-frame predictor 126 estimates the surrounding area of the reference block in each reference picture included in the L0 list and the L1 list based on the initial MVs that are MV candidates obtained from each MV candidate list. For example, as shown, the initial MV corresponding to the reference block in the L0 table is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets the search position for the reference picture in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set (specifically, the initial MV (i.e., InitMV_L0)), the vector difference from the search position is MVd_L0. The inter-frame predictor 126 then determines the estimated position in the reference picture in the L1 list. This search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference as MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines the position symmetric to the position indicated by the initial MV as the search position in each reference picture in the L0 list and the L1 list. The inter-frame predictor 126 calculates the sum of the absolute differences (SAD) between the pixel values at the search positions in the block as the cost for each search position and finds the search position that produces the minimum cost.
[0615] is a conceptual diagram for illustrating an example of motion estimation in DMVR, and is a flowchart showing an example of the process of motion estimation.
[0616] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at a search position other than the starting point is the minimum, the inter-frame predictor 126 changes the target to the search position that obtains the minimum cost and executes the process in step 2. When the cost at the starting point is the minimum, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.
[0617] In step 2, the inter-frame predictor 126 performs a search similar to the process in step 1, and regards the search position after the target is changed as the new starting point according to the result of the process in step 1. Then the inter-frame predictor 126 determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at a search position other than the starting point is the minimum, the inter-frame predictor 126 executes the process in step 4. When the cost at the starting point is the minimum, the inter-frame predictor 126 executes the process in step 3.
[0618] In step 4, the inter-frame predictor 126 regards the search position at the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0619] In step 3, the inter-frame predictor 126 determines the pixel position with sub-pixel accuracy that obtains the minimum cost based on the costs at four points located above, below, left, and right with respect to the starting point in step 1 or step 2, and regards this pixel position as the final search position. The pixel position at sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) above, below, left, and right using the cost at the corresponding search position among the four search positions as a weight. The inter-frame predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0620] (Motion compensation > BIO / OBMC / LIC)
[0621] Motion compensation involves modes for generating a predicted image and correcting the predicted image. The modes are, for example, bidirectional optical flow (BIO), overlapping block motion compensation (OBMC), local illumination compensation (LIC), etc., which will be described later.
[0622] Figure 59 is a flowchart showing an example of the process of generating a predicted image.
[0623] The inter-frame predictor 126 generates a predicted image (step Sm_1), and corrects the predicted image according to any one of the above modes, for example (step Sm_2).
[0624] Figure 60 is a flowchart showing another example of the process of generating a predicted image.
[0625] The inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame predictor 126 generates a predicted image using the motion vector (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter-frame predictor 126 generates a final predicted image by correcting the predicted image (step Sn_4). Note that in LIC described later, the luminance and chrominance can be corrected in step Sn_4. When it is determined not to perform the correction process (No in step Sn_3), the inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).
[0626] (Motion compensation > OBMC)
[0627] Note that, in addition to the motion information of the current block obtained by motion estimation, the motion information of adjacent blocks can also be used to generate an inter-predicted image. More specifically, an inter-predicted image can be generated for each sub-block in the current block by performing a weighted addition of a predicted image (in a reference picture) based on the motion information obtained by motion estimation and a predicted image (in the current picture) based on the motion information of adjacent blocks. Such inter-prediction (motion compensation) is also referred to as Overlapped Block Motion Compensation (OBMC) or OBMC mode.
[0628] In the OBMC mode, information indicating the sub-block size of OBMC (e.g., referred to as the OBMC block size) can be signaled at the sequence level. In addition, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Note that the signaling of such information does not necessarily need to be performed at the sequence level and the CU level, and can be performed at another level (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0629] The OBMC mode will be described in more detail. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for showing an outline of a predicted image correction process performed by OBMC.
[0630] First, as Figure 62 shown, a predicted image (Pred) obtained by conventional motion compensation is obtained using the MV assigned to the current block. In Figure 62 , the arrow "MV" points to the reference picture and indicates what the current block of the current picture refers to in order to obtain the predicted image.
[0631] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) that has been derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which indicates the reference picture from the current block. A first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0632] Similarly, a predicted image (Pred_U) is obtained by applying an MV (MV_U) that has been derived for an encoded block adjacent to the current block above (reusing the MV of the current block) to the current block. The MV (MV_U) is indicated by an arrow "MV_U" that indicates the reference picture from the current block. A second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) for which a first correction has already been performed. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is thus the final predicted image of the current block.
[0633] Although the above example is a two-way correction method using left and upper adjacent blocks, note that the correction method can also be a three-way or more-way correction method that also uses right and / or lower adjacent blocks.
[0634] Note that the region where such overlapping is performed can be only a part of the region near the block boundary, rather than the pixel region of the entire block.
[0635] Note that the predicted image correction process for obtaining one predicted image Pred from one reference picture by overlapping additional predicted images Pred_L and Pred_U according to OBMC has been described above. However, when correcting a predicted image based on multiple reference images, a similar process can be applied to each of the multiple reference pictures. In this case, after obtaining corrected predicted images from the corresponding reference pictures by performing OBMC image correction based on multiple reference pictures, the obtained corrected predicted images are further overlapped to obtain a final predicted image.
[0636] Note that in OBMC, the current block unit can be a PU, or a sub-block unit obtained by further dividing a PU.
[0637] An example of a method for determining whether to apply OBMC is a method for using a signal obmc_flag that indicates whether to apply OBMC. As a specific example, the encoder 100 can determine whether the current block belongs to a region with complex motion. The encoder 100 sets the obmc_flag to the value "1" when the block belongs to a region with complex motion and applies OBMC during encoding, and sets the obmc_flag to the value "0" when the block does not belong to a region with complex motion and encodes the block without applying OBMC. The decoder 200 switches between the application and non-application of OBMC by decoding the obmc_flag written in the stream.
[0638] (Motion Compensation > BIO)
[0639] Next, the MV derivation method is described. First, a mode of deriving MV based on a model assuming uniform linear motion is described. This mode is also referred to as the bidirectional optical flow (BIO) mode. Additionally, this bidirectional optical flow can be written as BDOF instead of BIO.
[0640] Figure 63 is a conceptual diagram for showing a model assuming uniform linear motion. In Figure 63 , (v x , v y ) represents a velocity vector, and τ0 and τ1 represent the time distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) represents the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) represents the MV corresponding to the reference picture Ref1.
[0641] Here, under the assumption that uniform linear motion is exhibited by the velocity vector (v x , v y ), (MV x0 , MV y0 ) and (MV x1 , MV y1 ) are respectively expressed as (v xτ0 , v yτ0 ) and (-v xτ1 , -v yτ1 ), and the following optical flow equation (2) is given.
[0642] [Mathematical Expression 3]
[0643]
[0644] Here, I(k) represents the motion-compensated luminance value of the reference picture k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of the following is zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected in units of pixels.
[0645] Note that a method different from the model based on the assumption of uniform linear motion can be used on the decoder side 200 to derive the motion vector. For example, the motion vector can be derived in units of sub-blocks based on the motion vectors of multiple adjacent blocks.
[0646] Figure 64is a flowchart showing an example of an inter-frame prediction process according to BIO. Figure 65 is a functional block diagram showing an example of a functional configuration of an inter-frame predictor 126 that can perform inter-frame prediction according to BIO.
[0647] As Figure 65 shown, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolated image derivator 126b, a gradient image derivator 126c, an optical flow derivator 126d, a correction value derivator 126e, and a predicted image corrector 126f. Note that the memory 126a can be a frame memory 122.
[0648] The inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) including a current block. Then, the inter-frame predictor 126 derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.
[0649] Next, the interpolated image derivator 126b derives an interpolated image I 0 of the current block by referring to the memory 126a using the motion vector M0 and the reference picture L0. Next, the interpolated image derivator 126b derives an interpolated image I 1 of the current block by referring to the memory 126a using the motion vector M1 and the reference picture L1 (step Sy_2). Here, the interpolated image I 0 is an image included in the reference picture Ref0 and to be derived for the current block, and the interpolated image I 1 is an image included in the reference picture Ref1 and to be derived for the current block. Each of the interpolated image I 0 and the interpolated image I 1 can be the same size as the current block. Alternatively, each of the interpolated image I 0 and the interpolated image I 1 can be an image larger than the current block. In addition, the interpolated image I 0 and the interpolated image I 1 can include a predicted image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.
[0650] In addition, the gradient image derivator 126c derives from the interpolated image I 0and the interpolated image I 1 Derive the gradient image of the current block (Ix 0 ,Ix 1 ,Iy 0 ,Iy 1 )(step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 ,Ix 1 ), and the gradient image in the vertical direction is (Iy 0 ,Iy 1 ). The gradient image derivator 126c can derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image can indicate the amount of spatial variation of pixel values in the horizontal direction, in the vertical direction, or both.
[0651] Next, the optical flow derivator 126d uses the interpolated image (I 0 ,I 1 ) and the gradient image (Ix 0 ,Ix 1 ,Iy 0 ,Iy 1 ) to derive the optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_4). The optical flow indicates the coefficient for correcting the amount of spatial pixel movement and can be referred to as a local motion estimate value, a corrected motion vector, or a corrected weighted vector. As an example, the sub-block can be a 4×4 pixel sub-CU. Note that the optical flow derivation can be performed for each pixel unit or the like, rather than for each sub-block.
[0652] Next, the inter-frame predictor 126 uses the optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value derivator 126e uses the optical flow (vx, vy) to derive the correction value of the pixel values included in the current block (step Sy_5). The predicted image corrector 126f then uses the correction value to correct the predicted image of the current block (step Sy_6). Note that the correction value can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.
[0653] Note that the BIO process flow is not limited to Figure 64 the processes disclosed in Figure 64 . For example, only a part of the processes disclosed in
[0654] (Motion Compensation > LIC)
[0655] Next, an example of a mode for generating a predicted image (prediction) using the local illumination compensation (LIC) process is described.
[0656] Figure 66A It is a conceptual diagram showing an example of a process of a prediction image generation method that uses a luminance correction process performed by LIC. Figure 66B It is a flowchart showing an example of a process of a prediction image generation method using LIC.
[0657] First, the inter-frame predictor 126 derives an MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).
[0658] Next, the inter-frame predictor 126 extracts information indicating how the luminance value changes between the current block and the reference picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region) in the current picture, and the luminance pixel values at the corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 126 uses the information indicating how the luminance value changes to calculate a luminance correction parameter (step Sz_3).
[0659] The inter-frame predictor 126 generates a prediction image of the current block by performing a luminance correction process in which the luminance correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_4). In other words, the prediction image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the chrominance can be corrected, or both. In other words, information indicating how the chrominance changes can be used to calculate a chrominance correction parameter, and a chrominance correction process can be performed.
[0660] Note that Figure 66A the shape of the surrounding reference region shown in
[0661] is an example; another shape can be used.
[0662] An example of a method for determining whether to apply the LIC is a method using a lic_flag as a signal indicating whether to apply the LIC. As a specific example, the encoder 100 determines whether the current block belongs to a region with a luminance change. When the block belongs to a region with a luminance change, the encoder 100 sets the lic_flag to the value "1" and applies the LIC during encoding, and when the block does not belong to a region with a luminance change, sets the lic_flag to the value "0" and performs encoding without applying the LIC. The decoder 200 can decode the lic_flag in the written stream and decode the current block by switching between application and non-application of the LIC according to the flag value.
[0663] An example of a different method for determining whether to apply the LIC process is a determination method based on whether the LIC process has been applied to surrounding blocks. As a specific example, when the current block has been processed in the merge mode, the inter-frame predictor 126 determines whether the encoded surrounding blocks selected in the MV derivation in the merge mode have been encoded using the LIC. The inter-frame predictor 126 performs encoding by switching between application and non-application of the LIC according to the result. Note that also in this example, the same process is applied in the process on the decoder 200 side.
[0664] The luminance correction (LIC) process has been described with reference to Figure 66A and Figure 66B and is further described below.
[0665] First, the inter-frame predictor 126 derives an MV for obtaining a reference image corresponding to the current block to be encoded from a reference picture that is an encoded picture.
[0666] Next, the inter-frame predictor 126 uses the luminance pixel values of the encoded surrounding reference regions adjacent to the left and above of the current block and the luminance values at the corresponding positions in the reference picture specified by the MV to extract information indicating how the luminance value of the reference picture changes to the luminance value of the current picture, and calculates a luminance correction parameter. For example, assume that the luminance pixel value of a given pixel in the surrounding reference region in the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region in the reference picture is p1. The inter-frame predictor 126 calculates the coefficients A and B for optimizing A×p1 + B = p0 as the luminance correction parameters for multiple pixels in the surrounding reference region.
[0667] Next, the inter-frame predictor 126 performs a luminance correction process using the luminance correction parameter of the reference image in the reference picture specified by the MV to generate a predicted image of the current block. For example, assume that the luminance pixel value in the reference image is p2, and the luminance pixel value after luminance correction of the predicted image is p3. The inter-frame predictor 126 generates a predicted image after undergoing the luminance correction process by calculating A×p2 + B = p3 for each pixel in the reference image.
[0668] For example, a region having a determined number of pixels extracted from each of the upper adjacent pixel and the left adjacent pixel can be used as a surrounding reference region. Additionally, the surrounding reference region is not limited to a region adjacent to the current block, and can be a region not adjacent to the current block. In Figure 66A the example shown, the surrounding reference region in the reference picture can be a region specified by another MV in the surrounding reference region in the current picture. For example, the another MV can be an MV in the surrounding reference region in the current picture.
[0669] Although the operations performed by the encoder 100 have been described here, it should be noted that the decoder 200 performs similar operations.
[0670] Note that the LIC can be applied not only to luminance but also to chrominance. At this time, the correction parameters can be separately derived for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of Y, Cb, and Cr.
[0671] Furthermore, the LIC process can be applied in units of sub-blocks. For example, the correction parameters can be derived using the surrounding reference region in the current sub-block and the surrounding reference region in the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0672] (Prediction Controller)
[0673] The prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from the intra-frame predictor 124) and the inter-frame prediction signal (the image or signal output from the inter-frame predictor 126), and outputs the selected predicted image to the subtractor 104 and the adder 116 as a prediction signal.
[0674] (Prediction Parameter Generator)
[0675] The prediction parameter generator 130 may output information related to intra prediction, inter prediction, selection of a predicted image in the prediction controller 128, etc. as prediction parameters to the entropy encoder 110. The entropy encoder 110 may generate a stream based on the prediction parameters input from the prediction parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameters may be used in the decoder 200. The decoder 200 may receive and decode the stream, and perform the same processes as the prediction processes performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameters may include, for example, (i) selection of a prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) an optional index, flag, or value based on the prediction processes performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128 or indicating the prediction processes.
[0676] (Decoder)
[0677] Next, the decoder 200 capable of decoding the stream output from the above-described encoder 100 will be described. Figure 67 FIG. is a block diagram showing a functional configuration of the decoder 200 according to the present embodiment. The decoder 200 is a device that decodes a stream that is an encoded image in units of blocks.
[0678] As Figure 67 shown, the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformer 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a prediction parameter generator 222, and a segmentation determiner 224. Note that the intra predictor 216 and the inter predictor 218 are configured as part of prediction executors.
[0679] (Installation example of decoder)
[0680] Figure 68 FIG. is a functional block diagram showing an installation example of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, Figure 67 as shown, a plurality of components of the decoder 200 are installed on Figure 68 the processor b1 and the memory b2 shown.
[0681] The processor b1 is a circuit that performs information processing and is coupled to the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 may be a processor such as a CPU. In addition, the processor b1 may be an aggregate of a plurality of electronic circuits. Further, for example, the processor b1 may undertake Figure 67The roles of two or more components among the components of the decoder 200 shown, excluding the components for storing information.
[0682] The memory b2 is a dedicated or general-purpose memory for storing information used by the processor b1 to decode the stream. The memory b2 can be an electronic circuit and can be connected to the processor b1. In addition, the memory b2 can be included in the processor b1. In addition, the memory b2 can be an aggregate of multiple electronic circuits. Additionally, the memory b2 can be a magnetic disk, an optical disk, etc., or can be represented as a storage device, a recording medium, etc. In addition, the memory b2 can be a non-volatile memory or a volatile memory.
[0683] For example, the memory b2 can store an image or a stream. Additionally, the memory b2 can store a program for causing the processor b1 to decode the stream.
[0684] In addition, for example, the memory b2 can act as Figure 67 the roles of two or more components for storing information among the components of the decoder 200 shown in etc. More specifically, the memory b2 can act as Figure 67 the roles of the block memory 210 and the frame memory 214 shown. More specifically, the memory b2 can store the reconstructed image (specifically, the reconstructed block, the reconstructed picture, etc.).
[0685] Note that in the decoder 200, not all elements of the multiple components shown in Figure 67 can be implemented, and not all processes described herein can be executed. Figure 67 A part of the components shown, etc., can be included in another device, or a part of the processes described herein can be executed by another device.
[0686] Hereinafter, the overall flow of the process executed by the decoder 200 will be described, and then each component included in the decoder 200 will be described. Note that some components included in the decoder 200 execute the same processes as some of those in the encoder 100, so the same processes will not be described in detail again. For example, the inverse quantizer 204, the inverse transformer 206, the adder 208, the block memory 210, the frame memory 214, the intra predictor 216, the inter predictor 218, the prediction controller 220, and the loop filter 212 included in the decoder 200 respectively execute processes similar to those executed by the inverse quantizer 112, the inverse transformer 114, the adder 116, the block memory 118, the frame memory 122, the intra predictor 124, the inter predictor 126, the prediction controller 128, and the loop filter 120 included in the decoder 200.
[0687] (Overall Flow of the Decoding Process)
[0688] Figure 69 It is a flowchart showing an example of the overall decoding process performed by the decoder 200.
[0689] First, the segmentation determiner 224 in the decoder 200 determines the segmentation pattern of each of the plurality of fixed-size blocks (e.g., 128×128 pixels) included in the picture based on the parameters input from the entropy decoder 202 (step Sp_1). This segmentation pattern is the segmentation pattern selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks of the segmentation pattern.
[0690] The entropy decoder 202 decodes (specifically, entropy decodes) the encoded and quantized coefficients and the prediction parameters of the current block (step Sp_2).
[0691] Next, the inverse quantizer 204 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 206 performs an inverse transform on the result to recover the prediction residual (i.e., the difference block) (step Sp_3).
[0692] Next, the prediction executor including all or part of the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220 generates a prediction signal for the current block (step Sp_4).
[0693] Next, the adder 208 adds the predicted image and the prediction residual to generate a reconstructed image of the current block (also referred to as a decoded image block) (step Sp_5).
[0694] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).
[0695] The decoder 200 then determines whether the decoding of the entire picture has ended (step Sp_7). When it is determined that the decoding has not ended (No in step Sp_7), the decoder 200 repeats the process starting from step Sp_1.
[0696] Note that the processes of these steps Sp_1 to Sp_7 can be sequentially executed by the decoder 200, or two or more processes can be executed in parallel. The processing order of two or more processes can be modified.
[0697] (Segmentation Determiner)
[0698] Figure 70 It is a conceptual diagram for showing the relationship between the segmentation determiner 224 and other components in the embodiment. As an example, the segmentation determiner 224 can perform the following processes.
[0699] For example, the segmentation determiner 224 collects block information from the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoder 202. Then, the segmentation determiner 224 can determine the segmentation pattern of the fixed-size block based on the block information and the parameters. The segmentation determiner 224 can then output information indicating the determined segmentation pattern to the inverse transformer 206, the intra predictor 216, and the inter predictor 218. The inverse transformer 206 can perform inverse transformation of the transform coefficients based on the segmentation pattern indicated by the information from the segmentation determiner 224. The intra predictor 216 and the inter predictor 218 can generate a prediction image based on the segmentation pattern indicated by the information from the segmentation determiner 224.
[0700] (Entropy decoder)
[0701] Figure 71 is a block diagram showing an example of the functional configuration of the entropy decoder 202.
[0702] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the segmentation pattern by performing entropy decoding on the stream. For example, CABAC is used for entropy decoding. More specifically, the entropy decoding 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a de-binarizer 202c. The binary arithmetic decoder 202a arithmetically decodes the stream into a binary signal using the context value derived by the context controller 202b. The context controller 202b derives the context value in the same manner as the context controller 110b of the encoder 100, based on the characteristics of the syntax element or the surrounding state (i.e., the occurrence probability of the binary signal). The de-binarizer 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. This binarization can be performed according to the binarization method described above.
[0703] In this way, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can output the prediction parameters included in the stream (see Figure 1 ) to the intra predictor 216, the inter predictor 218, and the prediction controller 220. The intra predictor 216, the inter predictor 218, and the prediction controller 220 are capable of performing the same prediction process as the prediction process performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.
[0704] Figure 72 is a conceptual diagram of the process flow for showing an exemplary CABAC process in the entropy decoder 202.
[0705] First, initialization is performed in the CABAC in the entropy decoder 202. In the initialization, initialization in the binary arithmetic decoder 202a and setting of initial context values are performed. The binary arithmetic decoder 202a and the de-binarizer 202c then perform arithmetic decoding and de-binarization of the encoded data of, for example, a CTU. At this time, the context controller 202b updates the context value each time arithmetic decoding is performed. The context controller 202b then saves the context value as post-processing. For example, the saved context value is used to initialize the context value of the next CTU.
[0706] (Inverse Quantizer)
[0707] The inverse quantizer 204 inverse quantizes the quantized coefficients of the current block input from the entropy decoder 202. More specifically, the inverse quantizer 204 inverse quantizes the quantized coefficients of the current block based on the quantization parameter corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inverse quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.
[0708] Figure 73 is a block diagram showing an example of the functional configuration of the inverse quantizer 204.
[0709] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage device 204d, and an inverse quantization executor 204e.
[0710] Figure 74 is a flowchart showing an example of the inverse quantization process performed by the inverse quantizer 204.
[0711] The inverse quantizer 204 can perform an inverse quantization process for each CU based on Figure 74 the process shown as an example. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameter of the current block from the entropy decoder 202 (step Sv_12).
[0712] Next, the predicted quantization parameter generator 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage device 204d (step Sv_13). The predicted quantization parameter generator 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).
[0713] The quantization parameter generator 204a then generates the quantization parameter for the current block based on the differential quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the differential quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b can be added together to generate the quantization parameter for the current block. In addition, the quantization parameter generator 204a stores the quantization parameter of the current block in the quantization parameter storage device 204d (step Sv_16).
[0714] Next, the inverse quantization executor 204e inverse quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).
[0715] Note that the differential quantization parameter can be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the initial value of the quantization parameter and the differential quantization parameter can be used to generate the quantization parameter.
[0716] Note that the inverse quantizer 204 may include multiple inverse quantizers, and the inverse quantization method selected from multiple inverse quantization methods can be used to inverse quantize the quantized coefficients.
[0717] (Inverse transformer)
[0718] The inverse transformer 206 restores the prediction residual by inverse-transforming the transform coefficients that are the input to the inverse quantizer 204.
[0719] For example, when the information parsed from the stream indicates that EMT or AMT is to be applied (e.g., when the AMT flag is true), the inverse transformer 206 inverse-transforms the transform coefficients of the current block based on the information indicating the parsed transform type.
[0720] In addition, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.
[0721] Figure 75 is a flowchart showing an example of the process performed by the inverse transformer 206.
[0722] For example, the inverse transformer 206 determines whether there is information in the stream indicating that the orthogonal transformation has not been performed (step St_11). Here, when it is determined that such information does not exist (No in step St_11) (for example: there is no indication as to whether the orthogonal transformation is to be performed; there is an indication that the orthogonal transformation is to be performed), the inverse transformer 206 obtains information indicating the type of transformation decoded by the entropy decoder 202 (step St_12). Next, based on this information, the inverse transformer 206 determines the type of transformation used for the orthogonal transformation in the encoder 100 (step St_13). The inverse transformer 206 then performs an inverse orthogonal transformation using the determined type of transformation (step St_14). As Figure 75 shown, when it is determined that there is information indicating that the orthogonal transformation has not been performed (Yes in step St_11) (for example, an explicit indication that the orthogonal transformation has not been performed; there is no indication to perform the orthogonal transformation), the orthogonal transformation is not performed.
[0723] Figure 76 is a flowchart showing an example of the process performed by the inverse transformer 206.
[0724] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a determined value (step Su_11). The determined value can be pre-determined. Here, when it is determined that the transform size is less than or equal to the determined value (Yes in step Su_11), the inverse transformer 206 obtains from the entropy decoder 202 information indicating which type of transformation among at least one type of transformation included in the first transformation type group was used by the encoder 100 (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transformer 206.
[0725] Based on this information, the inverse transformer 206 determines the type of transformation used for the orthogonal transformation in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transformation on the transform coefficients of the current block using the determined type of transformation (step Su_14). When it is determined that the transform size is not less than or equal to the determined value (No in step Su_11), the inverse transformer 206 performs an inverse transformation on the transform coefficients of the current block using the second transformation type group (step Su_15).
[0726] Note that, as an example, the inverse orthogonal transformation of the inverse transformer 206 can be performed for each TU according to Figure 75 or Figure 76 shown. In addition, the inverse orthogonal transformation can be performed by using a defined type of transformation without decoding the information indicating the type of transformation used for the orthogonal transformation. The defined type of transformation can be a pre-defined type of transformation or a default type of transformation. Additionally, the type of transformation can specifically be DST7, DCT8, etc. In the inverse orthogonal transformation, the inverse transformation basis function corresponding to the type of transformation is used.
[0727] (Adder)
[0728] The adder 208 reconstructs the current block by adding the prediction residual which is the input from the inverse transformer 206 and the prediction residual which is the input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.
[0729] (Block Memory)
[0730] The block memory 210 is a storage device for storing blocks included in the current picture and can be referenced in intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.
[0731] (Loop Filter)
[0732] The loop filter 212 applies the loop filter to the reconstructed image generated by the adder 208, outputs the filtered reconstructed image to the frame memory 214, and provides the output of the decoder 200, for example, and outputs it to a display device, etc.
[0733] When the information indicating the turn-on or turn-off of the ALF parsed from the stream indicates that the ALF is turned on, one filter can be selected from multiple filters based on, for example, the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0734] Figure 77 is a block diagram showing an example of the functional configuration of the loop filter 212. Note that the configuration of the loop filter 212 is similar to the configuration of the loop filter 120 of the encoder 100.
[0735] For example, as Figure 77 shown, the loop filter 212 includes a deblocking filter executor 212a, a SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filter process on the reconstructed image. The SAO executor 212b performs a SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include Figure 77 all the constituent elements disclosed in Figure 77 and can include only a part of the constituent elements. In addition, the loop filter 212 can be configured to execute the above processes in a processing order different from the processing order disclosed in Figure 77 and can not execute all the processes shown in
[0736] (Frame Memory)
[0737] The frame memory 214 is, for example, a storage device for storing reference pictures used in inter-frame prediction, and may also be referred to as a frame buffer. More specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter 212.
[0738] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))
[0739] Figure 78 is a flowchart showing an example of the process performed by the predictor of the decoder 200. Note that the prediction executor may include all or part of the following constituent elements: the intra-frame predictor 216; the inter-frame predictor 218; and the prediction controller 220. The prediction executor includes, for example, the intra-frame predictor 216 and the inter-frame predictor 218.
[0740] The predictor generates a predicted image of the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra-frame prediction signal or an inter-frame prediction signal. More specifically, the predictor uses the reconstructed image that has been obtained for another block through the generation of the prediction image, the recovery of the prediction residue, and the addition of the prediction images to generate the predicted image of the current block. The predictor of the decoder 200 generates the same predicted image as the predicted image generated by the predictor of the encoder 100. In other words, the predicted image is generated according to a method common to or corresponding to each other between the predictors.
[0741] The reconstructed image may be, for example, an image in a reference picture, or an image of a decoded block (i.e., the above-mentioned other block) in the current picture including the current block. The decoded block in the current picture is, for example, an adjacent block of the current block.
[0742] Figure 79 is a flowchart showing another example of the process performed by the predictor of the decoder 200.
[0743] The predictor determines a method or mode for generating the predicted image (step Sr_1). For example, the method or mode may be determined based on, for example, prediction parameters, etc.
[0744] When the first method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the third method (step Sr_2c).
[0745] The first method, the second method, and the third method may be mutually different methods for generating a predicted image. Each of the first to third methods may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.
[0746] Figures 80A to 80C (Collectively referred to as FIG. 80) is a flowchart showing another example of the process executed by the predictor of the decoder 200.
[0747] As an example, the predictor may execute a prediction process according to the flow shown in FIG. 80. Note that the intra-block copy shown in FIG. 80 is a mode belonging to inter-frame prediction, and the blocks included in the current picture are referred to as reference images or reference blocks. In other words, in intra-block copy, a picture different from the current picture is not referred to. Additionally, the PCM mode shown in FIG. 80 is a mode belonging to intra-frame prediction, and transformation and quantization are not performed therein.
[0748] (Intra-frame predictor)
[0749] The intra-frame predictor 216 performs intra-frame prediction by referring to the blocks in the current picture stored in the block memory 210, based on the intra-frame prediction mode parsed from the stream, to generate a predicted image of the current block (i.e., the intra-frame prediction block). More specifically, the intra-frame predictor 216 performs intra-frame prediction by referring to the pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-frame prediction image, and then outputs the intra-frame prediction image to the prediction controller 220.
[0750] Note that when an intra-frame prediction mode in which the luminance block is referred to in the intra-frame prediction of the chrominance block is selected, the intra-frame predictor 216 may predict the chrominance component of the current block based on the luminance component of the current block.
[0751] Furthermore, when the information parsed from the stream indicates that PDPC is to be applied, the intra-frame predictor 216 corrects the pixel values of the intra-frame prediction based on the horizontal / vertical reference pixel gradients.
[0752] Figure 81 is a diagram showing an example of the process executed by the intra-frame predictor 216 of the decoder 200.
[0753] The intra-frame predictor 216 first determines whether to adopt MPM. As Figure 81As shown, the intra predictor 216 determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 exists (Yes in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 among the MPMs from the entropy decoder 202. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode (step Sw_14), which is included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_12.
[0754] When it is determined that the MPM flag indicating 1 does not exist (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains from the entropy decoder 202 information indicating the intra prediction mode selected from at least one intra prediction mode that has never been included in the MPM in the encoder 100. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Then, the intra predictor 216 determines the intra prediction mode (step Sw_17), which is not included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15.
[0755] The intra predictor 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0756] (Inter - frame predictor)
[0757] The inter - frame predictor 218 predicts the current block by referring to the reference pictures stored in the frame memory 214. Prediction is performed in units of the current block or the current sub - block in the current block. Note that the sub - block is included in the block and is a unit smaller than the block. The size of the sub - block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub - block can be switched in units such as slices, bricks, pictures, etc.
[0758] For example, the inter - frame predictor 218 generates an inter - frame prediction image of the current block or the current sub - block by performing motion compensation using the motion information (e.g., MV) parsed from the stream (e.g., the prediction parameters output from the entropy decoder 202), and outputs the inter - frame prediction image to the prediction controller 220.
[0759] When the information parsed from the stream indicates that the OBMC mode is to be applied, in addition to the motion information of the current block obtained through motion estimation, the inter-frame predictor 218 also uses the motion information of adjacent blocks to generate an inter-frame predicted image.
[0760] In addition, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives motion information by performing motion estimation according to the pattern matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter-frame predictor 218 then performs motion compensation (prediction) using the derived motion information.
[0761] In addition, when the BIO mode is to be applied, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. In addition, when the information parsed from the stream indicates that the affine mode is to be applied, the inter-frame predictor 218 derives the MV of each sub-block based on the MVs of multiple adjacent blocks.
[0762] (MV Derivation Process)
[0763] Figure 82 is a flowchart showing an example of the MV derivation process in the decoder 200.
[0764] For example, the inter-frame predictor 218 determines whether to decode motion information (e.g., MV). For example, the inter-frame predictor 218 can make the determination according to the prediction mode included in the stream, or can make the determination based on other information included in the stream. Here, when it is determined to decode motion information, the inter-frame predictor 218 derives the MV of the current block in the mode of decoding the motion information. When it is determined not to decode motion information, the inter-frame predictor 218 derives the MV in the mode of not decoding the motion information.
[0765] Here, the MV derivation modes include the conventional inter-frame mode, the conventional merge mode, the FRUC mode, the affine mode, etc., which will be described later. The modes of decoding motion information in the modes include the conventional inter-frame mode, the conventional merge mode, the affine mode (specifically, the affine inter-frame mode and the affine merge mode), etc. Note that the motion information can include not only the MV but also the MV predictor selection information described later. The modes of not decoding motion information include the FRUC mode, etc. The inter-frame predictor 218 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.
[0766] Figure 83 is a flowchart showing an example of the process of MV derivation in the decoder 200.
[0767] For example, the inter-frame predictor 218 may determine whether to decode the MV difference, i.e., for example, the determination may be made according to the prediction mode included in the stream, or the determination may be made based on other information included in the stream. Here, when it is determined to decode the MV difference, the inter-frame predictor 218 may derive the MV of the current block in the mode of decoding the MV difference. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.
[0768] When it is determined not to decode any MV difference, the inter-frame predictor 218 derives the MV in the mode of not decoding the MV difference. In this case, the encoded MV difference is not included in the stream.
[0769] Here, as described above, the MV derivation modes include the conventional inter-frame mode, the conventional merge mode, the FRUC mode, the affine mode, etc., which will be described later. The modes in which the MV difference is encoded include the conventional inter-frame mode and the affine mode (specifically, the affine inter-frame mode), etc. The modes in which the MV difference is not encoded include the FRUC mode, the conventional merge mode, the affine mode (specifically, the affine merge mode), etc. The inter-frame predictor 218 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.
[0770] (MV Derivation > Conventional Inter-Frame Mode)
[0771] For example, when the information parsed from the stream indicates that the conventional inter-frame mode is to be applied, the inter-frame predictor 218 derives the MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.
[0772] Figure 84 is a flowchart showing an example of the process of inter-frame prediction by the conventional inter-frame mode in the decoder 200.
[0773] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sg_11). In other words, the inter-frame predictor 218 generates an MV candidate list.
[0774] Next, the inter-frame predictor 218 extracts N (an integer of 2 or greater) MV candidates as motion vector predictor candidates (also referred to as MV predictor candidates) from the multiple MV candidates obtained in step Sg_11 according to the determined ranking in the priority order (step Sg_12). Note that the ranking in the priority order can be determined in advance for the corresponding N MV predictor candidates, and the ranking can be pre-determined.
[0775] Next, the inter-frame predictor 218 decodes the MV predictor selection information from the input stream, and uses the decoded MV predictor selection information to select one MV predictor candidate from among N MV predictor candidates as the MV predictor for the current block (step Sg_13).
[0776] Next, the inter-frame predictor 218 decodes the MV difference from the input stream, and derives the MV of the current block by adding the difference, which is the decoded MV difference, to the selected MV predictor (step Sg_14).
[0777] Finally, the inter-frame predictor 218 generates a predicted image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Sg_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed for each block among all the blocks in a slice, the inter-frame prediction of the slice using the normal inter-frame mode ends. For example, when the processes in steps Sg_11 to Sg_15 are performed for each block among all the blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode ends. Note that not all the blocks included in a slice undergo the processes in steps Sg_11 to Sg_15, and when some blocks undergo the processes, the inter-frame prediction of the slice using the normal inter-frame mode can end. This also applies to the picture in steps Sg_11 to Sg_15. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode can end.
[0778] (MV Derivation > Normal Merge Mode)
[0779] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter-frame predictor 218 derives the MV and performs motion compensation (prediction) using the MV.
[0780] Figure 85 is a flowchart showing an example of the process of inter-frame prediction by the normal merge mode in the decoder 200.
[0781] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sh_11). In other words, the inter-frame predictor 218 generates an MV candidate list.
[0782] Next, the inter-frame predictor 218 selects one MV candidate from among the multiple MV candidates obtained in step Sh_11 and derives the MV of the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains the MV selection information included in the stream as a prediction parameter, and selects the MV candidate identified by the MV selection information as the MV of the current block.
[0783] Finally, the inter-frame predictor 218 generates a predicted image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed for each block among all the blocks in a slice, the inter-frame prediction of the slice using the regular merge mode ends. Additionally, when the processes in steps Sh_11 to Sh_13 are performed for each block among all the blocks in a picture, the inter-frame prediction of the picture using the regular merge mode ends. Note that not all the blocks included in a slice undergo the processes in steps Sh_11 to Sh_13, and when some blocks undergo the processes, the inter-frame prediction of the slice using the regular merge mode can end. This also applies to the picture in steps Sh_11 to Sh_13. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the regular merge mode can end.
[0784] (MV derivation > FRUC mode)
[0785] For example, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives the MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is derived on the decoder 200 side without being signaled from the encoder 100 side. For example, the decoder 200 can derive the motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values in the current block.
[0786] Figure 86 is a flowchart showing an example of the process of inter-frame prediction by the FRUC mode in the decoder 200.
[0787] First, the inter - frame predictor 218 generates a list of MVs indicating decoded blocks adjacent to the current block spatially or temporally as MV candidates by referring to the MVs (this list is an MV candidate list and can also be used as an MV candidate list for a normal merge mode, for example (step Si_11)). Next, the best MV candidate is selected from among the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter - frame predictor 218 calculates the evaluation value of each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter - frame predictor 218 then derives the MV of the current block (step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV of the current block. Additionally, for example, the MV of the current block is derived using pattern matching in the surrounding area of the position corresponding to the selected best MV candidate included in the reference picture. In other words, estimation using pattern matching and evaluation values in the reference picture can be performed in the surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces a better evaluation value, and the updated MV can be determined as the final MV of the current block. In an embodiment, the update to the MV that produces a better evaluation value may not be performed.
[0788] Finally, the inter - frame predictor 218 generates a predicted image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all the blocks of a slice, the inter - frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all the blocks of a picture, the inter - frame prediction of the picture using the FRUC mode ends. Each sub - block can be processed similarly to the case of each block.
[0789] (MV Derivation > FRUC Mode)
[0790] For example, when the information parsed from the stream indicates that the affine merge mode is to be applied, the inter - frame predictor 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0791] Figure 87 is a flowchart showing an example of the process of inter - frame prediction in the decoder 200 through the affine merge mode.
[0792] In the affine merge mode, first, the inter - frame predictor 218 derives the MV at the corresponding control points of the current block (step Sk_11). As Figure 46AAs shown, the control points are the upper left point and the upper right point of the current block, or as Figure 46B shown, are the upper left point, the upper right point, and the lower left point of the current block.
[0793] For example, when using the Figures 47A to 47C shown MV derivation method, as Figure 47A shown, the inter - frame predictor 218 checks the decoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) in this order and identifies the first valid block decoded according to the affine mode. The inter - frame predictor 218 uses the identified first valid block decoded according to the affine mode to derive the MV at the control points. For example, when block A is identified and block A has two control points, as Figure 47B shown, the inter - frame predictor 218 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left and upper right corners of the decoded block including block A. In this way, the MV at each control point is derived.
[0794] Note that, as Figure 49A shown, when block A is identified and block A has two control points, the MV at three control points can be calculated, and as Figure 49B shown, when block A is identified and block A has three control points, the MV at two control points can be calculated.
[0795] In addition, when MV selection information is included in the stream as a prediction parameter, the inter - frame predictor 218 can use the MV selection information to derive the MV at each control point of the current block.
[0796] Next, the inter - frame predictor 218 performs motion compensation on each block included in the current block. In other words, the inter - frame predictor 218 uses two motion vectors v0 and v1 and the above - mentioned expression (1A) or three motion vectors v0, v1, and v2 and the above - mentioned expression (1B) to calculate the MV for each of the multiple sub - blocks as the affine MV (step Sk_12). The inter - frame predictor 218 then uses these affine MVs and the decoded reference picture to perform motion compensation on the sub - blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all the sub - blocks included in the current block, the inter - frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate the predicted image of the current block.
[0797] Note that the above - mentioned MV candidate list can be generated in step Sk_11. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for exampleFigures 47A to 47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49B the MV derivation method shown in
[0798] Note that, except for the affine mode, the MV candidate list may include MV candidates in the mode that performs prediction in units of sub-blocks.
[0799] Note that, for example, an MV candidate list including MV candidates in the affine merge mode using two control points and MV candidates in the affine merge mode using three control points can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points can be generated.
[0800] (MV Derivation > Affine Inter-Frame Mode)
[0801] For example, when the information parsed from the stream indicates that the affine inter-frame mode will be applied, the inter-frame predictor 218 derives an MV in the affine inter-frame mode and performs motion compensation (prediction) using the MV.
[0802] Figure 88 is a flowchart showing an example of the process of inter-frame prediction by the affine inter-frame mode in the decoder 200.
[0803] In the affine inter-frame mode, first, the inter-frame predictor 218 derives an MV predictor (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_11). The control points are the upper left corner point, the upper right corner point, and the lower left corner point of the current block, as Figure 46A or Figure 46B shown.
[0804] The inter-frame predictor 218 obtains MV predictor selection information included in the stream as a prediction parameter, and derives an MV predictor at each control point of the current block using the MV identified by the MV predictor selection information. For example, when using the Figure 48A and Figure 48B shown MV derivation methods, the inter-frame predictor 218 selects by Figure 48A or Figure 48BThe motion vectors of the blocks identified by the motion vector predictor selection information in the coded blocks near the corresponding control points of the current block shown in the figure are used to derive the motion vector predictors (v0, v1) or (v0, v1, v2) at the control points of the current block.
[0805] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block and the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point of the current block is derived.
[0806] Next, the inter-frame predictor 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter-frame predictor 218 calculates the MV of each of the plurality of sub-blocks as an affine MV using two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) (step Sj_13). The inter-frame predictor 218 then uses these affine MVs and the decoded reference pictures to perform motion compensation on the sub-blocks (step Sj_14). When the processes in steps Sj_13 and Sj_14 are performed for each sub-block included in the current block, the inter-frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate a predicted image of the current block.
[0807] Note that the above MV candidate list can be generated in step Sj_11 as in step Sk_11.
[0808] (MV Derivation > Triangle Mode)
[0809] For example, when the information parsed from the stream indicates that the triangle mode is to be applied, the inter-frame predictor 218 derives the MV in the triangle mode and performs motion compensation (prediction) using the MV.
[0810] Figure 89 is a flowchart showing an example of the process of inter-frame prediction in the decoder 200 by the triangle mode.
[0811] In the triangle mode, first, the inter-frame predictor 218 divides the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as a prediction parameter, which is information related to the division. The inter-frame predictor 218 can then divide the current block into a first partition and a second partition according to the partition information.
[0812] Next, the inter-frame predictor 218 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0813] The inter-frame predictor 218 then respectively selects an MV candidate for the first partition and an MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV (step Sx_13). At this time, the inter-frame predictor 218 may obtain MV selection information from the stream for identifying each selected MV candidate as a prediction parameter. The inter-frame predictor 218 can then select the first MV and the second MV according to the MV selection information.
[0814] Next, the inter-frame predictor 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame predictor 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).
[0815] Finally, the inter-frame predictor 218 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_16).
[0816] (MV estimation > DMVR)
[0817] For example, if the information parsed from the stream indicates that DMVR is to be applied, the inter-frame predictor 218 performs motion estimation using DMVR.
[0818] Figure 90 is a flowchart showing an example of the process of motion estimation performed by DMVR in the decoder 200.
[0819] The inter-frame predictor 218 derives the MV for the current block according to the merge mode (step S1_11). Next, the inter-frame predictor 218 derives the final MV for the current block by searching the area around the reference picture indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined according to DMVR.
[0820] Figure 91 is a flowchart showing an example of the motion estimation process performed by DMVR in the decoder 200 and is the same as Figure 58B the same.
[0821] First, in Figure 58AIn step 1 as shown, the inter-frame predictor 218 calculates the cost between the search position (also known as the starting point) indicated by the initial MV and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 changes the target to the search position that obtains the minimum cost and performs the process in step 2 shown in FIG. 58. When the cost at the starting point is the minimum, the inter-frame predictor 218 skips Figure 58A the process in step 2 shown in
[0822] and performs the process in step 3. Figure 58A In step 2 as shown, the inter-frame predictor 218 performs a search similar to the process in step 1, and regards the search position after the target is changed as the new starting point according to the result of the process in step 1. Then, the inter-frame predictor 218 determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 performs the process in step 4. When the cost at the starting point is the minimum, the inter-frame predictor 218 performs the process in step 3.
[0823] In step 4, the inter-frame predictor 218 regards the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0824] In Figure 58A step 3 as shown, the inter-frame predictor 218 determines the pixel position at the sub-pixel accuracy that obtains the minimum cost based on the costs at four points located at the upper, lower, left, and right positions relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position.
[0825] The pixel position at the sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) of up, down, left, and right by using the cost at the corresponding search position among the four search positions as the weight. The inter-frame predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0826] (Motion Compensation > BIO / OBMC / LIC)
[0827] For example, when the information parsed from the stream indicates that the predicted image is to be corrected, when the predicted image is generated, the inter-frame predictor 218 corrects the predicted image based on the correction mode. This mode is, for example, one of the above-mentioned BIO, OBMC, and LIC.
[0828] Figure 92It is a flowchart showing an example of the process of generating a predicted image in the decoder 200.
[0829] The inter-frame predictor 218 generates a predicted image (step Sm_11) and corrects the predicted image according to any of the above patterns (step Sm_12).
[0830] Figure 93 It is a flowchart showing another example of the process of generating a predicted image in the decoder 200.
[0831] The inter-frame predictor 218 derives the MV of the current block (step Sn_11). Next, the inter-frame predictor 218 generates a predicted image using the MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, the inter-frame predictor 218 obtains the prediction parameters included in the stream and determines whether to perform a correction process based on the prediction parameters. For example, the prediction parameter is a flag indicating whether one or more of the above patterns are to be applied. Here, when it is determined to perform the correction process (Yes in step Sn_13), the inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). Note that in LIC, the luminance and chrominance can be corrected in step Sn_14. When it is determined not to perform the correction process (No in step Sn_13), the inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).
[0832] (Motion Compensation>OBMC)
[0833] For example, when the information parsed from the stream indicates that OBMC is to be performed, when the predicted image is generated, the inter-frame predictor 218 corrects the predicted image according to OBMC.
[0834] Figure 94 It is a flowchart showing an example of the process of correcting a predicted image by OBMC in the decoder 200. Note that Figure 94 the flowchart in Figure 62 shows the correction process of the predicted image using the current picture and the reference picture shown in
[0835] First, as Figure 62 shown, a predicted image (Pred) is obtained by conventional motion compensation using the MV assigned to the current block.
[0836] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) that has been derived for the decoded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). Then, the inter-frame predictor 218 performs a first correction of the predicted image by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0837] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) that has been derived for a decoded block adjacent above the current block to the current block (reusing the motion vector for the current block). Then, the inter-frame predictor 218 performs a second correction of the predicted image by overlapping the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) for which the first correction has already been performed. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and thus is the final predicted image of the current block.
[0838] (Motion Compensation > BIO)
[0839] For example, when the information parsed from the stream indicates that BIO is to be performed, when the predicted image has been generated, the inter-frame predictor 218 corrects the predicted image according to BIO.
[0840] Figure 95 is a flowchart showing an example of the process of correcting the predicted image performed by BIO in the decoder 200.
[0841] As Figure 63 shown, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter-frame predictor 218 derives the predicted image of the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is the motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is the motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.
[0842] Next, the inter-frame predictor 218 derives the interpolated image I 0 of the current block using the motion vector M0 and the reference picture L0. Additionally, the inter-frame predictor 218 derives the interpolated image I 1 of the current block using the motion vector M1 and the reference picture L1 (step Sy_12). Here, the interpolated image I 0 is the image included in the reference picture Ref0 and to be derived for the current block, and the interpolated image I 1 is the image included in the reference picture Ref1 and derived for the current block. Each of the interpolated image I 0 and the interpolated image I 1 may be the same size as the current block. Alternatively, the interpolated image I0 and the interpolated image I 1 each in may be an image larger than the current block. Further, the interpolated image I 0 and the interpolated image I 1 may include a predicted image obtained by using motion vectors (M0, M1) and reference pictures (L0, L1) and applying a motion compensation filter.
[0843] Further, the inter-frame predictor 218 derives a gradient image (Ix 0 and the interpolated image I 1 of the current block from the interpolated image I 0 Ix 1 Iy 0 Iy 1 )(step Sy_13). Note that the gradient image in the horizontal direction is (Ix 0 Ix 1 ), and the gradient image in the vertical direction is (Iy 0 Iy 1 ). The inter-frame predictor 218 may derive the gradient image by applying, for example, a gradient filter to the interpolated image. The gradient image may be an image in which each indicates the amount of spatial change of the pixel value in the horizontal direction or the amount of spatial change of the pixel value in the vertical direction.
[0844] Next, the inter-frame predictor 218 uses the interpolated images (I 0 I 1 ) and the gradient images (Ix 0 Ix 1 Iy 0 Iy 1 ) to derive an optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_14). As an example, the sub-block may be a 4×4 pixel sub-CU.
[0845] Next, the inter-frame predictor 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame predictor 218 derives a corrected value of the pixel values included in the current block using the optical flow (vx, vy) (step Sy_15). The inter-frame predictor 218 then corrects the predicted image of the current block using the corrected value (step Sy_16). Note that the corrected value may be derived in units of pixels, or may be derived in units of multiple pixels or in units of sub-blocks, etc.
[0846] Note that the BIO process flow is not limited to the Figure 95 process disclosed in. It is possible to execute only Figure 95 a part of the process disclosed in, or different processes may be added or used as an alternative, or the processes may be executed in a different processing order.
[0847] (Motion Compensation > LIC)
[0848] For example, when the information parsed from the stream indicates that LIC is to be performed, when a predicted image is generated, the inter - frame predictor 218 corrects the predicted image according to LIC.
[0849] Figure 96 is a flowchart showing an example of the process of correcting a predicted image by LIC in the decoder 200.
[0850] First, the inter - frame predictor 218 obtains a reference image corresponding to the current block from the decoded reference pictures using the MV (step Sz_11).
[0851] Next, the inter - frame predictor 218 extracts information indicating how the luminance values change between the current picture and the reference picture for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values of the decoded left - adjacent reference region (surrounding reference region) and the decoded upper - adjacent reference region (surrounding reference region), as well as the luminance pixel values at the corresponding positions in the reference picture specified by the derived MV. The inter - frame predictor 218 uses the information indicating how the luminance values change to calculate a luminance correction parameter (step Sz_13).
[0852] The inter - frame predictor 218 generates a predicted image for the current block by performing a luminance correction process in which the luminance correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_14). In other words, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the chrominance can be corrected.
[0853] (Prediction Controller)
[0854] The prediction controller 220 selects an intra - frame predicted image or an inter - frame predicted image and outputs the selected image to the adder 208. Generally speaking, the configurations, functions, and processes of the prediction controller 220, the intra - frame predictor 216, and the inter - frame predictor 218 on the decoder 200 side can correspond to the configurations, functions, and processes of the prediction controller 128, the intra - frame predictor 124, and the inter - frame predictor 126 on the encoder 100 side.
[0855] (First Aspect)
[0856] Figure 97 is a flowchart showing an example of the process flow 1000 for decoding an image using the CCALF (Cross - Component Adaptive Loop Filter) process according to the first aspect. The process flow 1000 can be executed, for example, by Figure 67 the decoder 200, etc.
[0857] In step S1001, a filtering process is applied to the reconstructed image samples of the first component. For example, the first component can be a luminance component. The luminance component can be represented as the Y component. The reconstructed image samples of luminance can be the output signal of the ALF process. The output signal of the ALF can be the reconstructed luminance samples generated by the SAO process. In some embodiments, the filtering process performed in step S1001 can be represented as the CCALF process. The number of reconstructed luminance samples can be the same as the number of coefficients of the filter to be used in the CCALF process. In other embodiments, a clipping process can be performed on the filtered reconstructed luminance samples.
[0858] In step S1002, the reconstructed image samples of the second component are modified. The second component can be a chrominance component. The chrominance component can be represented as the Cb and / or Cr components. The reconstructed image samples of chrominance can be the output signal of the ALF process. The output signal of the ALF can be the reconstructed chrominance samples generated by the SAO process. The modified reconstructed image samples can be the sum of the reconstructed samples of chrominance and the filtered reconstructed samples of luminance, i.e., the output of step S1001. In other words, the modification process can be performed by adding the filtered values of the reconstructed luminance samples generated by the CCALF process in step S1001 to the filtered values of the reconstructed chrominance samples generated by the ALF process. In some embodiments, a clipping process can be performed on the reconstructed chrominance samples. The first component and the second component can belong to the same block, or can belong to different blocks.
[0859] In step S1003, the values of the modified reconstructed image samples of the chrominance component are clipped. By performing the clipping process, it can be ensured that the values of the samples are within a certain range. Further, clipping can promote better convergence in processes such as least squares optimization to minimize the difference between the residual (the difference between the original sample value and the reconstructed sample value) and the filtered value of the chrominance samples, so as to facilitate the determination of the filter coefficients.
[0860] In step S1004, the image is decoded using the clipped reconstructed image samples of the chrominance component. In some embodiments, step S1003 does not need to be performed. In this case, the image is decoded using the unclipped modified reconstructed chrominance samples.
[0861] Figure 98 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment. In this embodiment, a clipping process is applied to the modified reconstructed image samples of the chrominance component, as Figure 97In step S1003 of. For example, for a 10-bit output, the modified reconstructed image sample can be cropped to the range of [0, 1023]. In some embodiments, when cropping the filtered reconstructed image sample of the luminance component generated by the CCALF process, it may not be necessary to crop the modified reconstructed image sample of the chrominance component.
[0862] Figure 99 is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment. In this embodiment, as Figure 97 in step S1003, a cropping process is applied to the modified reconstructed image sample of the chrominance component. The cropping process does not apply to the filtered reconstructed luminance samples generated by the CCALF process. The filtered values of the reconstructed chrominance samples generated by the ALF process do not need to be cropped, as Figure 99 shown as "no cropping" in. In other words, the reconstructed image sample to be modified is generated using the filtered value (ALF chrominance) and the difference (CCALF Cb / Cr), and no cropping is applied to the output of the generated sample values.
[0863] Figure 100 is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment. In this embodiment, the cropping process is applied to the filtered reconstructed luminance samples ("cropped output samples") generated by the CCALF process and the modified reconstructed image samples of the chrominance component ("cropped after summation"). The filtered values of the reconstructed chrominance samples generated by the ALF process are not cropped ("no cropping"). For example, the cropping range applied to the filtered reconstructed image sample of the luminance component can be [-2^15, 2^15 - 1] or [-2^7, 2^7 - 1].
[0864] Figure 101 shows another example where the cropping process is applied to: the filtered reconstructed luminance samples ("cropped output samples") generated by the CCALF process, the modified reconstructed image samples of the chrominance component ("cropped after summation"), and the filtered reconstructed chrominance samples ("cropped") generated by the ALF process. In other words, the output values of the CCALF process and the ALF chrominance process are cropped separately and then cropped again after they are summed. In this embodiment, the modified reconstructed image sample of the chrominance component does not need to be cropped. For example, the final output of the ALF chrominance process may be cropped to a 10-bit value. For example, the cropping range applied to the filtered reconstructed image sample of the luminance component can be [-2^15, 2^15 - 1] or [-2^7, 2^7 - 1]. This range can be fixed or can be determined adaptively. In either case, this range can be signaled in the header information, for example, in the SPS (Sequence Parameter Set) or APS (Adaptation Parameter Set). In the case when using non-linear ALF, it can be forFigure 101 The "clipping after summation" in Figure 101 defines the clipping parameters.
[0865] The reconstructed image samples of the luminance component to be filtered by the CCALF process can be neighboring samples adjacent to the current reconstructed image samples of the chrominance component. That is, a modified current reconstructed image sample can be generated by adding the filtered value of the neighboring image samples of the luminance component located adjacent to the current image sample to the filtered value of the current image sample of the chrominance component. The filtered value of the image samples of the luminance component can be represented as a difference.
[0866] The process disclosed in this aspect can reduce the internal hardware memory size required to store the filtered image sample values.
[0867] (Second aspect)
[0868] Figure 102 is a flowchart of an example of the process flow 2000 for decoding an image by applying the CCALF process using the information defined in the second aspect. The process flow 2000 can be executed, for example, by Figure 67 a decoder 200 of
[0869] In step S2001, the clipping parameters are parsed from the bitstream. The clipping parameters can be parsed from the VPS, APS, SPS, PPS, slice header at the CTU or TU level, as Figure 103 described in Figure 103 is a conceptual diagram showing the location of the clipping parameters. Figure 103 The parameters described in
[0870] Figures 98 - 101 can be replaced by different types of clipping parameters, flags, or indices. Two or more clipping parameters can be parsed from two or more parameter sets in the bitstream. In step S2002, the difference is clipped using the clipping parameters. A difference is generated based on the reconstructed image samples of the first component (e.g.,
[0871] the difference (CCALF Cb / Cr) in
[0872] For example, the first component is the luminance component, and the difference is the filtered reconstructed luminance sample generated by the CCALF process. In this case, the clipping process is applied to the filtered reconstructed luminance sample using the parsed clipping parameters. Figure 104(i) as shown. In this example, ccalf_luma_clip_idx[] is the index, -range_array[] is the lower limit, and range_array[] is the upper limit. In this example, range_array[] is the determined range array, which may be different from the range array used for ALF. The determined range array can be pre-determined.
[0873] The clipping parameter can indicate the lower and upper limits, as Figure 104 (ii) as shown. In this example, -ccalf_luma_clip_low_range[] is the lower limit range, and ccalf_luma_clip_up_range[] is the upper limit range.
[0874] The clipping parameter can indicate the common range of both the lower limit range and the upper limit range, as Figure 104 (iii) as shown. In this example, -ccalf_luma_clip_range is the lower limit, and ccalf_luma_clip_range is the upper limit.
[0875] Generate a difference by multiplying, dividing, adding, or subtracting at least two reconstructed image samples of the first component. For example, the two reconstructed image samples can be from the current and adjacent image samples or two adjacent image samples. The positions of the current and adjacent image samples can be pre-determined.
[0876] In step S2003, modify the reconstructed image samples of the second component different from the first component using the clipped value. The clipped value can be the clipped value of the reconstructed image samples of the luminance component. The second component can be the chrominance component. The modification can include operations for multiplying, dividing, adding, or subtracting the clipped value with respect to the reconstructed image samples of the second component.
[0877] In step S2004, decode the image using the modified reconstructed image samples.
[0878] In the present disclosure, one or more clipping parameters for cross-component adaptive loop filtering are signaled in the bitstream. Through this signaling, the syntax of cross-component adaptive loop filtering and the syntax of the adaptive loop filter can be combined for syntax simplification. In addition, through this signaling, the design of cross-component adaptive loop filtering can be more flexible to improve the coding efficiency.
[0879] Clipping parameters can be defined or predefined for both the encoder and decoder without signaling. The clipping parameters can also be derived using luminance information without signaling. For example, if strong gradients or edges are detected in the luminance reconstructed image, clipping parameters corresponding to a large clipping range can be derived, and if weak gradients or edges are detected in the luminance reconstructed image, clipping parameters corresponding to a short clipping range can be derived.
[0880] (Third aspect)
[0881] Figure 105 FIG. 3000 is a flowchart of an example of a process flow 3000 for decoding an image by applying a CCALF process using filter coefficients according to the third aspect. The process flow 3000 can be executed, for example, by Figure 67 a decoder 200 or the like. Filter coefficients are used in the filtering step of the CCALF process to generate filtered reconstructed image samples of the luminance component.
[0882] In step S3001, it is determined whether the filter coefficients are within a defined symmetric region of the filter. Optionally, an additional step of determining whether the shape of the filter coefficients is symmetric can be performed. Information indicating whether the samples of the filter coefficients are symmetric can be encoded into the bitstream. If the shape is symmetric, the positions of the coefficients within the symmetric region can be determined or predetermined.
[0883] In step S3002, if the filter coefficients are within the defined symmetric region (''Yes'' in step S3001), the filter coefficients are copied to symmetric positions and a set of filter coefficients is ge...
Claims
1. A coding method, comprising: In response to a first reconstructed image sample being outside a virtual boundary, copying a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample; Generating a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to the first reconstructed image sample of a luminance component, the first reconstructed image sample being located adjacent to a second reconstructed image sample of a chrominance component; In response to the first coefficient value being less than 64, setting the first coefficient value to zero; Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image sample of the chrominance component; Generating a third coefficient value by adding the first coefficient value and the second coefficient value; And Outputting a third reconstructed image sample of the chrominance component using the third coefficient value.
2. A decoding method, comprising: In response to a first reconstructed image sample being outside a virtual boundary, copying a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample; Generating a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to the first reconstructed image sample of a luminance component, the first reconstructed image sample being located adjacent to a second reconstructed image sample of a chrominance component; In response to the first coefficient value being less than 64, setting the first coefficient value to zero; Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image sample of the chrominance component; Generating a third coefficient value by adding the first coefficient value and the second coefficient value; And Outputting a third reconstructed image sample of the chrominance component using the third coefficient value.
3. A non-transitory computer-readable medium storing a bitstream, characterized in that The bitstream includes filtering information that causes a decoder to perform a filtering process, the filtering process comprising: In response to a first reconstructed image sample being outside a virtual boundary, copying a reconstructed sample that is inside the virtual boundary and adjacent to the virtual boundary to generate the first reconstructed image sample; Generating a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to the first reconstructed image sample of a luminance component, the first reconstructed image sample being located adjacent to a second reconstructed image sample of a chrominance component; In response to the first coefficient value being less than 64, setting the first coefficient value to zero; Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image sample of the chrominance component; Generating a third coefficient value by adding the first coefficient value and the second coefficient value; and Outputting a third reconstructed image sample of the chrominance component using the third coefficient value.