Systems and methods for video coding
By using the CCALF and ALF processes in the video encoder to encode video images, the problems of low encoding efficiency and poor image quality in the prior art are solved, and more efficient encoding and better image quality are achieved.
Patent Information
- Application Number
- CN202510332443.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-08
- Filing Date
- 2020-08-07
- Publication Date
- 2025-05-13
AI Technical Summary
When existing video encoding technology processes high resolution and large amounts of data, it has low encoding efficiency, poor image quality, large circuit scale and slow processing speed.
An encoder is adopted to generate coefficient values by applying a CCALF process to the reconstructed image samples of the luminance component, and to apply an ALF process to the reconstructed image samples of the chrominance component, trimming and summing coefficient values for encoding.
It improves encoding efficiency, enhances image quality, reduces circuit scale, and improves encoding and decoding processing speed.
Smart Images

Figure CN119996683A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent filed on August 7, 2020, with application number 202080051926.7 and title "System and Method for Video Coding". Technical Field
[0002] This disclosure relates to video coding, and more particularly to video coding and decoding systems, components and methods in video coding and decoding, such as those for performing the CCALF (Cross Component Adaptive Loop Filtering) process. Background Technology
[0003] With advancements in video coding technologies, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High-Efficiency Video Coding), and H.266 / VVC (Multi-Functional Video Codec), there remains a need for continuous improvement and optimization of video coding techniques to handle the ever-increasing volume of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in the CCALF (Cross-Component Adaptive Loop Filtering) process. Summary of the Invention
[0004] According to one aspect, an encoder is provided, comprising circuitry and a memory coupled to the circuitry. In operation, the circuitry generates first coefficient values by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of the luminance component. The circuitry generates second coefficient values by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of the chrominance component, and clips the second coefficient values. The circuitry generates third coefficient values by adding the first coefficient values to the clipped second coefficient values, and clips the third coefficient values. The circuitry encodes a third reconstructed image sample of the chrominance component using the clipped third coefficient values.
[0005] According to another aspect, the first reconstructed image sample is located adjacent to the second reconstructed image sample.
[0006] According to another aspect, the circuit sets the first coefficient value to zero in operation in response to the first coefficient value being less than 64.
[0007] According to another aspect, an encoder is provided, comprising: a block segmenter that segments a first image into a plurality of blocks in operation; an intra-frame predictor that predicts blocks included in the first image using reference blocks included in the first image in operation; an inter-frame predictor that predicts blocks included in the first image using reference blocks included in a second image different from the first image in operation; a loop filter that filters the blocks included in the first image in operation; a transformer that transforms a prediction error between an original signal and a prediction signal generated by the intra-frame predictor or the inter-frame predictor in operation to generate transform coefficients; a quantizer that quantizes the transform coefficients in operation to generate quantized coefficients; and an entropy encoder that variably encodes the quantized coefficients in operation to generate an encoded bitstream including the encoded quantized coefficients and control information. The loop filter performs the following operations:
[0008] The first coefficient values are generated by applying the CCALF (Cross-Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component;
[0009] The second coefficient values are generated by applying the ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma components;
[0010] Trim the second coefficient value;
[0011] A third coefficient value is generated by adding the first coefficient value to the cropped second coefficient value;
[0012] Trim the third coefficient value; and
[0013] The third reconstructed image sample of the chroma component is encoded using the cropped third coefficient values.
[0014] According to another aspect, a decoder is provided, comprising circuitry and a memory coupled to the circuitry. In operation, the circuitry generates first coefficient values by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of the luminance component. The circuitry generates second coefficient values by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of the chrominance component, and then clips the second coefficient values. The circuitry generates third coefficient values by adding the first coefficient values to the clipped second coefficient values, and then clips the third coefficient values. The circuitry uses the clipped third coefficient values to decode a third reconstructed image sample of the chrominance component.
[0015] According to another aspect, a decoding apparatus is provided, comprising: a decoder that decodes an encoded bitstream in operation to output quantized coefficients; an inverse quantizer that in operation inversely quantizes the quantized coefficients to output transform coefficients; an inverse transformer that in operation inversely transforms the transform coefficients to output a prediction error; an intra-frame predictor that in operation uses a reference block included in a first image to predict a block included in the first image; an inter-frame predictor that in operation uses a reference block included in a second image different from the first image to predict a block included in the first image; a loop filter that in operation filters the block included in the first image; and an output that in operation outputs an image including the first image. The loop filter performs the following operations:
[0016] The first coefficient values are generated by applying the CCALF (Cross-Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component;
[0017] The second coefficient values are generated by applying the ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma components;
[0018] Trim the second coefficient value;
[0019] A third coefficient value is generated by adding the first coefficient value to the trimmed second coefficient value;
[0020] Trim the third coefficient value; and
[0021] The third reconstructed image sample of the chroma component is decoded using the cropped third coefficient value.
[0022] According to another aspect, an encoding method is provided, including:
[0023] The first coefficient values are generated by applying the CCALF (Cross-Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component;
[0024] The second coefficient values are generated by applying the ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma components;
[0025] Trim the second coefficient value;
[0026] A third coefficient value is generated by adding the first coefficient value to the trimmed second coefficient value;
[0027] Trim the third coefficient value; and
[0028] The cropped third coefficient values are used to encode the third reconstructed image sample of the chroma component.
[0029] According to another aspect, a decoding method is provided, including:
[0030] The first coefficient values are generated by applying the CCALF (Cross-Component Adaptive Loop Filtering) process to the first reconstructed image sample of the luminance component;
[0031] The second coefficient values are generated by applying the ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma components;
[0032] Trim the second coefficient value;
[0033] A third coefficient value is generated by adding the first coefficient value to the trimmed second coefficient value;
[0034] Trim the third coefficient value; and
[0035] The third reconstructed image sample of the chroma component is decoded using the cropped third coefficient values.
[0036] In video coding technology, new methods are needed to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of embodiments of this disclosure, including individual or combined elements of embodiments of this disclosure, can facilitate one or more of the following: improved coding efficiency, enhanced image quality, reduced utilization of processing resources associated with encoding / decoding, reduced circuit size, and increased encoding / decoding processing speed.
[0037] Furthermore, some implementations of the embodiments of this disclosure, including constituent elements of the embodiments of this disclosure considered individually or in various combinations, can facilitate the appropriate selection of one or more elements during encoding and decoding, such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations. Note that this disclosure includes disclosures regarding configurations and methods that can provide advantages beyond those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing the increase in processing resource usage.
[0038] Additional benefits and advantages of the disclosed embodiments will become apparent from the specification and accompanying drawings. Benefits and / or advantages may be obtained individually from the various embodiments and features in the specification and drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.
[0039] It should be noted that general or specific embodiments may be implemented as systems, methods, integrated circuits, computer programs, storage media, or any alternative combination thereof. Attached Figure Description
[0040] [ Figure 1 ] Figure 1This is a schematic diagram illustrating an example of the functional configuration of a transmission system according to an embodiment.
[0041] [ Figure 2 ] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream.
[0042] [ Figure 3 ] Figure 3 This is a conceptual diagram used to illustrate an example of slice configuration.
[0043] [ Figure 4 ] Figure 4 This is a conceptual diagram used to illustrate an example of a tile configuration.
[0044] [ Figure 5 ] Figure 5 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.
[0045] [ Figure 6 ] Figure 6 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.
[0046] [ Figure 7 ] Figure 7 This is a block diagram illustrating the functional configuration of an encoder according to an embodiment.
[0047] [ Figure 8 ] Figure 8 This is a functional block diagram illustrating an example of encoder installation.
[0048] [ Figure 9 ] Figure 9 It is a flowchart that indicates an example of the overall encoding process performed by the encoder.
[0049] [ Figure 10 ] Figure 10 This is a conceptual diagram illustrating an example of block partitioning.
[0050] [ Figure 11 ] Figure 11 This is a block diagram illustrating an example of the functional configuration of a splitter according to an embodiment.
[0051] [ Figure 12 ] Figure 12 This is a conceptual diagram used to illustrate an example of a splitting pattern.
[0052] [ Figure 13A ] Figure 13A This is a conceptual diagram used to illustrate an example of a syntactic tree for segmentation patterns.
[0053] [ Figure 13B ] Figure 13B This is a conceptual diagram used to illustrate another example of a syntactic tree for segmentation patterns.
[0054] [ Figure 14 ] Figure 14 It is a graph indicating example transformation basis functions used for various transformation types.
[0055] [ Figure 15 ] Figure 15 This is a conceptual diagram used to illustrate an example spatial transformation (SVT).
[0056] [ Figure 16 ] Figure 16 This is a flowchart illustrating an example of a process performed by a converter.
[0057] [ Figure 17 ] Figure 17 This is a flowchart illustrating another example of a process performed by a converter.
[0058] [ Figure 18 ] Figure 18 This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment.
[0059] [ Figure 19 ] Figure 19 This is a flowchart illustrating an example of the quantization process performed by a quantizer.
[0060] [ Figure 20 ] Figure 20 This is a block diagram illustrating an example of the functional configuration of an entropy encoder according to an embodiment.
[0061] [ Figure 21 ] Figure 21 This is a conceptual diagram illustrating an example flow of the context-based adaptive binary arithmetic coding (CABAC) process in an entropy encoder.
[0062] [ Figure 22 ] Figure 22 This is a block diagram illustrating an example of the functional configuration of a loop filter according to an embodiment.
[0063] [ Figure 23A ] Figure 23A This is a conceptual diagram used to illustrate an example of the filter shape used in an adaptive loop filter (ALF).
[0064] [ Figure 23B ] Figure 23B This is a conceptual diagram used to illustrate another example of the filter shape used in ALF.
[0065] [ Figure 23C ] Figure 23CThis is a conceptual diagram used to illustrate another example of the filter shape used in ALF.
[0066] [ Figure 23D ] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF).
[0067] [ Figure 23E ] Figure 23E This is a conceptual diagram used to illustrate an example of the filter shape used in CC-ALF.
[0068] [ Figure 23F ] Figure 23F This is a conceptual diagram used to illustrate an example process of Joint Chromaticity CCALF (JC-CCALF).
[0069] [ Figure 23G ] Figure 23G This is a table showing example weight index candidates that can be used in JC-CCALF.
[0070] [ Figure 24 ] Figure 24 This is a block diagram illustrating an example of a specific configuration of a loop filter used as a deblocking filter (DBF).
[0071] [ Figure 25 ] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics about the block boundaries.
[0072] [ Figure 26 ] Figure 26 It is a conceptual diagram used to illustrate the block boundaries for performing the deblocking filtering process.
[0073] [ Figure 27 ] Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.
[0074] [ Figure 28 ] Figure 28 This is a flowchart illustrating an example of the process performed by the encoder's predictor.
[0075] [ Figure 29 ] Figure 29 This is a flowchart illustrating another example of the process performed by the encoder's predictor.
[0076] [ Figure 30 ] Figure 30 This is a flowchart illustrating another example of the process performed by the encoder's predictor.
[0077] [ Figure 31 ] Figure 31This is a conceptual diagram used to illustrate the sixty-seven intra-prediction modes used in the intra-prediction in the embodiments.
[0078] [ Figure 32 ] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor.
[0079] [ Figure 33 ] Figure 33 This is a concept diagram used to illustrate a reference image.
[0080] [ Figure 34 ] Figure 34 This is a concept diagram used to illustrate a list of reference images.
[0081] [ Figure 35 ] Figure 35 This is a flowchart illustrating the basic process flow of an example inter-frame prediction.
[0082] [ Figure 36 ] Figure 36 This is a flowchart illustrating an example of the derivation process of the motion vector.
[0083] [ Figure 37 ] Figure 37 This is a flowchart illustrating another example of the derivation process of the motion vector.
[0084] [ Figure 38A ] Figure 38A This is a conceptual diagram used to illustrate example representations of patterns used for MV derivation.
[0085] [ Figure 38B ] Figure 38B This is a conceptual diagram used to illustrate example representations of patterns used for MV derivation.
[0086] [ Figure 39 ] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in a regular inter-frame mode.
[0087] [ Figure 40 ] Figure 40 This is a flowchart illustrating an example of the inter-frame prediction process in a conventional merging mode.
[0088] [ Figure 41 ] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process in a merging mode.
[0089] [ Figure 42 ] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process for the current image using the HMVP merging pattern.
[0090] [ Figure 43 ] Figure 43 This is a flowchart illustrating an example of the Frame Rate Upconversion (FRUC) process.
[0091] [ Figure 44 ] Figure 44 This is a conceptual diagram used to illustrate an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.
[0092] [ Figure 45 ] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image.
[0093] [ Figure 46A ] Figure 46A This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block based on the motion vectors of multiple adjacent blocks.
[0094] [ Figure 46B ] Figure 46B This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block in an affine pattern that uses three control points.
[0095] [ Figure 47A ] Figure 47A This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.
[0096] [ Figure 47B ] Figure 47B This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.
[0097] [ Figure 47C ] Figure 47C This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.
[0098] [ Figure 48A ] Figure 48A This is a conceptual diagram used to illustrate an affine pattern that uses two control points.
[0099] [ Figure 48B ] Figure 48B This is a conceptual diagram used to illustrate an affine pattern that uses three control points.
[0100] [ Figure 49A ] Figure 49A This is a conceptual diagram illustrating an example of a method for deriving the MV at control points when the number of control points used for encoding a block and the number of control points used for the current block are different from each other.
[0101] [ Figure 49B ] Figure 49BThis is a conceptual diagram illustrating another example of a method for deriving the MV at control points when the number of control points used for encoding blocks and the number of control points used for the current block are different from each other.
[0102] [ Figure 50 ] Figure 50 This is a flowchart illustrating an example of the process in the affine merge pattern.
[0103] [ Figure 51 ] Figure 51 This is a flowchart illustrating an example of a process in affine inter-frame mode.
[0104] [ Figure 52A ] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangle prediction images.
[0105] [ Figure 52B ] Figure 52B This is a conceptual diagram used to illustrate the first part of the first partition that overlaps with the second partition, as well as an example of the first and second sample sets that can be weighted as part of the correction process.
[0106] [ Figure 52C ] Figure 52C This is a conceptual diagram used to illustrate the first part of the first partition, which is the portion of the first partition that overlaps with a portion of an adjacent partition.
[0107] [ Figure 53 ] Figure 53 This is a flowchart illustrating an example of a process in a triangle pattern.
[0108] [ Figure 54 ] Figure 54 This is a conceptual diagram used to illustrate an example of an Advanced Temporal Motion Vector Prediction (ATMVP) pattern in which MV is derived on a sub-block basis.
[0109] [ Figure 55 ] Figure 55 This is a flowchart illustrating the relationship between merge mode and Dynamic Motion Vector Refresh (DMVR).
[0110] [ Figure 56 ] Figure 56 This is a conceptual diagram used to illustrate an example of DMVR.
[0111] [ Figure 57 ] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV.
[0112] [ Figure 58A ] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in DMVR.
[0113] [ Figure 58B ] Figure 58B This is a flowchart illustrating an example of the motion estimation process in a DMVR.
[0114] [ Figure 59 ] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.
[0115] [ Figure 60 ] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.
[0116] [ Figure 61 ] Figure 61 This is a flowchart illustrating an example of the correction process for a predicted image using Overlapping Block Motion Compensation (OBMC).
[0117] [ Figure 62 ] Figure 62 This is a conceptual diagram used to illustrate an example of the predictive image correction process via OBMC.
[0118] [ Figure 63 ] Figure 63 It is a conceptual diagram used to illustrate a model that assumes uniform linear motion.
[0119] [ Figure 64 ] Figure 64 This is a flowchart illustrating an example of the inter-frame prediction process based on BIO.
[0120] [ Figure 65 ] Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor that can perform inter-frame prediction based on BIO.
[0121] [ Figure 66A ] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses a brightness correction process performed by a LIC.
[0122] [ Figure 66B ] Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.
[0123] [ Figure 67 ] Figure 67 This is a block diagram illustrating the functional configuration of the decoder according to an embodiment.
[0124] [ Figure 68 ] Figure 68This is a functional block diagram illustrating an example of decoder installation.
[0125] [ Figure 69 ] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by the decoder.
[0126] [ Figure 70 ] Figure 70 It is a conceptual diagram used to illustrate the relationship between the segmentation determiner and other constituent elements.
[0127] [ Figure 71 ] Figure 71 This is a block diagram illustrating an example of the functional configuration of an entropy decoder.
[0128] [ Figure 72 ] Figure 72 This is a conceptual diagram used to illustrate an example flow of the CABAC process in an entropy decoder.
[0129] [ Figure 73 ] Figure 73 This is a block diagram illustrating an example of the functional configuration of an inverse quantizer.
[0130] [ Figure 74 ] Figure 74 This is a flowchart illustrating an example of the inverse quantization process performed by an inverse quantizer.
[0131] [ Figure 75 ] Figure 75 This is a flowchart illustrating an example of a process performed by an inverse transformer.
[0132] [ Figure 76 ] Figure 76 This is a flowchart illustrating another example of the process performed by the inverse converter.
[0133] [ Figure 77 ] Figure 77 This is a block diagram illustrating an example of the functional configuration of a loop filter.
[0134] [ Figure 78 ] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of the decoder.
[0135] [ Figure 79 ] Figure 79 This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0136] [ Figure 80A ] Figure 80A This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0137] [ Figure 80B ] Figure 80B This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0138] [ Figure 80C ] Figure 80C This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0139] [ Figure 81 ] Figure 81 This is a diagram illustrating an example of the process performed by the decoder's intra-frame predictor.
[0140] [ Figure 82 ] Figure 82 This is a flowchart illustrating an example of the MV derivation process in the decoder.
[0141] [ Figure 83 ] Figure 83 This is a flowchart illustrating another example of the MV derivation process in the decoder.
[0142] [ Figure 84 ] Figure 84 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using a regular inter-frame mode.
[0143] [ Figure 85 ] Figure 85 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using a regular merging mode.
[0144] [ Figure 86 ] Figure 86 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using FRUC mode.
[0145] [ Figure 87 ] Figure 87 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using an affine merging mode.
[0146] [ Figure 88 ] Figure 88 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using affine inter-frame patterns.
[0147] [ Figure 89 ] Figure 89 This is a flowchart illustrating an example of the process of inter-frame prediction using triangular patterns in the decoder.
[0148] [ Figure 90 ] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed by DMVR in the decoder.
[0149] [ Figure 91 ] Figure 91 This is a flowchart illustrating an example process of motion estimation via DMVR in the decoder.
[0150] [ Figure 92 ] Figure 92 This is a flowchart illustrating an example of the process of generating a predicted image in a decoder.
[0151] [ Figure 93 ] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in the decoder.
[0152] [ Figure 94 ] Figure 94 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via OBMC.
[0153] [ Figure 95 ] Figure 95 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via BIO.
[0154] [ Figure 96 ] Figure 96 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via LIC.
[0155] [ Figure 97 ] Figure 97 It is a flowchart of the sample process for decoding an image by applying the CCALF (Cross Component Adaptive Loop Filtering) process according to the first aspect.
[0156] [ Figure 98 ] Figure 98 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment.
[0157] [ Figure 99 ] Figure 99 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment.
[0158] [ Figure 100 ] Figure 100 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment.
[0159] [ Figure 101 ] Figure 101 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment.
[0160] [ Figure 102 ] Figure 102 This is a flowchart of a sample process for decoding an image based on the CCALF process applied in the second aspect.
[0161] [ Figure 103 ] Figure 103 The diagram illustrates the sample locations of clipping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice header, CTU, or TU of a bitstream.
[0162] [ Figure 104 ] Figure 104 An example of clipping parameters is illustrated.
[0163] [ Figure 105 ] Figure 105 This is a flowchart of the sample process for decoding an image using the CCALF procedure based on the filter coefficients from the third aspect.
[0164] [ Figure 106 ] Figure 106 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0165] [ Figure 107 ] Figure 107 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0166] [ Figure 108 ] Figure 108 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0167] [ Figure 109 ] Figure 109 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0168] [ Figure 110 ] Figure 110 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0169] [ Figure 111 ] Figure 111 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0170] [ Figure 112 ] Figure 112 This is a conceptual diagram indicating the location of filter coefficients that will be used in the CCALF process.
[0171] [ Figure 113 ] Figure 113 This is a block diagram illustrating the functional configuration of the CCALF process performed by the encoder and decoder according to an embodiment.
[0172] [ Figure 114 ] Figure 114This is a flowchart of a sample process for decoding an image by applying the CCALF process using a filter selected from multiple filters, according to the fourth aspect.
[0173] [ Figure 115 ] Figure 115 The diagram illustrates an example of the filter selection process.
[0174] [ Figure 116 ] Figure 116 An example of a filter is illustrated.
[0175] [ Figure 117 ] Figure 117 An example of a filter is illustrated.
[0176] [ Figure 118 ] Figure 118 This is a flowchart of a sample process for decoding an image by applying the CCALF procedure using parameters, based on the fifth aspect.
[0177] [ Figure 119 ] Figure 119 The illustration shows an example of the number of coefficients to be parsed from a bitstream.
[0178] [ Figure 120 ] Figure 120 This is a flowchart of a sample process for decoding an image by applying the CCALF procedure using parameters, based on the sixth aspect.
[0179] [ Figure 121 ] Figure 121 This is a conceptual diagram illustrating an example of generating the CCALF value of the luminance component of the current chroma sample by calculating a weighted average of neighboring samples.
[0180] [ Figure 122 ] Figure 122 This is a conceptual diagram illustrating an example of generating the CCALF value of the luminance component of the current chroma sample by calculating a weighted average of neighboring samples.
[0181] [ Figure 123 ] Figure 123 This is a conceptual diagram illustrating an example of generating the CCALF value of the luminance component of the current chroma sample by calculating a weighted average of neighboring samples.
[0182] [ Figure 124 ] Figure 124 This is a conceptual diagram illustrating an example of generating the CCALF value of the luminance component of the current sample by calculating a weighted average of neighboring samples, where the positions of neighboring samples are adaptively determined as chromaticity types.
[0183] [ Figure 125 ] Figure 125This is a conceptual diagram illustrating an example of generating the CCALF value of the luminance component of the current sample by calculating a weighted average of neighboring samples, where the positions of neighboring samples are adaptively determined according to the chromaticity type.
[0184] [ Figure 126 ] Figure 126 This is a conceptual diagram illustrating an example of generating CCALF values for the luminance component by applying displacement bits to the output value of a weighted calculation.
[0185] [ Figure 127 ] Figure 127 This is a conceptual diagram illustrating an example of generating CCALF values for the luminance component by applying displacement bits to the output value of a weighted calculation.
[0186] [ Figure 128 ] Figure 128 This is a flowchart of a sample process for decoding an image by applying the CCALF procedure using parameters, based on the seventh aspect.
[0187] [ Figure 129 ] Figure 129 The diagram shows the sample locations of one or more parameters to be parsed from the bit stream.
[0188] [ Figure 130 ] Figure 130 The sample procedure for retrieving one or more parameters is shown.
[0189] [ Figure 131 ] Figure 131 Sample values for the second parameter are shown.
[0190] [ Figure 132 ] Figure 132 An example of parsing the second parameter using arithmetic encoding is shown.
[0191] [ Figure 133 ] Figure 133 This is a conceptual diagram of a variation of this embodiment applied to rectangular partitions and non-rectangular partitions (e.g., triangular partitions).
[0192] [ Figure 134 ] Figure 134 This is a flowchart of an example process for decoding an image by applying the CCALF procedure using parameters, according to aspect eight.
[0193] [ Figure 135 ] Figure 135 This is a flowchart of a sample process for decoding an image by applying the CCALF procedure using parameters, based on aspect eight.
[0194] [ Figure 136 ] Figure 136 Example locations for chroma sample types 0 to 5 are shown.
[0195] [ Figure 137 ] Figure 137 This is a conceptual diagram illustrating symmetrical filling of samples.
[0196] [ Figure 138 ] Figure 138 This is a conceptual diagram illustrating symmetrical filling of samples.
[0197] [ Figure 139 ] Figure 139 This is a conceptual diagram illustrating symmetrical filling of samples.
[0198] [ Figure 140 ] Figure 140 This is a conceptual diagram illustrating asymmetric filling of samples.
[0199] [ Figure 141 ] Figure 141 This is a conceptual diagram illustrating asymmetric filling of samples.
[0200] [ Figure 142 ] Figure 142 This is a conceptual diagram illustrating asymmetric filling of samples.
[0201] [ Figure 143 ] Figure 143 This is a conceptual diagram illustrating asymmetric filling of samples.
[0202] [ Figure 144 ] Figure 144 This is a conceptual diagram illustrating further sample asymmetric filling.
[0203] [ Figure 145 ] Figure 145 This is a conceptual diagram illustrating further sample asymmetric filling.
[0204] [ Figure 146 ] Figure 146 This is a conceptual diagram illustrating further sample asymmetric filling.
[0205] [ Figure 147 ] Figure 147 This is a conceptual diagram illustrating further sample asymmetric filling.
[0206] [ Figure 148 ] Figure 148 This is a conceptual diagram illustrating further sample symmetry filling.
[0207] [ Figure 149 ] Figure 149 This is a conceptual diagram illustrating further sample symmetry filling.
[0208] [ Figure 150 ] Figure 150 This is a conceptual diagram illustrating further sample symmetry filling.
[0209] [ Figure 151 ] Figure 151 This is a conceptual diagram illustrating further sample asymmetric filling.
[0210] [ Figure 152 ] Figure 152 This is a conceptual diagram illustrating further sample asymmetric filling.
[0211] [ Figure 153 ] Figure 153 This is a conceptual diagram illustrating further sample asymmetric filling.
[0212] [ Figure 154 ] Figure 154 This is a conceptual diagram illustrating further sample asymmetric filling.
[0213] [ Figure 155 ] Figure 155 A further example of padding with horizontal and vertical virtual boundaries is illustrated.
[0214] [ Figure 156 ] Figure 156 This is a block diagram illustrating the functional configuration of the encoder and decoder according to the example, where symmetrical padding is used on the virtual boundary locations of the ALF and symmetrical or asymmetrical padding is used on the virtual boundary locations of the CC-ALF.
[0215] [ Figure 157 ] Figure 157 This is a block diagram illustrating the functional configuration of the encoder and decoder according to another example, where symmetrical padding is used on the virtual boundary locations of the ALF and unilateral padding is used on the virtual boundary locations of the CC-ALF.
[0216] [ Figure 158 ] Figure 158 This is a conceptual diagram showing an example of a single-sided fill with a horizontal or vertical virtual boundary.
[0217] [ Figure 159 ] Figure 159 This is a conceptual diagram illustrating an example of a single-sided fill with horizontal and vertical virtual boundaries.
[0218] [ Figure 160 ] Figure 160 This is a diagram illustrating an example overall configuration of a content delivery system used to implement a content distribution service.
[0219] [ Figure 161 ] Figure 161 This is a conceptual diagram illustrating an example of a display screen used to show a webpage.
[0220] [ Figure 162 ] Figure 162 This is a conceptual diagram illustrating an example of a display screen used to show a webpage.
[0221] [ Figure 163 ] Figure 163 This is a block diagram illustrating an example of a smartphone.
[0222] [ Figure 164 ] Figure 164 This is a block diagram illustrating an example of the functional configuration of a smartphone. Detailed Implementation
[0223] In the accompanying drawings, unless the context otherwise indicates, the same reference numerals denote similar elements. The size and relative position of the elements in the drawings are not necessarily drawn to scale.
[0224] Hereinafter, embodiments will be described with reference to the accompanying drawings. Note that the embodiments described below each illustrate a general or specific example. The numerical values, shapes, materials, components, arrangements and connections of components, steps, relationships and sequences of steps indicated in the following embodiments are merely examples and are not intended to limit the scope of the claims.
[0225] Embodiments of the encoder and decoder will now be described. These embodiments are examples of encoders and decoders, and the processes and / or configurations presented in the description of aspects of this disclosure can be applied to said encoder and decoder. The processes and / or configurations can also be implemented in encoders and decoders different from those according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following can be implemented:
[0226] (1) Any component of the encoder or decoder of an embodiment presented in the description of an aspect of this disclosure may be replaced or combined with another component presented anywhere in the description of an aspect of this disclosure.
[0227] (2) In the encoder or decoder according to the embodiment, any function or process performed by one or more components of the encoder or decoder may be arbitrarily changed, such as by adding, replacing, or removing a function or process. For example, any function or process may be replaced by or combined with another function or process presented anywhere in the description of this aspect of the disclosure.
[0228] (3) In the method implemented by the encoder or decoder according to the embodiments, any changes may be made, such as adding, replacing, and removing one or more processes included in the method. For example, any process in the method may be replaced by or combined with another process presented anywhere in the description of the aspects of this disclosure.
[0229] (4) One or more components included in the encoder or decoder according to the embodiment may be combined with components presented anywhere in the description of the aspects of this disclosure, may be combined with components presenting one or more functions presented anywhere in the description of the aspects of this disclosure, and may be combined with components that implement one or more processes implemented by components presented in the description of the aspects of this disclosure.
[0230] (5) A component that includes one or more functions of an encoder or decoder according to an embodiment, or a component that implements one or more processes of an encoder or decoder according to an embodiment, may be combined with or replace components presented anywhere in the description of this disclosure, or may be combined with or replace components that include one or more functions presented anywhere in the description of this disclosure, or may be combined with or replace components that implement one or more processes presented anywhere in the description of this disclosure.
[0231] (6) In a method implemented by an encoder or decoder according to an embodiment, any process included in the method may be replaced or combined with a process presented anywhere in the description of aspects of this disclosure or with any corresponding or equivalent process.
[0232] (7) One or more processes included in the method implemented by the encoder or decoder according to the embodiment may be combined with processes presented anywhere in the description of aspects of this disclosure.
[0233] (8) The implementation of the processes and / or configurations presented in the description of aspects of this disclosure is not limited to the encoder or decoder according to the embodiments. For example, the processes and / or configurations may be implemented in a device for a different purpose than the motion image encoder or motion image decoder disclosed in the embodiments.
[0234] (Terminology Definition)
[0235] The corresponding terms can be defined as follows, as indicated by the example below.
[0236] An image is a data unit configured with a set of pixels; it is a picture, or a block of data smaller than pixels. Besides video, images also include still images.
[0237] An image is an image processing unit configured with a set of pixels, and can also be called a frame or field. For example, an image can be in the form of a monochrome luminance sample array or a luminance sample array and two corresponding chrominance sample arrays in 4:2:0, 4:2:2 and 4:4:4 color formats.
[0238] A block is a processing unit, which is a defined set of pixels. Blocks can have any number of different shapes. For example, a block can be a rectangle of M×N (M columns × N rows) pixels, a square of M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, pieces, bricks, CTUs, superblocks, basic segmentation units, VPDUs, processing segmentation units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and subblocks. Blocks can take the form of an M×N sample array or an M×N transform coefficient array. For example, a block can be a square or rectangular pixel region comprising a luminance matrix and two chrominance matrices.
[0239] A pixel or sample is the smallest point in an image. A pixel or sample includes pixels at integer positions and pixels at sub-pixel positions, such as those generated based on pixels at integer positions.
[0240] Pixel values or sample values are the feature values of a pixel. Pixel values or sample values can include one or more of the following: luminance value, chrominance value, RGB grayscale level, depth value, binary value of zero or 1, etc.
[0241] Chroma, or chrominance, is the intensity of a color, usually represented by the symbols Cb and Cr. These symbols specify the values of an array of samples or a single sample value representing the value of one of two color difference signals associated with the primary color.
[0242] Luminance, or luminance, is the brightness of an image, usually represented by the symbol or subscript Y or L. These specify the values of an array of samples or a single sample value representing the value of a monochromatic signal associated with the primary color.
[0243] A flag consists of one or more bits that indicate the value of a parameter or index. A flag can be a binary flag, indicating the binary value of the flag, or it can indicate the non-binary value of a parameter.
[0244] Signals convey information, which is symbolized or encoded into the signal. Signals include discrete digital signals and continuous analog signals.
[0245] A stream or bitstream is a string of digital data. A stream or bitstream can be a single stream or multiple streams configured with hierarchical layers. Streams or bitstreams can be transmitted serially via a single transmission path or packet-based via multiple transmission paths.
[0246] Difference refers to various mathematical differences, such as simple difference (xy), absolute value of difference (|xy|), difference of squares (x^2-y^2), square root of difference (√(x–y)), weighted difference (ax-by: a and b are constants), offset difference (x-y+a: a is the offset), etc. In the case of scalars, simple difference is sufficient, and difference calculation is included.
[0247] The term "sum" refers to various mathematical sums, such as the simple sum (x+y), the absolute value of the sum (|x+y|), the sum of squares (x^2+y^2), the square root of the sum (√(x+y)), the weighted difference (ax+by: a and b are constants), the offset sum (x+y+a: a is the offset), etc. In the case of scalars, the simple sum is sufficient, and sum calculations are included.
[0248] A frame is a combination of a top field and a bottom field, where sample lines 0, 2, 4, ... originate from the top field, and sample lines 1, 3, 5, ... originate from the bottom field.
[0249] A slice is an integer number of coded tree units contained in an independent slice and all subsequent dependent slices (if any) before the next independent slice (if any) within the same access unit.
[0250] A slice is a rectangular region of a coded tree block within a specific slice column and row in an image. A slice can be a rectangular region of a frame designed to be decoded and encoded independently, although loop filtering across slice edges can still be applied.
[0251] A coding tree unit (CTU) can be a coding tree block of luminance samples from an image with three sample arrays, or two corresponding coding tree blocks of chrominance samples. Alternatively, a CTU can be a coding tree block of samples from a monochrome image and an image encoded using three separate color planes and a syntax structure for encoding the samples. A superblock can be a 64×64 pixel square block consisting of one or two pattern information blocks, or recursively divided into four 32×32 blocks, which themselves can be further subdivided.
[0252] (System Configuration)
[0253] First, the transmission system according to an embodiment will be described. Figure 1 This is a schematic diagram illustrating an example configuration of a transmission system 400 according to an embodiment.
[0254] Transmission system 400 is a system for transmitting a stream generated by encoding an image and decoding the transmitted stream. As shown in the figure, transmission system 400 includes, for example, […]. Figure 1 The encoder 100, network 300, and decoder 200 are shown.
[0255] An image is input to encoder 100. Encoder 100 generates a stream by encoding the input image and outputs the stream to network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed through encoding.
[0256] It should be noted that the image before encoding by encoder 100 is also referred to as the raw image, raw signal, or raw sample. The image can be video or a still image. Image is a general concept encompassing sequences, pictures, and blocks, and therefore, unless otherwise specified, it is not limited to a spatial region of a specific size or a temporal region of a specific size. An image is an array of pixels or pixel values, and the signal representing the image or pixel values is also called a sample. This stream can be referred to as a bitstream, encoded bitstream, compressed bitstream, or encoded signal. Furthermore, encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.
[0257] Network 300 transmits the stream generated by encoder 100 to decoder 200. Network 200 can be any combination of the Internet, wide area network (WAN), local area network (LAN), or network. Network 300 is not limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasts and satellite broadcasts. Alternatively, network 300 can be replaced by a recording medium such as a digital multifunction disc (DVD) and Blu-ray disc (BD) on which the stream is recorded.
[0258] Decoder 200 generates a decoded image as an uncompressed image, for example, by decoding the stream transmitted by network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method adopted by encoder 100.
[0259] It should be noted that decoder 200 can also be called image decoder or video decoder, and the decoding method performed by decoder 200 can also be called decoding method, image decoding method or video decoding method.
[0260] (Data Structures)
[0261] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream. For convenience, reference will be made to... Figure 1 The transmission system 400 is used to describe Figure 2 Streams include, for example, video sequences. (e.g.) Figure 2 As shown in (a), the video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplementary enhancement information (SEI), and multiple pictures.
[0262] In a video with multiple layers, a VPS may include encoding parameters shared between some of the layers, as well as encoding parameters associated with some of the layers included in the video or with a single layer.
[0263] The SPS includes parameters for the sequence, that is, the encoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the image. It should be noted that multiple SPSs can exist.
[0264] PPS includes parameters for the images; that is, encoding parameters that the decoder 200 references in order to decode each image in the sequence. For example, encoding parameters may include a reference value for the quantization width used to decode the images and flags indicating the application of weighted prediction. It should be noted that multiple PPSs can exist. Each of the SPS and PPS can be simply referred to as a parameter set.
[0265] like Figure 2 As shown in (b), an image may include an image header and one or more slices. The image header includes encoding parameters referenced by the decoder 200 for decoding one or more slices.
[0266] like Figure 2 As shown in (c), a slice includes a slice header and one or more bricks. The slice header includes encoding parameters referenced by the decoder 200 for decoding one or more bricks.
[0267] like Figure 2 As shown in (d), a brick comprises one or more coding tree units (CTUs).
[0268] Note that an image may not include any slices and may include groups of slices instead of slices. In this case, a group of slices includes at least one slice. Furthermore, bricks may include slices.
[0269] CTU is also called a superblock or basis splitting unit. For example... Figure 2 As shown in (e), the CTU includes a CTU header and at least one coding unit (CU). As shown, the CTU includes four coding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes coding parameters referenced by the decoder 200 for decoding at least one CU.
[0270] A CU can be divided into multiple smaller CUs. As shown in the figure, CU(10) is not divided into smaller coding units; CU(11) is divided into four smaller coding units CU(110), CU(111), CU(112), and CU(113); CU(12) is not divided into smaller coding units; and CU(13) is divided into seven smaller coding units CU(1310), CU(1311), CU(1312), CU(1313), CU(132), CU(133), and CU(134). Figure 2 As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is used to predict the CU, and the residual coefficient information represents the prediction residuals, which will be described later. Although the CU is essentially the same as the prediction unit (PU) and transform unit (TU), it should be noted that, for example, a subblock transform (SBT), which will be described later, may include multiple TUs smaller than the CU. Furthermore, the CU can be processed for each virtual pipelined decoding unit (VPDU) included in the CU. The VPDU is, for example, a fixed unit that can be processed at one stage when pipelined processing is performed in the hardware.
[0271] It should be noted that a flow may not include... Figure 2 All layers are shown. The order of the layers can be interchanged, or any layer can be replaced by another layer. Here, the image that is the target of a process to be performed by a device such as encoder 100 or decoder 200 is called the current image. When the process is an encoding process, the current image represents the current image to be encoded, and when the process is a decoding process, the current image represents the current image to be decoded. Similarly, for example, a CU or CU block that is the target of a process to be performed by a device such as encoder 100 or decoder 200 is called the current block. When the process is an encoding process, the current block represents the current block to be encoded, and when the process is a decoding process, the current block represents the current block to be decoded.
[0272] (Image structure: slice / partition)
[0273] An image can be configured with one or more slice units or one or more fragment units to facilitate parallel encoding / decoding of the image.
[0274] A slice is a basic coding unit included in an image. An image may include, for example, one or more slices. Furthermore, a slice includes one or more coding tree units (CTUs).
[0275] Figure 3 This is a conceptual diagram used to illustrate an example of slice configuration. For example, in Figure 3The image contains 11×8 CTUs and is divided into four slices (slices 1 to 4). Slice 1 contains 16 CTUs, slice 2 contains 21 CTUs, slice 3 contains 29 CTUs, and slice 4 contains 22 CTUs. Each CTU in the image belongs to one of the slices. The shape of each slice is obtained by horizontally segmenting the image. The boundaries of each slice do not need to coincide with the edges of the image and can coincide with any boundaries between CTUs in the image. The processing order (encoding or decoding order) of the CTUs in the slice is, for example, the raster scan order. Each slice includes a slice header and encoded data. Slice characteristics can be written into the slice header. Characteristics may include the CTU address of the top CTU in the slice, the slice type, etc.
[0276] A tile is a unit of rectangular area included in an image. Image tiles can be assigned a number called TileId according to the raster scan order.
[0277] Figure 4 This is a conceptual diagram used to illustrate an example of a sharding configuration. For example, in Figure 4 The image contains 11×8 CTUs and is divided into four rectangular regions (partitions 1 to 4). When using partitioning, the processing order of the CTUs may differ from that without partitioning. Without partitioning, multiple CTUs in the image are typically processed in raster scan order. When using multiple partitions, at least one CTU in each of the multiple partitions is processed in raster scan order. For example, as... Figure 4 As shown, the processing order of the CTUs included in shard 1 is from the left end of the first column of shard 1 to the right end of the first column of shard 1, and then continues from the left end of the second column of shard 1 to the right end of the second column of shard 1.
[0278] It should be noted that a slice can include one or more slices, and a slice can include one or more slices.
[0279] It is important to note that an image can be configured with one or more fragment sets. A fragment set can include one or more fragment groups, or one or more fragments. An image can be configured with one of fragment sets, fragment groups, and fragments. For example, suppose the order in which multiple fragments are scanned for each fragment set in raster scan order is the basic encoding order of the fragments. Suppose the set of one or more fragments consecutive in basic encoding order within each fragment set is a fragment group. Such an image can be segmented by segmenter 102, described later (see [link to documentation]). Figure 7 Configure it using ).
[0280] (Scalable encoding)
[0281] Figure 5 and Figure 6This is a conceptual diagram illustrating an example of a scalable flow structure, and for convenience, references are provided. Figure 1 Describe it.
[0282] like Figure 5 As shown, encoder 100 can generate a temporally / spatially scalable stream by dividing each of multiple images into any of multiple layers and encoding the images within those layers. For example, encoder 100 encodes the images for each layer, thus achieving scalability even when enhancement layers exist on top of the base layer. This encoding of each image is also called scalable encoding. In this way, decoder 200 is able to switch the image quality of the images displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as the processing power of decoder 200 and external factors such as the state of communication bandwidth. As a result, decoder 200 is able to decode content while freely switching between low and high resolutions. For example, a user of the stream might watch halfway through a video stream on a smartphone on their way home and continue watching the video at home on a device (e.g., a TV connected to the Internet). It should be noted that each of the aforementioned smartphones and devices includes decoder 200 with the same or different performance. In this case, the user can watch high-quality video at home while the device decodes layers to higher layers in the stream. In this way, encoder 100 does not need to generate multiple streams with different image qualities having the same content, and thus the processing load can be reduced.
[0283] Furthermore, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 may generate a video whose image quality has been enhanced by performing super-resolution imaging on the images in the base layer based on the metadata. Super-resolution imaging may include, for example, an increase in the SN ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values in filtering processes, machine learning, or least squares methods (used in super-resolution processing).
[0284] In an embodiment, a configuration can be provided in which an image is divided into segments, for example, based on the meaning of objects within it. In this case, the decoder 200 can decode only a portion of the image by selecting the segments to be decoded. Furthermore, the attributes of objects (people, cars, balls, etc.) and their positions within the image (coordinates within the same image) can be stored as metadata. In this case, the decoder 200 can identify the location of the desired object based on the metadata and determine the segment that includes that object. For example, as... Figure 6As shown, metadata can be stored using a different data storage structure than image data, such as the SEI (Supplemental Enhancement Information) message in HEVC. This metadata indicates, for example, the location, size, or color of the main object.
[0285] Metadata can be stored in units of multiple images (e.g., streams, sequences, or random access units). In this way, decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information to the image unit information, it can identify the images in which the object (person) appears and determine the object's position in the image.
[0286] (encoder)
[0287] An encoder according to an embodiment will be described. Figure 7 This is a block diagram illustrating the functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes video in blocks.
[0288] like Figure 7 As shown, encoder 100 is a device for encoding images in blocks, and includes a segmenter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra-frame predictor 124 and the inter-frame predictor 126 are part of the prediction controller.
[0289] The encoder 100 is implemented, for example, as a general-purpose processor and memory. In this case, when the software program stored in memory is executed by the processor, the processor acts as a divider 102, a subtractor 104, a converter 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse converter 114, an adder 116, a loop filter 120, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the divider 102, subtractor 104, converter 106, quantizer 108, entropy encoder 110, inverse quantizer 112, inverse converter 114, adder 116, loop filter 120, intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128.
[0290] (Encoder installation example)
[0291] Figure 8 This is a functional block diagram illustrating an example installation of encoder 100. Encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The encoder 100 shown has multiple constituent elements mounted on it. Figure 8 The processor a1 and memory a2 are shown.
[0292] Processor a1 is a circuit that performs information processing and is coupled to memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes images. Processor a1 can be a processor such as a CPU. Furthermore, processor a1 can be an assembly of multiple electronic circuits. Additionally, for example, processor a1 can perform... Figure 7 The roles of two or more of the constituent elements in the encoder 100 shown.
[0293] Memory a2 is a dedicated or general-purpose memory used by processor a1 to encode images. Memory a2 can be an electronic circuit and can be connected to processor a1. Furthermore, memory a2 can be included within processor a1. Additionally, memory a2 can be an assembly of multiple electronic circuits. Furthermore, memory a2 can be a disk, optical disk, etc., or can be represented as a storage device, recording medium, etc. Furthermore, memory a2 can be non-volatile memory or volatile memory.
[0294] For example, memory a2 can store the image to be encoded or the bit stream corresponding to the encoded image. Furthermore, memory a2 can store a program for instructing processor a1 to encode the image.
[0295] Furthermore, for example, memory a2 can act as Figure 7 The encoder 100 shown illustrates the roles of two or more constituent elements among its multiple constituent elements for storing information. For example, memory a2 can act as... Figure 7 The roles of block memory 118 and frame memory 122 are shown. More specifically, memory a2 can store reconstructed blocks, reconstructed images, etc.
[0296] It should be noted that in encoder 100, implementation is not required. Figure 7 All of the multiple constituent elements shown are included, and it is not necessary to perform all the processes described here. Figure 7 A portion of the constituent elements shown may be included in another device, or a portion of the process described herein may be performed by another device.
[0297] The following describes the overall flow of the process performed by encoder 100, and then describes each constituent element included in encoder 100.
[0298] (The overall flow of the coding process)
[0299] Figure 9This is a flowchart illustrating an example of the overall encoding process performed by encoder 100, and for convenience, references... Figure 7 Describe it.
[0300] First, the segmenter 102 of the encoder 100 segments each image contained in the input image into multiple blocks of a fixed size (e.g., 128 × 128 pixels) (step Sa_1). The segmenter 102 then selects a segmentation pattern for the fixed-size blocks (also called block shapes) (step Sa_2). In other words, the segmenter 102 further segments the fixed-size blocks into multiple blocks forming the selected segmentation pattern. For each of the multiple blocks, the encoder 100 performs steps Sa_3 to Sa_9 for that block (i.e., the current block to be encoded).
[0301] The prediction controller 128 and the prediction actuator (which includes an intra-frame predictor 124 and an inter-frame predictor 126) generate a prediction image of the current block (step Sa-3). The prediction image may also be referred to as a prediction signal, a prediction block, or a prediction sample.
[0302] Next, subtractor 104 generates the difference between the current block and the predicted image as the prediction residual (step Sa_4). The prediction residual can also be called the prediction error.
[0303] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate multiple quantized coefficients (step Sa_5). The multiple quantized coefficients are sometimes referred to as a coefficient block.
[0304] Next, the entropy encoder 110 encodes (specifically, entropy coding) multiple quantized coefficients and prediction parameters related to the generation of the predicted image to generate a stream (step Sa_6). This stream may sometimes be referred to as an encoded bitstream or a compressed bitstream.
[0305] Next, the inverse quantizer 112 performs inverse quantization on the multiple quantized coefficients, and the inverse transformer 114 performs inverse transformation on the result to recover the prediction residual (step Sa_7).
[0306] Next, adder 116 adds the predicted image to the recovered prediction residual to reconstruct the current block (step Sa_8). This generates the reconstructed image. The reconstructed image can also be called a reconstructed block or a decoded image block.
[0307] When the reconstructed image is generated, the loop filter 120 performs filtering on the reconstructed image as needed (step Sa_9).
[0308] Encoder 100 then determines whether the encoding of the entire image has been completed (step Sa_10). If it is determined that the encoding has not been completed (no in step Sa_10), the processing starting from step Sa_2 is repeated for the next block of the image.
[0309] Although in the example above, encoder 100 selects a segmentation pattern for fixed-size blocks and encodes each block according to the segmentation pattern, it should be noted that each block can be encoded according to a corresponding segmentation pattern from multiple segmentation patterns. In this case, encoder 100 can evaluate the cost of each of the multiple segmentation patterns and, for example, select the stream that can be obtained by encoding according to the segmentation pattern that produces the minimum cost as the output stream.
[0310] As shown in the figure, the processes in steps Sa_1 to Sa_10 are executed sequentially by encoder 100. Alternatively, two or more processes can be executed in parallel, processes can be reordered, and so on.
[0311] The encoder 100 employs a hybrid encoding process using predictive coding and transform coding. Furthermore, the predictive coding is performed by an encoding loop configured with a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, a block memory 118, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. In other words, the prediction executor configured with the intra-frame predictor 124 and the inter-frame predictor 126 is part of the encoding loop.
[0312] (Divider)
[0313] Segmenter 102 divides each image contained in the original image into multiple blocks and outputs each block to subtractor 104. For example, segmenter 102 first segments the image into fixed-size blocks (e.g., 128×128 pixels). Other fixed block sizes may be used. Fixed-size blocks are also called coding tree units (CTUs). Segmenter 102 then segments each fixed-size block into variable-size blocks (e.g., 64×64 pixels or smaller) based on recursive quadtree and / or binary tree block segmentation. In other words, segmenter 102 selects a segmentation mode. Variable-size blocks may also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in the image can be processed in units of CUs, PUs, or TUs.
[0314] Figure 10 This is a conceptual diagram used to illustrate an example of block segmentation according to an embodiment. Figure 10In the diagram, solid lines represent the block boundaries of blocks partitioned by quadtree block partitioning, while dashed lines represent the block boundaries of blocks partitioned by binary tree block partitioning.
[0315] Here, block 10 is a square block with 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into four square blocks of 64×64 pixels (quadtree block partitioning).
[0316] The 64×64 pixel block in the upper left corner is further vertically divided into two rectangular 32×64 pixel blocks, and the 32×64 pixel block on the left is further vertically divided into two rectangular 16×64 pixel blocks (binary tree block partitioning). As a result, the 64×64 pixel block in the upper left corner is divided into two 16×64 pixel blocks 11 and 12 and one 32×64 pixel block 13.
[0317] The 64×64 pixel block in the upper right corner is horizontally divided into two rectangular 64×32 pixel blocks, 14 and 15 (binary tree block division).
[0318] The 64×64 pixel block in the bottom left corner is first divided into four 32×32 pixel square blocks (quadtree block partitioning). The top left and bottom right blocks of these four 32×32 pixel square blocks are further partitioned. The top left 32×32 pixel square block is vertically divided into two 16×32 pixel rectangular blocks, and the right 16×32 pixel block is further horizontally divided into two 16×16 pixel blocks (binary tree block partitioning). The bottom right 32×32 pixel block is horizontally divided into two 32×16 pixel blocks (binary tree block partitioning). The top right 32×32 pixel square block is horizontally divided into two 32×16 pixel rectangular blocks (binary tree block partitioning). As a result, the 64×64 pixel square in the lower left corner was divided into a 16×32 pixel rectangular block 16, two 16×16 pixel square blocks 17 and 18, two 32×32 pixel square blocks 19 and 20, and two 32×16 pixel rectangular blocks 21 and 22.
[0319] The 64×64 pixel block 23 in the bottom right corner was not segmented.
[0320] As mentioned above, in Figure 10 In this example, based on recursive quadtree and binary tree block partitioning, block 10 is divided into 13 variable-size blocks 11 to 23. This type of partitioning is also known as quadtree plus binary tree (QTBT) partitioning.
[0321] It is important to note that, in Figure 10 In this context, a block is divided into four or two blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these examples. For instance, a block can be partitioned into three blocks (ternary block partitioning). Partitioning that includes this type of ternary block partitioning is also known as multi-type tree (MBT) partitioning.
[0322] Figure 11 This is a block diagram illustrating an example of the functional configuration of a segmenter 102 according to one embodiment. Figure 11 As shown, the segmenter 102 may include a block segmentation determiner 102a. As an example, the block segmentation determiner 102a may perform the following process.
[0323] For example, block segmentation determiner 102a can obtain or retrieve block information from block memory 118 and / or frame memory 122, and determine a segmentation mode (e.g., the segmentation mode described above) based on the block information. Segmenter 102 segments the original image according to the segmentation mode and outputs at least one block obtained by segmentation to subtractor 104.
[0324] Furthermore, for example, the block segmentation determiner 102a outputs one or more parameters indicating the determined segmentation pattern (e.g., the segmentation pattern described above) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 can transform the prediction residual based on one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 can generate a prediction image based on one or more parameters. Furthermore, the entropy encoder 110 can entropy encode one or more parameters.
[0325] Parameters related to the segmentation mode can be written in a stream, as shown in the following example.
[0326] Figure 12 This is a conceptual diagram used to illustrate examples of segmentation patterns. Examples of segmentation patterns include: segmentation into four regions (QT), where one block is divided into two regions in both the horizontal and vertical directions; segmentation into three regions (HT or VT), where one block is divided in the same direction at a 1:2:1 ratio; segmentation into two regions (HB or VB), where one block is divided in the same direction at a 1:1 ratio; and no segmentation (NS).
[0327] It should be noted that the segmentation mode does not have block segmentation direction when it is divided into four regions or not segmented, and the segmentation mode has segmentation direction information when it is divided into two regions or three regions.
[0328] Figure 13A This is a conceptual diagram used to illustrate an example of a syntactic tree for segmentation patterns.
[0329] Figure 13B This is a conceptual diagram used to illustrate another example of a syntactic tree for segmentation patterns.
[0330] Figure 13A and Figure 13BThis is a conceptual diagram used to illustrate an example of a syntactic tree for segmentation patterns. Figure 13A In the example, first, there is information indicating whether to perform a split (S: split flag), and then there is information indicating whether to split into 4 regions (QT: QT flag). Next, there is information indicating which of the three or two regions to split into (TT: TT flag or BT: BT flag), and then there is information indicating the direction of the split (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of the at least one block obtained by splitting according to this splitting pattern can be further split in a similar process. In other words, as an example, whether to perform a split, whether to split into four regions, which of the horizontal and vertical directions is the direction to perform the splitting method, whether to split into three regions, and which of the two regions to split into can be recursively determined and can be based on... Figure 13A The encoding order revealed by the syntax tree shown will determine the encoding of the result in the stream.
[0331] Additionally, although the information items indicating S, QT, TT, and Ver are respectively in Figure 13A The syntax tree shown is arranged in the listed order, indicating that the information items for S, QT, Ver, and BT can also be arranged in the listed order. In other words, in Figure 13B In the example, first, there is information indicating whether to perform a split (S: split flag), and then there is information indicating whether to split into 4 regions (QT: QT flag). Next, there is information indicating the split direction (Ver: vertical flag, or Hor: horizontal flag), and then there is information indicating which of the three regions to split into (BT: BT flag or TT: TT flag).
[0332] It should be noted that the segmentation patterns described above are examples, and segmentation patterns other than those described can be used, or a portion of the described segmentation patterns can be used.
[0333] (Subtractor)
[0334] Subtractor 104 subtracts the predicted image (predicted samples input from the prediction controller 128 indicated below) from the original image, which is input from and segmented by segmenter 102, on a block-by-block basis. In other words, subtractor 104 calculates the prediction residual (also referred to as error) for the current block. Subtractor 104 then outputs the calculated prediction residual to transformer 106.
[0335] The original image can be an image that has been input to encoder 100 as a signal (e.g., luminance signal and two chrominance signals) representing each picture included in the video. The signal representing the image can also be called a sample.
[0336] (Transformer)
[0337] Transformer 106 transforms the prediction residual in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to quantizer 108. More specifically, transformer 106 applies, for example, a defined Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) to the prediction residual in the spatial domain. The defined DCT or DST can be predefined.
[0338] It should be noted that the transformer 106 can adaptively select a transformation type from multiple transformation types and transform the prediction residuals into transformation coefficients by using transformation basis functions corresponding to the selected transformation type. This transformation is also called explicit multi-kernel transformation (EMT) or adaptive multi-kernel transformation (AMT). The transformation basis functions can also be called bases.
[0339] Transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transformation types can also be represented as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 This is a graph indicating the example transformation basis functions of the example transformation type. Figure 14 In this context, N represents the number of input pixels. For example, the choice of a transform type from multiple transform types can depend on the prediction type (one of intra-frame prediction and inter-frame prediction) and can also depend on the intra-frame prediction mode.
[0340] Information indicating whether to apply this EMT or AMT (e.g., referred to as EMT flags or AMT flags) and information indicating the selected transformation type are typically signaled at the CU level. It should be noted that signaling for this type of information does not necessarily need to be performed at the CU level and can also be performed at another level (e.g., sequence level, image level, slice level, fragment level, or CTU level).
[0341] Furthermore, transformer 106 can re-transform the transform coefficients (which are the transform results). This re-transformation is also known as adaptive quadratic transform (AST) or non-separable quadratic transform (NSST). For example, transformer 106 performs the re-transformation on a sub-block (e.g., a 4×4 pixel sub-block) basis, which is included in the transform coefficient block corresponding to the intra-frame prediction residual. Information indicating whether NSST is applied, as well as information related to the transform matrix used in NSST, is typically signaled at the CU level. It should be noted that such signaling does not necessarily need to be performed at the CU level and can also be performed at another level (e.g., sequence level, picture level, slice level, fragment level, or CTU level).
[0342] Transformer 106 can employ both separable and non-separable transformations. A separable transformation is a method in which the transformation is performed multiple times by individually applying the transformation to each of the multiple directions according to the dimension of the input. A non-separable transformation is a method of performing a collective transformation, in which two or more dimensions of the multidimensional input are collectively treated as a single dimension.
[0343] In one example of an inseparable transformation, when the input is a 4×4 pixel block, the 4×4 pixel block is considered as a single array containing 16 elements, and the transformation applies a 16×16 transformation matrix to that array.
[0344] In another example of an inseparable transformation, the 4×4 pixel input block is treated as a single array containing 16 elements, and a transformation (hypercube given transformation) can then be performed on that array with multiple given rotations.
[0345] In the transformation within transformer 106, the transformation type of the transform basis functions to be transformed into the frequency domain can be switched according to the region in the CU. Examples include spatially varying transform (SVT).
[0346] Figure 15 This is a conceptual diagram used to illustrate an example of SVT.
[0347] In SVT, such as Figure 15As shown, the CU is divided into two equal regions horizontally or vertically, and only one of these regions is transformed to the frequency domain. The transform base type can be set for each region. For example, DST7 and DST8 can be used. For instance, in the two regions obtained by vertically dividing the CU into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, in the two regions, DST7 can be used for the region at position 1. Similarly, in the two regions obtained by horizontally dividing the CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Although in Figure 15 In the example shown, one of the two regions in the CU is transformed while the other is not, but each of the two regions can be transformed. Furthermore, the segmentation method can include not only segmentation into two regions but also segmentation into four regions. Moreover, the segmentation method can be more flexible. For example, information indicating the segmentation method can be encoded and signaled in the same way as CU segmentation. Note that SVT can also be called Subblock Transform (SBT).
[0348] The AMT and EMT described above can be referred to as MTS (Multiple Transform Selection). When applying MTS, transform types such as DST7 and DCT8 can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. There is another process called IMTS (Implicit MTS) for selecting the transform type to be used for orthogonal transforms performed without encoded index information. When applying IMTS, for example, when the CU has a rectangular shape, the orthogonal transform for the rectangular shape can be performed using DST7 (for the short side) and DST2 (for the long side). Alternatively, for example, when the CU has a square shape, the orthogonal transform for the rectangular shape can be performed by using DCT2 when MTS is valid in the sequence and using DST7 when MTS is invalid in the sequence. DCT2 and DST7 are just examples. Other transform types can be used, and the combination of transform types used can also be changed to different combinations of transform types. IMTS can be used only for intra-prediction blocks, or it can be used for both intra-prediction blocks and inter-prediction blocks.
[0349] The three processes MTS, SBT, and IMTS have been described above as selection processes for selectively switching the transformation type used for orthogonal transformation. However, all three selection processes can be used, or only some selection processes can be selectively used. For example, it can be determined whether to use one or more selection processes based on flag information in a header such as SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transformation of the CU is performed. It should be noted that the selection process for selectively switching the transformation type can be a different selection process from the three selection processes mentioned above, or each of the three selection processes can be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transformation of the entire CU and encoding information indicating the transformation type used in the transformation. Function [2] is a function for performing the orthogonal transformation of the entire CU and determining the transformation type based on a defined rule without encoding the information indicating the transformation type. Function [3] is a function for performing the orthogonal transformation of a portion of the CU and encoding information indicating the transformation type used in the transformation. The function [4] is used to perform orthogonal transformations on a portion of the CU and determine the transformation type based on predetermined rules without encoding information indicating the type of transformation used in the transformation. The predetermined rules can be pre-determined.
[0350] It is important to note that the application of MTS, IMTS, and / or SBT can be determined for each processing unit. For example, the application of MTS, IMTS, and / or SBT can be determined for each sequence, image, brick, slice, CTU, or CU.
[0351] It should be noted that the tool for selectively switching transformation types in this invention can be described as a method, selection process, or procedure for selectively selecting the basis used in the transformation process. Furthermore, the tool for selectively switching transformation types can be described as a mode for adaptively selecting transformation types.
[0352] Figure 16 This is a flowchart illustrating an example of the process performed by converter 106, and for convenience, references will be made... Figure 7 Describe it.
[0353] For example, transformer 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined that an orthogonal transformation should be performed (Yes in step St_1), transformer 106 selects a transformation type for orthogonal transformation from multiple transformation types (step St_2). Next, transformer 106 performs orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Transformer 106 then outputs information indicating the selected transformation type to entropy encoder 110 to allow entropy encoder 110 to encode this information (step St_4). On the other hand, when it is determined that an orthogonal transformation should not be performed (No in step St_1), transformer 106 outputs information indicating that an orthogonal transformation should not be performed to allow entropy encoder 110 to encode this information (step St_5). It should be noted that whether to perform an orthogonal transformation in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, a defined transformation type can be used to perform the orthogonal transformation without encoding the information indicating the transformation type used in the orthogonal transformation. The defined transformation type can be predefined.
[0354] Figure 17 This is a flowchart illustrating an example of the process performed by converter 106, and for convenience, references will be made... Figure 7 Please describe it. It should be noted that... Figure 17 The example shown is in the case where the type of transformation used in orthogonal transformations is selectively switched (such as in...). Figure 16 Examples of orthogonal transformations (as shown in the example).
[0355] As an example, the first transform type group may include DCT2, DST7, and DCT8. As another example, the second transform type group may include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group may partially overlap with each other or may be completely different from each other.
[0356] Transformer 106 determines whether the transform size is less than or equal to a predetermined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the predetermined value (Yes in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform types included in the first transform type group (step Su_2). Next, transformer 106 outputs information to entropy encoder 110 indicating the transform type to be used from at least one transform type included in the first transform type group, so that entropy encoder 110 can encode this information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to a predetermined value (No in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The predetermined value can be a threshold and can be a predetermined value.
[0357] In step Su_3, the information indicating the transformation type used in the orthogonal transformation can be a combination of information indicating the transformation type to be applied vertically in the current block and the transformation type to be applied horizontally in the current block. The first type group may include only one transformation type and may not encode the information indicating the transformation type used for the orthogonal transformation. The second transformation type group may include multiple transformation types and may encode the information indicating the transformation type used for the orthogonal transformation that is included among one or more transformation types in the second transformation type group.
[0358] Alternatively, the transformation type can be indicated based on the transformation size without encoding the information indicating the transformation type. It should be noted that such determination is not limited to determining whether the transformation size is less than or equal to a determined value, and other procedures are possible for determining the transformation type used in orthogonal transformations based on the transformation size.
[0359] (Quantizer)
[0360] Quantizer 108 quantizes the transform coefficients output from converter 106. More specifically, quantizer 108 scans the transform coefficients of the current block in a defined scan order and quantizes the scanned transform coefficients based on the quantization parameters (QP) corresponding to the transform coefficients. Quantizer 108 then outputs the quantized transform coefficients of the current block (hereinafter also referred to as quantized coefficients) to entropy encoder 110 and inverse quantizer 112. The defined scan order can be predetermined.
[0361] A defined scan order is the order in which the quantization / inverse quantization transform coefficients are applied. For example, a defined scan order can be defined as ascending frequency (from low to high frequency) or descending frequency (from high to low frequency).
[0362] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, when the value of the quantization parameter increases, the quantization step size also increases. In other words, when the value of the quantization parameter increases, the error of the quantized coefficients (quantization error) increases.
[0363] Furthermore, a quantization matrix can be used for quantization. For example, various quantization matrices can be used corresponding to the frequency transform size (e.g., 4×4, 8×8), the prediction mode (e.g., intra-frame prediction, inter-frame prediction), and the pixel components (e.g., luma, chroma pixel components). It should be noted that quantization means digitizing values sampled at determined intervals corresponding to a defined level. In this art, quantization can be referred to using other expressions, such as rounding and scaling, and rounding and scaling can be employed. The determined intervals and determined levels can be predetermined.
[0364] Methods for using a quantization matrix can include: using a quantization matrix directly set on the encoder 100 side, and using a quantization matrix that has been set as a default (default matrix). On the encoder 100 side, a quantization matrix suitable for image features can be set by directly setting the quantization matrix. However, this may have the disadvantage of increasing the amount of coding required to encode the quantization matrix. It should be noted that a quantization matrix for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.
[0365] There exists a method for quantizing high-frequency and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be considered equivalent to using a quantization matrix (a flat matrix) whose coefficients have the same values.
[0366] The quantization matrix can be encoded at, for example, the sequence level, image level, slice level, brick level, or CTU level. The quantization matrix can be specified using, for example, a Sequence Parameter Set (SPS) or a Picture Parameter Set (PPS). The SPS includes parameters for the sequence, and the PPS includes parameters for the image. Each of the SPS and PPS can be simply referred to as a parameter set.
[0367] When using a quantization matrix, quantizer 108 uses the values of the quantization matrix to scale the quantization width for each transform coefficient, for example, based on quantization parameters. A quantization process performed without using a quantization matrix can be a process for quantizing the transform coefficients based on a quantization width calculated based on quantization parameters. It should be noted that in a quantization process performed without using any quantization matrix, the quantization width can be multiplied by a predetermined value common to all transform coefficients in the block. This predetermined value can be pre-determined.
[0368] Figure 18This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment. For example, quantizer 108 includes a differential quantization parameter generator 108a, a predictive quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization actuator 108e.
[0369] Figure 19 This is a flowchart illustrating an example of the quantization process performed by quantizer 108, and for convenience, references... Figure 7 and 18 Describe it.
[0370] As an example, the quantizer 108 can be based on Figure 19 The flowchart shown performs quantization for each CU. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (yes in step Sv_1), the quantization parameter generator 108c generates the quantization parameters for the current block (step Sv_2) and stores the quantization parameters in the quantization parameter storage device 108d (step Sv_3).
[0371] Next, quantization executor 108e uses the quantization parameters generated in step Sv_2 to quantize the transform coefficients of the current block (step Sv_4). Predictive quantization parameter generator 108b then obtains the quantization parameters of the processing unit different from the current block from the quantization parameter storage device 108d (step Sv_5). Predictive quantization parameter generator 108b generates predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). Differential quantization parameter generator 108a calculates the difference between the quantization parameters of the current block generated by quantization parameter generator 108c and the predicted quantization parameters of the current block generated by predictive quantization parameter generator 108b (step Sv_7). Differential quantization parameters can be generated by calculating the difference. Differential quantization parameter generator 108a outputs the differential quantization parameters to entropy encoder 110 to allow entropy encoder 110 to encode the differential quantization parameters (step Sv_8).
[0372] It should be noted that differential quantization parameters can be encoded at, for example, the sequence level, image level, slice level, brick level, or CTU level. Furthermore, the initial values of the quantization parameters can be encoded at the sequence level, image level, slice level, brick level, or CTU level. During initialization, the initial values of the quantization parameters and the differential quantization parameters can be used to generate the quantization parameters.
[0373] It should be noted that quantizer 108 may include multiple quantizers and may apply dependent quantization, in which the transformation coefficients are quantized using a quantization method selected from multiple quantization methods.
[0374] (Entropy encoder)
[0375] Figure 20 This is a block diagram illustrating an example of the functional configuration of the entropy encoder 110 according to an embodiment, and for convenience, reference will be made to... Figure 7 The entropy encoder 110 generates a stream by entropy encoding quantized coefficients input from quantizer 108 and prediction parameters input from prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy encoding. More specifically, the entropy encoder 110, as shown, includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, where a multi-level signal, such as quantized coefficients and prediction parameters, is transformed into a binary signal. Examples of binarization methods include truncated Ricean binarization, exponential Golomb codes, and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntax elements or the surrounding state (i.e., the probability of occurrence of the binary signal). Examples of methods for deriving the context value include bypassing, referencing syntax elements, referencing upper and left adjacent blocks, referencing hierarchical information, etc. The binary arithmetic encoder 110c uses the derived context to perform arithmetic encoding on the binary signal.
[0376] Figure 21 This is a conceptual diagram illustrating an example flow of the CABAC process in entropy encoder 110. First, initialization is performed in CABAC within entropy encoder 110. During initialization, initialization and setting of the initial context value are performed in binary arithmetic encoder 110c. For example, binarizer 110a and binary arithmetic encoder 110c can sequentially perform binarization and arithmetic encoding of multiple quantization coefficients within the CTU. Each time arithmetic encoding is performed, context controller 110b can update the context value. Context controller 110b can then save the context value for post-processing. For example, the saved context value can be used to initialize the context value for the next CTU.
[0377] (Inverse quantizer)
[0378] Inverse quantizer 112 inverse quantizes the quantized coefficients input from quantizer 108. More specifically, inverse quantizer 112 inverse quantizes the quantized coefficients of the current block in a defined scan order. Inverse quantizer 112 then outputs the inverse quantized transform coefficients of the current block to inverse transformer 114. The defined scan order can be predetermined.
[0379] (Inverse Transformer)
[0380] Inverse transformer 114 recovers the prediction residual by performing an inverse transform on the transform coefficients input from inverse quantizer 112. More specifically, inverse transformer 114 recovers the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by transformer 106. Inverse transformer 114 then outputs the recovered prediction residual to adder 116.
[0381] It should be noted that, because information is typically lost during quantization, the recovered prediction residual does not match the prediction residual calculated by subtractor 104. In other words, the recovered prediction residual usually includes quantization error.
[0382] (Adder)
[0383] Adder 116 reconstructs the current block by adding the prediction residual input from inverse transformer 114 and the prediction image input from prediction controller 128. The reconstructed image is then generated. Adder 116 then outputs the reconstructed image to block memory 118 and loop filter 120. The reconstructed block can also be referred to as a local decoded block.
[0384] (Block memory)
[0385] Block memory 118 is a storage device for storing, for example, blocks in the current image used for intra-frame prediction. More specifically, block memory 118 stores the reconstructed image output from adder 116.
[0386] (Frame Memory)
[0387] Frame memory 122 is a storage device, for example, used to store reference images used in inter-frame prediction, and is also referred to as a frame buffer. More specifically, frame memory 122 stores reconstructed images filtered by loop filter 120.
[0388] (Loop filter)
[0389] Loop filter 120 applies a loop filter to the reconstructed image output by adder 116 and outputs the filtered reconstructed image to frame memory 122. A loop filter is a filter used in the coding loop (in-loop filter). Examples of loop filters include, for example, adaptive loop filter (ALF), deblocking filter (DB or DBF), sample adaptive offset (SAO) filter, etc.
[0390] Figure 22 This is a block diagram illustrating an example of the functional configuration of a loop filter 120 according to an embodiment. For example, as Figure 22As shown, the loop filter 120 includes a deblocking filter executor 120a, a SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs a deblocking filter process on the reconstructed image. The SAO executor 120b performs a SAO process on the reconstructed image after the deblocking filter process. The ALF executor 120c performs an ALF process on the reconstructed image after the SAO process. The ALF and deblocking filter will be described in detail later. The SAO process is used to improve image quality by reducing ringing (the phenomenon where pixel values are distorted like waves around edges) and correcting pixel value deviations. Examples of SAO processes include edge offset processes and band offset processes. It should be noted that in some embodiments, the loop filter 120 may not include... Figure 22 The loop filter 120 can be configured to interact with all the constituent elements disclosed herein, and may include some constituent elements and additional elements. Furthermore, the loop filter 120 can be configured to interact with... Figure 22 The above process may be executed in different processing orders as disclosed in the documentation, and it is possible that not all processes will be executed, etc.
[0391] (Loop filter > Adaptive loop filter)
[0392] In ALF, a least-squares error filter is applied to remove compression artifacts. For example, a filter selected from multiple filters is applied for each 2×2 pixel sub-block in the current block, based on the direction and activity of the local gradient.
[0393] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of several classes (e.g., fifteen or twenty-five classes). The classification of sub-blocks can be based on, for example, gradient directionality and activity. In the example, a category index C (e.g., C = 5D + A) is calculated or determined based on the gradient directionality D (e.g., 0 to 2 or 0 to 4) and the gradient activity A (e.g., 0 to 4). Then, based on the category index C, each sub-block is classified into one of several categories.
[0394] For example, gradient directionality D is calculated by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, gradient activity A is calculated, for example, by summing the gradients in multiple directions and quantizing the sum.
[0395] Based on such classification results, the filter to be used for each sub-block can be determined from multiple filters.
[0396] The filter shape to be used in ALF is, for example, a circularly symmetric filter shape. Figures 23A to 23C This is a conceptual diagram used to illustrate an example of the filter shape used in ALF. Figure 23A The diagram illustrates a 5×5 diamond filter. Figure 23B The diagram illustrates a 7×7 diamond filter, and Figure 23C The diagram illustrates a 9×9 diamond-shaped filter. Information indicating the filter shape is typically signaled at the picture level. It should be noted that this signaling indicating the filter shape does not necessarily need to be performed at the picture level and can be performed at another level (e.g., at the sequence level, slice level, fragment level, CTU level, or CU level).
[0397] For example, the activation or deactivation of ALF can be determined at the image level or the CU level. For instance, a decision on whether to apply ALF to luminance can be made at the CU level, and a decision on whether to apply ALF to chrominance can be made at the image level. The information indicating whether ALF is activated or deactivated is typically signaled at the image or CU level. It should be noted that the signaling indicating whether ALF is activated or deactivated does not necessarily need to be executed at the image or CU level, and can be executed at another level (e.g., at the sequence level, slice level, partition level, or CTU level).
[0398] Furthermore, as described above, a filter is selected from multiple filters, and the ALF procedure for the sub-block is executed. The coefficient set for each of the multiple filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that the signaling of the coefficient set does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, fragment level, CTU level, CU level, or sub-block level).
[0399] (Loop filter > Cross-component adaptive loop filter)
[0400] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF). Figure 23E This is a conceptual diagram used to illustrate an example of the filter shape used in CC-ALF, for example... Figure 23D CC-ALF. Figure 23D and Figure 23E The example CC-ALF operates by applying a linear diamond filter to the luminance channel of each chromaticity component. For example, the filter coefficients can be transmitted in the APS, scaled by a factor of 2^10, and rounded against the fixed-point representation. For example, in... Figure 23D In this context, the Y sample (first component) is used for CCALF for Cb and CCALF for Cr (a component different from the first component).
[0401] The application of filters can be controlled on a variable block size and signaled via context-coded flags received for each sample block. The block size, along with the CC-ALF enable flags, can be received at the slice level for each chroma component. CC-ALF can support various block sizes, such as (in chroma samples) 16×16 pixels, 32×32 pixels, 64×64 pixels, and 128×128 pixels.
[0402] (Loop filter > Joint chroma cross-component adaptive loop filter)
[0403] An example of Union Chromaticity-CCALF is in Figure 23F and Figure 23G As shown in the image. Figure 23F This is a conceptual diagram used to illustrate an example process for joint chromaticity CCALF. Figure 23G This is a table showing example weight index candidates. As shown, a CCALF filter is used to generate a CCALF filtered output as a chroma refinement signal for one color component, while a weighted version of the same chroma refinement signal is applied to another color component. This reduces the complexity of the existing CCALF by approximately half. Weight values can be encoded as a symbol flag and a weight index. The weight index (denoted as weight_index) can be encoded as 3 bits and specifies the size of the JC-CCALF weight JcCcWeight, which is non-zero. For example, the size of JcCcWeight can be determined as follows:
[0404] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index >> 2;
[0405] Otherwise, JcCcWeight equals 4 / (weight_index–4).
[0406] The block-level on / off control for Cb and Cr ALF filters can be separate. This is the same as in CCALF, and the block-level on / off control flags can be encoded for two separate groups. Unlike CCALF, the Cb and Cr on / off control block sizes are the same here, so only one block size variable needs to be encoded.
[0407] (Loop filter > Deblocking filter)
[0408] During the deblocking filtering process, the loop filter 120 performs a filtering process on the block boundaries in the reconstructed image in order to reduce the distortion that occurs at the block boundaries.
[0409] Figure 24 This shows the loop filter 120 acting as a deblocking filter (see...). Figure 7 and Figure 22A block diagram of an example configuration of the deblocking filter actuator 120a.
[0410] The deblocking filter actuator 120a includes: a boundary determiner 1201; a filter determiner 1203; a filter actuator 1205; a process determiner 1208; a filter characteristic determiner 1207; and switches 1202, 1204, and 1206.
[0411] Boundary determiner 1201 determines whether the pixel to be deblocked (i.e., the current pixel) exists around the block boundary. Boundary determiner 1201 then outputs the determination result to switch 1202 and processing determiner 1208.
[0412] If boundary determiner 1201 determines that the current pixel exists around the block boundary, switch 1202 outputs the unfiltered image to switch 1204. Conversely, if boundary determiner 1201 determines that the current pixel does not exist around the block boundary, switch 1202 outputs the unfiltered image to switch 1206. Note that the unfiltered image is an image containing the current pixel and at least one surrounding pixel located around it.
[0413] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0414] If the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, switch 1204 outputs the unfiltered image obtained through switch 1202 to filter executor 1205. Conversely, if the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel, switch 1204 outputs the unfiltered image obtained through switch 1202 to switch 1206.
[0415] When an unfiltered image is obtained via switches 1202 and 1204, filter actuator 1205 performs deblocking filtering on the current pixel, with filtering characteristics determined by filter characteristic determiner 1207. Filter actuator 1205 then outputs the filtered pixel to switch 1206.
[0416] Under the control of the processing determinant 1208, the switch 1206 selectively outputs one of the pixels that have not yet been deblocked and filtered and the pixels that have been deblocked and filtered by the filter actuator 1205.
[0417] The processing determiner 1208 controls the switch 1206 based on the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary and the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the processing determiner 1208 causes the switch 1207 to output the pixel that has undergone deblocking filtering. Furthermore, in addition to the above cases, the processing determiner 1208 causes the switch 1206 to output pixels that have not undergone deblocking filtering. By repeating the pixel output in this way, a filtered image is output from the switch 1206. It should be noted that... Figure 24 The configuration shown is an example of a configuration in the deblocking filter executor 120a. The deblocking filter executor 120a can have various configurations.
[0418] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics relative to the block boundary.
[0419] During deblocking filtering, pixel values and quantization parameters can be used to select one of two deblocking filters with different characteristics (i.e., a strong filter and a weak filter). In the case of a strong filter, when pixels p0 to p2 and q0 to q2 cross block boundaries, such as... Figure 25 As shown, by performing a calculation, for example, according to the following expression, the pixel values of the corresponding pixels q0 to q2 are changed to pixel values q'0 to q'2.
[0420] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0421] q'1=(p0+q0+q1+q2+2) / 4
[0422] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0423] It should be noted that in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the corresponding pixels p0 to p2 and q0 to q2. Additionally, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of pixel q2 relative to the block boundary. Furthermore, the coefficients multiplied on the right-hand side of each expression by the corresponding pixel values of the pixels to be used for deblocking filtering are the filter coefficients.
[0424] Furthermore, in deblocking filtering, clipping can be performed to ensure that the variation in calculated pixel values does not exceed a threshold. For example, during clipping, the pixel values calculated according to the above expression can be clipped to values obtained according to "calculate pixel value ± 2 × threshold" (using a threshold determined based on quantization parameters). This prevents over-smoothing.
[0425] Figure 26 It is a conceptual diagram used to illustrate the block boundaries for which the deblocking filtering process is performed. Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.
[0426] The block boundaries for which the deblocking filtering process is performed are, for example, the boundaries between CUs, Pus, or TUs with 8×8 pixel blocks, such as... Figure 26 As shown. The deblocking filtering process can be performed, for example, in units of four rows or four columns. First, as... Figure 27 The diagram shows the relationship between block P and block Q ( Figure 26 The boundary strength (Bs) value is determined as shown.
[0427] according to Figure 27 The Bs value determines whether to perform deblocking filtering on block boundaries belonging to the same image using different intensities. When the Bs value is 2, deblocking filtering is performed on the chroma signal. When the Bs value is 1 or greater and certain conditions are met, deblocking filtering is performed on the luma signal. These conditions can be predetermined. Note that the conditions used to determine the Bs value are not limited to... Figure 27 The ones shown, and the Bs value can be determined based on another parameter.
[0428] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))
[0429] Figure 28 This is a flowchart illustrating an example of the process performed by the predictor of encoder 100. Note that the predictor includes all or part of the following constituent elements: intra-frame predictor 124; inter-frame predictor 126; and prediction controller 128. The prediction executor includes, for example, intra-frame predictor 124 and inter-frame predictor 126.
[0430] The predictor generates a predicted image for the current block (step Sb_1). This predicted image can also be referred to as a predicted signal or a predicted block. Note that the predicted signal is, for example, an intra-frame predicted image (image prediction signal) or an inter-frame predicted image (inter-frame prediction signal). The predictor generates the predicted image for the current block using a reconstructed image obtained from another block through the generation of the predicted image, the generation of the prediction residual, the generation of the quantized coefficients, the recovery of the prediction residual, and the addition with the predicted image.
[0431] The reconstructed image can be, for example, an image in a reference image, or an image of a coded block in the current image (i.e., the other blocks mentioned above), where the current image includes the current block. A coded block in the current image can be, for example, a neighboring block of the current block.
[0432] Figure 29 This is a flowchart illustrating another example of the process performed by the predictor of encoder 100.
[0433] The predictor generates a predicted image using a first method (step Sc_1a), a second method (step Sc_1b), and a third method (step Sc_1c). The first, second, and third methods can be different from each other for generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0434] Next, the prediction processor evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates the cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the predicted images by comparing the cost C of the predicted images. It should be noted that the cost C can be calculated, for example, according to an expression of the RD optimization model, such as C = D + λ × R. In this expression, D represents the compression artifacts of the predicted image and is expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. Additionally, R represents the bit rate of the stream. Furthermore, λ represents, for example, the multiplier according to the Lagrange multiplier method.
[0435] Then, the predictor selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects the method or mode used to obtain the final predicted image. For example, the predictor selects the predicted image with the minimum cost C based on the cost C calculated for the predicted image. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 can be based on parameters used during the encoding process. The encoder 100 can transform information used to identify the selected predicted image, method, or mode into a stream. This information can be, for example, flags. In this way, the decoder 200 is able to generate a predicted image based on this information according to the method or mode selected by the encoder 100. It should be noted that in Figure 29 In the example shown, after generating the predicted image using the appropriate method, the predictor selects any predicted image. However, the predictor can select a method or mode based on the parameters used in the encoding process described above before generating the predicted image, and can generate the predicted image according to the selected method or mode.
[0436] For example, the first method and the second method can be intra-frame prediction and inter-frame prediction, respectively, and the predictor can select the final predicted image of the current block from the predicted images generated according to the prediction method.
[0437] Figure 30 This is a flowchart illustrating another example of the process performed by the predictor of encoder 100.
[0438] First, the predictor generates a prediction image using intra-frame prediction (step Sd_1a) and an inter-frame prediction (step Sd_1b). It should be noted that the prediction image generated by intra-frame prediction is also called the intra-frame prediction image, and the prediction image generated by inter-frame prediction is also called the inter-frame prediction image.
[0439] Next, the predictor evaluates each of the intra-frame and inter-frame predicted images (step Sd_2). The cost C described above can be used in the evaluation. The predictor can then select the predicted image for which the minimum cost C has been calculated from the intra-frame and inter-frame predicted images as the final predicted image for the current block (step Sd_3). In other words, a prediction method or mode for generating the predicted image for the current block is selected.
[0440] The prediction processor then selects the prediction image for which the minimum cost C has been calculated from the intra-frame prediction images and inter-frame prediction images as the final prediction image for the current block (step Sd_3). In other words, a prediction method or mode for generating the prediction image for the current block is selected.
[0441] (Intra-frame predictor)
[0442] Intra-predictor 124 generates a prediction signal (i.e., an intra-predicted image) by performing intra-prediction of the current block by referencing one or more blocks in the current image and stored in block memory 118. More specifically, intra-predictor 124 generates an intra-predicted image by performing intra-prediction by referencing pixel values (e.g., luminance and / or chrominance values) i of one or more blocks adjacent to the current block, and then outputs the intra-predicted image to prediction controller 128.
[0443] For example, the intra predictor 124 performs intra prediction by using one of a plurality of predefined intra prediction modes. Intra prediction modes typically include one or more non-directional prediction modes and multiple directional prediction modes. The defined modes can be predefined.
[0444] One or more non-directional prediction modes include, for example, the planar prediction mode and the DC prediction mode as defined in the H.265 / High Efficiency Video Coding (HEVC) standard.
[0445] Multiple directional prediction modes include, for example, the thirty-three directional prediction modes defined in the H.265 / HEVC standard. It should be noted that, in addition to the thirty-three directional prediction modes, multiple directional prediction modes may also include thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31This is a conceptual diagram illustrating the total of sixty-seven intra-prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra-prediction. Solid arrows represent the thirty-three directions defined in the H.265 / HEVC standard, and dashed arrows represent the additional thirty-two directions. Figure 31 (Two non-directional prediction modes are not shown in the image).
[0446] In various processing examples, the luma block can be referenced in the intra-frame prediction of the chroma block. In other words, the chroma component of the current block can be predicted based on the luma component of the current block. This intra-frame prediction is also called Cross-Component Linear Model (CCLM) prediction. An intra-frame prediction mode for the chroma block that references such a luma block (also known as, for example, CCLM mode) can be added as one of the intra-frame prediction modes for the chroma block.
[0447] Intra-predictor 124 can correct the intra-predicted pixel values based on horizontal / vertical reference pixel gradients. Intra-prediction accompanied by this correction is also called position-dependent intra-prediction combination (PDPC). Information indicating whether PDPC is applied (e.g., a PDPC flag) is typically signaled at the CU level. It should be noted that such signaling does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, or CTU level).
[0448] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor 124.
[0449] Intra-predictor 124 selects one intra-prediction mode from a plurality of intra-prediction modes (step Sw_1). Intra-predictor 124 then generates a predicted image based on the selected intra-prediction mode (step Sw_2). Next, intra-predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra-prediction modes. For example, two of the six intra-prediction modes may be a planar mode and a DC prediction mode, and the other four modes may be directional prediction modes. Intra-predictor 124 determines whether the intra-prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0450] Here, when it is determined that the intra-prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra-predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the intra-prediction mode selected in these MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra-prediction mode can be encoded into prediction parameters by the entropy encoder 110.
[0451] When it is determined that the selected intra-prediction mode is not included in the MPM (No in step Sw_4), the intra-predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra-predictor 124 does not set any MPM flag. The intra-predictor 124 then generates information indicating the selected intra-prediction mode among at least one intra-prediction mode not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra-prediction mode can be encoded into prediction parameters by the entropy encoder 110. The information indicating the intra-prediction mode indicates, for example, any one of 0 to 60.
[0452] (Inter-frame predictor)
[0453] Inter-frame prediction (also known as inter-frame prediction) of the current block is performed by referencing one or more blocks in a reference image. Inter-frame predictor 126 generates a predicted image (inter-frame prediction image), which is different from the current image and stored in frame memory 122. Inter-frame prediction is performed on a unit basis: the current block or the current sub-block within the current block (e.g., a 4×4 block). Sub-blocks are included within blocks and are smaller than blocks. The size of a sub-block can be in the form of slices, bricks, images, etc.
[0454] For example, inter-frame predictor 126 performs motion estimation in a reference image of the current block or sub-block and finds a reference block or reference sub-block that best matches the current block or sub-block. Inter-frame predictor 126 then obtains motion information (e.g., motion vectors) to compensate for motion or variation from the reference block or reference sub-block to the current block or sub-block. Inter-frame predictor 126 generates an inter-frame predicted image of the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. Inter-frame predictor 126 outputs the generated inter-frame predicted image to prediction controller 128.
[0455] Motion information used in motion compensation can be signaled in various forms as inter-frame prediction signals. For example, motion vectors can be signaled. As another example, the difference between motion vectors and motion vector predictors can be signaled.
[0456] (List of reference images)
[0457] Figure 33 This is a concept diagram used to illustrate a reference image. Figure 34 This is a conceptual diagram used to illustrate an example of a list of reference images. The list of reference images is a list indicating at least one reference image stored in frame memory 122. It is worth noting that, in Figure 33In this diagram, each rectangle represents an image, each arrow represents an image reference relationship, the horizontal axis represents time, and the I, P, and B symbols within the rectangles represent intra-frame predicted images, single-predicted images, and double-predicted images, respectively. The numbers within the rectangles indicate the decoding order. For example... Figure 33 As shown, the decoding order of the images is I0, P1, B2, B3, B4, and the display order of the images is I0, B3, B2, B4, P1. Figure 34 As shown, the reference image list is a list representing candidate reference images. For example, an image (or slice) may include at least one reference image list. For instance, one reference image list is used when the current image is a single-prediction image, and two reference image lists are used when the current image is a double-prediction image. Figure 33 and Figure 34 In the example, image B3, which is the current image currPic, has two lists of reference images: the L0 list and the L1 list. When the current image currPic is image B3, the candidate reference images for the current image currPic are I0, P1, and B2, and the list of reference images (i.e., the L0 list and the L1 list) indicates these images. The inter-frame predictor 126 or the prediction controller 128 specifies which image in each list to actually reference in the form of a reference image index refidxLx. Figure 34 In the text, reference images P1 and B2 are specified by reference image indices refIdxL0 and refIdxL1.
[0458] Such a list of reference images can be generated for each unit, such as a sequence, picture, slice, block, CTU, or CU. Furthermore, the reference image index indicating which reference image will be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, block level, CTU level, or CU level. Additionally, a common list of reference images can be used across multiple inter-frame prediction modes.
[0459] (Basic process of inter-frame prediction)
[0460] Figure 35 This is a flowchart illustrating an example of the basic processing flow for inter-frame prediction.
[0461] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates the difference between the current block and the prediction image as the prediction residual (step Se_4).
[0462] Here, in the generation of the predicted image, the inter-frame predictor 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Furthermore, in determining the MV, the inter-frame predictor 126 determines the MV by selecting motion vector candidates (MV candidates) (step Se_1) and deriving the MV (step Se_2). The selection of MV candidates is performed, for example, by the inter-frame predictor 126 generating a list of MV candidates and selecting at least one MV candidate from the list. It should be noted that previously derived MVs can be added to the MV candidate list. Alternatively, in MV derivation, the inter-frame predictor 126 may also select at least one MV candidate from at least one MV candidate and determine the selected at least one MV candidate as the MV of the current block. Alternatively, the inter-frame predictor 126 can determine the MV of the current block by performing estimation in a specified reference image region from each of the selected at least one MV candidate. It should be noted that the estimation in the reference image region can be referred to as motion estimation.
[0463] In addition, although steps Se_1 to Se_3 are performed by the inter-frame predictor 126 in the above example, the processes of steps Se_1, Se_2, etc., can be performed by another constituent element included in the encoder 100.
[0464] It should be noted that an MV candidate list can be generated for each process in inter-frame prediction mode, or a common MV candidate list can be used in multiple inter-frame prediction modes. The processes in steps Se_3 and Se_4 respectively correspond to... Figure 9 Steps Sa_3 and Sa_4 are shown. The process in step Sa_3 corresponds to... Figure 30 The process in step Sd_1b.
[0465] (Derivation process of motion vectors)
[0466] Figure 36 This is a flowchart illustrating an example of the derivation process of the motion vector.
[0467] The inter-frame predictor 126 can derive the MV of the current block in a mode used to encode motion information (e.g., MV). In this case, for example, the motion information can be encoded as prediction parameters and can be signaled. In other words, the encoded motion information is included in the stream.
[0468] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where motion information is not encoded. In this case, motion information is not included in the stream.
[0469] Here, the MV derivation mode can include regular inter-frame mode, regular merging mode, FRUC mode, affine mode, etc., as described later. Modes that encode motion information include regular inter-frame mode, regular merging mode, affine mode (specifically, affine inter-frame mode and affine merging mode), etc. It should be noted that the motion information can include not only MV, but also the motion vector predictor selection information described later. Modes that do not encode motion information include FRUC mode, etc. The inter-frame predictor 126 selects a mode from multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.
[0470] Figure 37 This is a flowchart illustrating another example of the derivation of motion vectors.
[0471] The inter-frame predictor 126 can derive the MV of the current block in a mode where the MV difference is encoded. In this case, for example, the MV difference can be encoded as prediction parameters and can be signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. It should be noted that the MV predictor is a motion vector predictor.
[0472] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where the MV difference is not encoded. In this case, the encoded MV difference is not included in the stream.
[0473] Here, as described above, the MV derivation modes include the regular inter-frame mode, regular merging mode, FRUC mode, affine mode, etc., as described later. Modes that encode the MV difference include the regular inter-frame mode and the affine mode (specifically, the affine inter-frame mode). Modes that do not encode the MV difference include the FRUC mode, the regular merging mode, and the affine mode (specifically, the affine merging mode). The inter-frame predictor 126 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0474] (Motion vector derivation mode)
[0475] Figure 38A and Figure 38B This is a conceptual diagram used to illustrate example classifications of patterns used for MV derivation. For example, such as... Figure 38A As shown, based on whether motion information and MV difference are encoded, MV derivation modes can be broadly classified into three modes: inter-frame mode, merge mode, and frame rate up-conversion (FRUC) mode. Inter-frame mode performs motion estimation and encodes both motion information and MV difference. For example, as... Figure 38BAs shown, inter-frame modes include affine inter-frame mode and regular inter-frame mode. Merging mode is a mode that does not perform motion estimation and selects the motion difference (MV) from the coded surrounding blocks, using that MV to derive the MV of the current block. Merging mode is a mode that essentially encodes motion information without encoding the MV difference. For example, as... Figure 38B As shown, the merging modes include regular merging mode (also known as normal merging mode or normal merging mode), merge with motion vector difference (MMVD) mode, combined inter-frame merging / intra-frame prediction (CIIP) mode, triangular mode, ATMVP mode, and affine merging mode. Here, in the MMVD mode included in the merging modes, the MV difference is exceptionally encoded. It should be noted that affine merging mode and affine inter-frame mode are modes included within affine modes. An affine mode is used to derive the MV of each of the multiple sub-blocks included in the current block as the MV of the current block, assuming an affine transformation. The FRUC mode is a mode used to derive the MV of the current block by performing estimations between coded regions, and neither encodes motion information nor any MV difference. Note that the corresponding modes will be described in more detail later.
[0476] It should be noted that Figure 38A and Figure 38B The classification of modes shown is illustrative, and the classification is not limited to this. For example, when MV difference is encoded in CIIP mode, CIIP mode is classified as an inter-frame mode.
[0477] (MV Derivation > Regular Inter-Frame Mode)
[0478] The regular inter-frame mode is an inter-frame prediction mode used to derive the MV of the current block from a reference picture region specified by MV candidates, based on blocks similar to the current block in the image. In this regular inter-frame mode, the MV difference is encoded.
[0479] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in a regular inter-frame mode.
[0480] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks surrounding the current block in time or space (step Sg_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0481] Next, the inter-frame predictor 126 extracts N (2 or larger integers) MV candidates from the multiple MV candidates obtained in step Sg_1 as motion vector predictor candidates (also called MV predictor candidates) according to the determined priority order (step Sg_2). It should be noted that the priority order can be predetermined for each of the N MV candidates.
[0482] Next, the inter-frame predictor 126 selects a motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor (also called the MV predictor) for the current block (step Sg_3). At this time, the inter-frame predictor 126 encodes motion vector predictor selection information in the stream to identify the selected motion vector predictor. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0483] Next, the inter-frame predictor 126 derives the motion vector prediction (MV) of the current block using the reference image encoded with reference (step Sg_4). At this point, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as the MV difference within the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130. It is important to note that the encoded reference image includes images of multiple blocks that have already been reconstructed after encoding.
[0484] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference image, the inter-frame predictor 126 generates the predicted image for the current block (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in a slice, the inter-frame prediction for the slice using the regular inter-frame mode ends. Similarly, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in an image, the inter-frame prediction for the image using the regular inter-frame mode ends. It should be noted that not all blocks included in the slice undergo these processes in steps Sg_1 to Sg_5, and the inter-frame prediction for the slice using the regular inter-frame mode can end when only some blocks undergo the process. This also applies to the processes in steps Sg_1 to Sg_5. When the process is performed on only some blocks in an image, the inter-frame prediction for the image using the regular inter-frame mode can end.
[0485] It should be noted that the predicted image is the inter-frame prediction signal as described above. Furthermore, information representing the inter-frame prediction mode used to generate the predicted image (the conventional inter-frame mode in the example above) is encoded, for example, as prediction parameters in the coded signal.
[0486] It is important to note that the MV candidate list can also be used as a list in another pattern. Furthermore, procedures related to the MV candidate list can be applied to list-related procedures for use in another pattern. Procedures related to the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering MV candidates, or deleting MV candidates.
[0487] (MV derivation > Standard merge mode)
[0488] The regular merging mode is used to derive the inter-frame prediction mode of the MV by selecting MV candidates from the MV candidate list as the MV of the current block. It should be noted that the regular merging mode is a type of merging mode and can be simply referred to as the merging mode. In this embodiment, the regular merging mode and the merging mode are distinguished, and the merging mode is used in a broader sense.
[0489] Figure 40 This is a flowchart illustrating an example of inter-frame prediction in a regular merging mode.
[0490] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks around the current block in time or space (step Sh_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0491] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter-frame predictor 126 encodes MV selection information in the stream to identify the selected MV candidate. In other words, the inter-frame predictor 126 outputs the MV selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0492] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference image, the inter-frame predictor 126 generates the predicted image for the current block (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a slice, the inter-frame prediction of the slice using the regular merging mode ends. Similarly, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in an image, the inter-frame prediction of the image using the regular merging mode ends. It should be noted that not all blocks included in a slice undergo the processes in steps Sh_1 to Sh_3, and when only some blocks undergo the process, the inter-frame prediction of the slice using the regular merging mode can end. This also applies to the processes in steps Sh_1 to Sh_3. When the process is performed on only some blocks in an image, the inter-frame prediction of the image using the regular merging mode can be completed.
[0493] Additionally, information contained in the encoded signal and used to generate the predicted image, representing the inter-frame prediction mode (in the example above, the regular merging mode), is encoded, for example, as prediction parameters in the stream.
[0494] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process for the current image through a regular merging pattern.
[0495] First, the inter-frame predictor 126 generates a list of MV candidates registered therein. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of multiple coded blocks located spatially around the current block; temporally adjacent MV candidates, which are MVs of surrounding blocks, the position of the current block in the coded reference picture being projected onto these surrounding blocks; combined MV candidates, which are MVs generated by combining the MV values of spatially adjacent MV predictors and temporally adjacent MV predictors; and zero MV candidates, which are MVs with a value of zero.
[0496] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates registered in the MV candidate list and determines that MV candidate as the MV of the current block.
[0497] In addition, the entropy encoder 110 writes and encodes merge_idx in the stream, which is a signal indicating which MV candidate has been selected.
[0498] It should be noted that registration is in Figure 41 The MV candidates in the MV candidate list described are examples. The number of MV candidates may differ from the number of MV candidates in the figure. The MV candidate list can be configured in such a way that it may exclude some types of MV candidates in the figure, or include one or more types of MV candidates other than those in the figure.
[0499] The final MV is determined by performing Dynamic Motion Vector Refresh (DMVR), which will be described later, using the MV of the current block derived from the regular merge mode. It's important to note that in the regular merge mode, motion information is encoded, but the MV difference is not. In MMVD mode, an MV candidate is selected from the MV candidate list, and the MV difference is encoded as in the regular merge mode. Figure 38B As shown, MMVD can be classified as a merge mode along with the regular merge mode. It should be noted that the MV difference in MMVD mode does not always need to be the same as the MV difference used in inter-frame mode. For example, MV difference derivation in MMVD mode can be a process requiring less processing power than MV difference derivation in inter-frame mode.
[0500] In addition, a combined inter-frame merge / intra-frame prediction (CIIP) mode can be executed. This mode is used to overlap the prediction images generated in inter-frame prediction and the prediction images generated in intra-frame prediction to generate the prediction image for the current block.
[0501] It's important to note that the MV candidate list can be simply called a candidate list. Additionally, merge_idx contains MV selection information.
[0502] (MV Derivation > HMVP Pattern)
[0503] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process for the current image using the HMVP merging pattern.
[0504] In the regular merge mode, the MV of, for example, the CU that is the current block is determined by selecting an MV candidate from the list of MV candidates generated from the reference coding block (e.g., CU). Here, another MV candidate can be registered in the list of MV candidates. The mode of registering such another MV candidate is called the HMVP mode.
[0505] In HMVP mode, HMVP's First-In-First-Out (FIFO) server is used to manage MV candidates, separate from the MV candidate list in regular merge mode.
[0506] In a FIFO buffer, motion information such as the MV (Motion Value) of the most recently processed block is stored first. To manage the FIFO buffer, each time a block is processed, the MV of the most recent block (i.e., the CU immediately following the previous block) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the earliest processed CU) is removed from the FIFO buffer. Figure 42 In the example shown, HMVP1 is the MV of the most recent block, and HMVP5 is the MV of the oldest MV.
[0507] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is a different MV from all MV candidates already registered in the MV candidate list for the regular merge mode starting from HMVP1. When it is determined that the MV is different from all MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer as an MV candidate to the MV candidate list for the regular merge mode. At this point, one or more MV candidates in the FIFO buffer can be registered (added to the MV candidate list).
[0508] By using the HMVP pattern in this way, not only can the MVs of blocks spatially or temporally adjacent to the current block be added, but also the MVs of blocks that have been processed in the past. As a result, the variation in MV candidates is expanded compared to the regular merge pattern, which increases the possibility of improving coding efficiency.
[0509] Note that MV can be motion information. In other words, the information stored in the MV candidate list and FIFO buffer can include not only MV values, but also reference image information, reference orientation, number of images, etc. Additionally, blocks can be, for example, CUs (Cubic Cutoffs).
[0510] Notice, Figure 42 The MV candidate list and FIFO buffer shown are examples. The size of the MV candidate list and FIFO buffer can be... Figure 42 The differences in, or can be configured to, with Figure 42 The MV candidates are registered in different orders. Furthermore, the process described herein can be common between encoder 100 and decoder 200.
[0511] Note that the HMVP pattern can be applied to patterns other than the regular merge pattern. For example, motion information such as the MV of blocks previously processed in the affine pattern can be stored up-to-date and used as MV candidates, which can promote better efficiency. The pattern obtained by applying the HMVP pattern to the affine pattern can be called the historical affine pattern.
[0512] (MV Derivation > FRUC Pattern)
[0513] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, motion estimation is performed on the decoder side without using any pixel values from the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values from the current block include Frame Rate Upconversion (FRUC) mode, Pattern Matching Motion Vector Derivation (PMMVD) mode, etc.
[0514] Figure 43 The diagram illustrates an example of the FRUC process in flowchart form. First, the MVs of each of the coded blocks that are spatially or temporally adjacent to the current block are indicated by referring to the MVs, as a list of MV candidates (this list can be an MV candidate list, or it can be used as an MV candidate list for the regular merge mode) (step Si_1).
[0515] Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation values of the corresponding MV candidates included in the MV candidate list are calculated, and an MV candidate is selected based on the evaluation value. Based on the selected motion vector candidate, the motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector of the current block. Alternatively, for example, pattern matching in the region surrounding the position in the reference image can be used to derive the motion vector of the current block, where the position in the reference image corresponds to the selected motion vector candidate. In other words, the estimation using pattern matching and evaluation values can be performed in the region surrounding the best MV candidate, and when an MV that produces a better evaluation value exists, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, updating the motion vector that produces a better evaluation value may not be performed.
[0516] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference image, the inter-frame predictor 126 generates a predicted image for the current block (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in an image, the inter-frame prediction of the image using the FRUC mode ends. Note that not all blocks contained in a slice undergo the processes in steps Si_1 to Si_5, and the inter-frame prediction of the slice using the FRUC mode can end when only some blocks undergo the process. When the processes in steps Si_1 to Si_5 are performed on only some blocks included in an image in a similar manner, the inter-frame prediction of the image using the FRUC mode can end.
[0517] A similar process can be performed on a sub-block basis.
[0518] Evaluation values can be calculated using various methods. For example, a comparison can be made between a reconstructed image in a region of a reference image corresponding to a motion vector and a reconstructed image in a determined region (which could be, for example, a region in another reference image or a region in a neighboring block of the current image, as shown below). The determined region can be predetermined.
[0519] The difference between pixel values in two reconstructed images can be used to evaluate the motion vector. Note that information other than the difference can be used to calculate the evaluation.
[0520] Next, an example of pattern matching is described in detail. First, a candidate MV contained in the candidate MV list (e.g., a merged list) is selected as the starting point for estimation by pattern matching. For example, as pattern matching, either first pattern matching or second pattern matching can be used. First pattern matching and second pattern matching can be referred to as bilateral matching and template matching, respectively.
[0521] (MV derivation > FRUC > bilateral matching)
[0522] In the first pattern matching, pattern matching is performed between two blocks located along the motion trajectory of the current block and contained in two different reference images. Therefore, in the first pattern matching, a region in another reference image along the motion trajectory of the current block is used as the region to determine the evaluation value of the aforementioned candidates. This determined region can be predetermined.
[0523] Figure 44 This is a conceptual diagram used to illustrate an example of a first pattern matching (bilateral matching) between two blocks in two reference images along a motion trajectory. (See diagram below.) Figure 44 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the best matching pair among pairs of two blocks included in two different reference images (Ref0, Ref1) and located along the motion trajectory of the current block (Cur block). More specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference image (Ref0) specified by the MV candidate and the reconstructed image at a specified position in the second coded reference image (Ref1) specified by the symmetric MV obtained by scaling the MV candidate at display time intervals is derived, and the obtained difference is used to calculate an evaluation value. The MV candidate that produces the best evaluation value and is likely to produce good results can be selected as the final MV from among the multiple MV candidates.
[0524] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) of two reference blocks are specified to be proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, when the current image is located temporally between two reference images and the temporal distances from the current image to the corresponding two reference images are equal, a mirror-symmetric bidirectional motion vector is derived in the first pattern matching.
[0525] (MV derivation > FRUC > Template matching)
[0526] In the second pattern matching (template matching), pattern matching is performed between a block in the reference image and a template in the current image (the template being a block in the current image adjacent to the current block (adjacent blocks being, for example, the upper and / or left adjacent blocks)). Therefore, in the second pattern matching, the adjacent blocks of the current block in the current image are used as the determined regions for calculating the evaluation values of the aforementioned MV candidates.
[0527] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 45 As shown, in the second pattern matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference image (Ref0) that best matches the neighboring blocks of the current block in the current image (Cur Pic). More specifically, the difference between the reconstructed image in the coded region adjacent to the left and above, or to the left and above, and the reconstructed image in the corresponding region in the coded reference image (Ref0) specified by the MV candidate is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.
[0528] Information indicating whether FRUC mode is applied (e.g., referred to as the FRUC flag) can be signaled at the CU level. Furthermore, when FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (e.g., first pattern matching or second pattern matching) can be signaled at the CU level. It should be noted that signaling for this type of information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0529] (MV derivation > Affine mode)
[0530] Affine patterns are patterns that use affine transformations to generate motion vectors (MVs). For example, MVs can be derived on a sub-block basis based on the motion vectors of multiple neighboring blocks. This pattern is also known as an affine motion compensation prediction pattern.
[0531] Figure 46A This is a conceptual diagram used to illustrate an example of MV derivation based on motion vectors of multiple adjacent blocks, on a sub-block basis. Figure 46A In this context, the current block comprises, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the top-left control point of the current block is derived based on the motion vectors of neighboring blocks, and similarly, the motion vector V1 at the top-right control point of the current block is derived based on the motion vectors of neighboring sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (v1, v2, v3, v4, v1) of the corresponding sub-blocks in the current block can be derived.x ,v y ).
[0532] [Mathematical Expression 1]
[0533]
[0534] This information indicating the affine pattern (e.g., called an affine flag) can be signaled at the CU level. Note that the signaling indicating the affine pattern does not necessarily need to be executed at the CU level and can be executed at another level (e.g., at the sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0535] Furthermore, affine modes can include several modes that use different methods to derive motion vectors at the top-left and top-right control points. For example, affine modes include two modes: affine inter-frame mode (also known as affine regular inter-frame mode) and affine merge mode.
[0536] (MV derivation > Affine mode)
[0537] Figure 46B This is a conceptual diagram used to illustrate an example of MV derivation in a sub-block manner within an affine pattern that uses three control points. Figure 46B In this context, the current block comprises, for example, sixteen 4×4 blocks. Here, the motion vector V0 at the top-left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector V1 at the top-right control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V2 at the bottom-left control point of the current block is derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v1, v2, v2) of the corresponding sub-blocks in the current block can be derived. x ,v y ).
[0538] [Mathematical Expression 2]
[0539]
[0540] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w and h can be weighting coefficients, which can be predetermined weighting coefficients. In this embodiment, w can represent the width of the current block, and h can represent the height of the current block.
[0541] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0542] Furthermore, this affine mode using three control points can include different methods for deriving motion vectors at the top-left, top-right, and bottom-left control points. For example, similar to the affine mode using two control points, the affine mode using three control points can include both affine inter-frame mode and affine merge mode.
[0543] Note that in affine mode, the size of each sub-block contained within the current block is not limited to 4×4 pixels and can be other sizes. For example, the size of each sub-block can be 8×8 pixels.
[0544] (MV derivation > Affine pattern > Control point)
[0545] Figure 47A , Figure 47B and Figure 47C This is a conceptual diagram used to illustrate an example of MV derivation at control points in an affine mode.
[0546] like Figure 47A As shown, in affine mode, for example, based on multiple motion vectors corresponding to blocks encoded according to the affine mode among the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, a motion vector predictor at the corresponding control point of the current block is calculated. More specifically, the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in the listed order, and the first valid block encoded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on multiple motion vectors corresponding to the identified block.
[0547] For example, such as Figure 47B As shown, when block A, which is adjacent to the left of the current block, has been encoded according to an affine pattern using two control points, motion vectors v3 and v4 projected onto the upper left and upper right corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the upper left control point and motion vector v1 at the upper right control point of the current block are calculated.
[0548] For example, such as Figure 47CAs shown, when block A, which is adjacent to the left of the current block, has been encoded according to an affine pattern using three control points, motion vectors v3, v4, and v5 projected onto the top-left, top-right, and bottom-left positions of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, motion vector v0 at the top-left control point of the current block, motion vector v1 at the top-right control point of the current block, and motion vector v2 at the bottom-left control point of the current block are calculated.
[0549] Figures 47A to 47C The MV derivation method shown can be used in Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown is used, or can be used for what will be described later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.
[0550] Figure 48A and Figure 48B This is a conceptual diagram used to illustrate an example of MV derivation at control points in an affine mode.
[0551] Figure 48A This is a conceptual diagram used to illustrate an exemplary affine pattern using two control points.
[0552] In affine mode, such as Figure 48A As shown, the MV selected from the MVs of coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block.
[0553] Figure 48B This is a conceptual diagram used to illustrate an exemplary affine pattern using three control points.
[0554] In affine mode, such as Figure 48B As shown, the MV selected from the MVs of coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block. Furthermore, the MV selected from the MVs of coded blocks F and G adjacent to the current block is used as the motion vector v2 at the bottom-left control point of the current block.
[0555] Notice, Figure 48A and Figure 48B The MV derivation method shown can be used for the purposes described later. Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in the figure, or which can be used for the description later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.
[0556] Here, when affine modes using different numbers of control points (e.g., two and three control points) can be switched at the CU level and signaled, the number of control points in the coded block and the number of control points in the current block can be different from each other.
[0557] Figure 49A and Figure 49B This is a conceptual diagram illustrating an example of a method for deriving the MV at a control point when the number of control points in the coded block differs from the number of control points in the current block.
[0558] For example, such as Figure 49A As shown, the current block has three control points at the top left, top right, and bottom left corners, and block A, which is adjacent to the left of the current block, has already been encoded using an affine pattern with two control points. In this case, motion vectors v3 and v4 projected onto the top left and top right corners of the encoded block including block A are derived. Then, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated based on the derived motion vectors v3 and v4. Furthermore, motion vector v2 at the bottom left control point is calculated based on the derived motion vectors v0 and v1.
[0559] For example, such as Figure 49B As shown, the current block has two control points at its top left and top right corners, and the block A adjacent to the left of the current block has been encoded according to an affine pattern using three control points. In this case, motion vectors v3, v4, and v5 projected onto the top left, top right, and bottom left corners of the encoded block including block A are derived. Then, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0560] Notice, Figure 49A and Figure 49B The MV derivation method shown can be used in the descriptions that follow. Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown, or which can be used for descriptions later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.
[0561] (MV Derivation > Affine Pattern > Affine Merging Pattern)
[0562] Figure 50 This is a flowchart illustrating an example of the process in the affine merge pattern.
[0563] In the affine merging mode shown in the figure, firstly, the inter-frame predictor 126 derives the MV at the corresponding control point of the current block (step Sk_1). For example... Figure 46A As shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 46B The image shows the top-left corner, top-right corner, and bottom-left corner of the current block. The inter-frame predictor 126 can encode MV selection information to identify two or three derived MVs in the stream.
[0564] For example, when using Figures 47A to 47C When using the MV derivation method shown, such as Figure 47A As shown, the inter-frame predictor 126 examines the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in the listed order and identifies the first valid block encoded according to the affine pattern.
[0565] Inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47B As shown, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the coded block including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block by projecting the motion vectors v3 and v4 at the top-left and top-right control points of the coded block onto the current block.
[0566] Alternatively, when block A is identified and block A has three control points, such as Figure 47C As shown, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point, the motion vector v1 at the top-right control point, and the motion vector v2 at the bottom-left control point of the current block based on the motion vectors v3, v4, and v5 at the top-left, top-right, and bottom-left control points of the coded block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point, the motion vector v1 at the top-right control point, and the motion vector v2 at the bottom-left control point of the current block by projecting the motion vectors v3, v4, and v5 at the top-left, top-right, and bottom-left control points of the coded block onto the current block.
[0567] Note that, as mentioned above, Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, and as described above... Figure 49B As shown, when block A is identified and block A has three control points, the MV at two control points can be calculated.
[0568] Next, the inter-frame predictor 126 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the motion vector (MV) of each of the multiple sub-blocks as an affine MV, for example, using two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter-frame predictor 126 then uses these affine MVs and the encoded reference image to perform motion compensation for the sub-blocks (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all sub-blocks included in the current block, the process of generating a predicted image using the affine merging pattern of the current block ends. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.
[0569] Note that the MV candidate list described above can be generated in step Sk_1. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown Figure 49A and Figure 49B The MV derivation method shown is any combination of other MV derivation methods.
[0570] Note that, in addition to affine mode, the MV candidate list can include MV candidates in modes where prediction is performed on a sub-block basis.
[0571] Note that, for example, an MV candidate list (which includes MV candidates in affine merge patterns using two control points and affine merge patterns using three control points) can be generated as an MV candidate list. Alternatively, an MV candidate list including MV candidates in affine merge patterns using two control points and an MV candidate list including MV candidates in affine merge patterns using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge patterns using two control points and affine merge patterns using three control points can be generated. MV candidates can be, for example, the MVs of blocks A (left), B (top), C (top right), D (bottom left), and E (top left) used for encoding, or the MVs of valid blocks within a block.
[0572] Note that the index indicating one of the MVs in the MV candidate list can be transmitted as MV selection information.
[0573] (MV Derivation > Affine Mode > Affine Inter-Frame Mode)
[0574] Figure 51 This is a flowchart illustrating an example of a process in affine inter-frame mode.
[0575] In affine inter-frame mode, firstly, the inter-frame predictor 126 derives the MV predictor (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the top-left corner, the top-right corner, and the top-right corner of the current block, such as... Figure 46A or Figure 46B As shown.
[0576] For example, when using Figure 48A and Figure 48B When the MV derivation method is shown, the inter-frame predictor 126 selects... Figure 48A or Figure 48B The MV of any block in the coded block near the corresponding control point of the current block is used to derive the MV predictor (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block. At this time, the inter-frame predictor 126 encodes MV predictor selection information in the stream to identify the selected two or three MV predictors.
[0577] For example, the inter-frame predictor 126 can use cost evaluation or the like to determine the block from which the MV predictor is selected as the control point from the coded blocks adjacent to the current block, and can write a flag indicating which MV predictor has been selected into the bitstream. In other words, the inter-frame predictor 126 outputs MV predictor selection information, such as the flag, as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.
[0578] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 uses the above expression (1A) or expression (1B) to calculate the MV of each sub-block corresponding to the updated MV predictor as an affine MV (step Sj_3). The inter-frame predictor 126 then uses these affine MVs and encoded reference images to perform motion compensation for the sub-blocks (step Sj_4). When the MV predictor is updated in step Sj_2, the process in steps Sj_3 and Sj_4 is performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor as an MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0579] Finally, the inter-frame predictor 126 generates a predicted image of the current block by performing motion compensation for the current block using the determined MV and the encoded reference image (step Sj_6).
[0580] Note that the MV candidate list described above can be generated in step Sj_1. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown Figure 49A and Figure 49B The MV derivation method shown is any combination of other MV derivation methods.
[0581] Note that, in addition to affine mode, the MV candidate list can include MV candidates in modes that perform prediction on a sub-block basis.
[0582] Note that, for example, an MV candidate list can be generated that includes MV candidates in affine inter-frame modes using two control points and affine inter-frame modes using three control points. Alternatively, an MV candidate list can be generated separately, including MV candidates in affine inter-frame modes using two control points and MV candidate lists that include MV candidates in affine inter-frame modes using three control points. Alternatively, an MV candidate list can be generated that includes MV candidates in one of the affine inter-frame modes using two control points and affine inter-frame modes using three control points. MV candidates can be, for example, the MVs of blocks A (left), B (top), C (top right), D (bottom left), and E (top left) used for encoding, or the MVs of valid blocks within a block.
[0583] Note that the index indicating one of the MV candidates in the MV candidate list can be transmitted as MV predictor selection information.
[0584] (MV Derivation > Triangle Pattern)
[0585] In the example above, the inter-frame predictor 126 generates a rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 can generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and can combine multiple prediction images to generate a final rectangular prediction image. The shape other than a rectangle can be, for example, a triangle.
[0586] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangular predicted images.
[0587] Inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a first partition with a triangular shape in the current block using a first MV of a first partition. Similarly, inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a second partition with a triangular shape in the current block using a second MV of a second partition. Then, inter-frame predictor 126 combines these prediction images to generate a prediction image with a rectangular shape that is the same as the rectangular shape of the current block.
[0588] Note that a first prediction image with a rectangular shape corresponding to the current block can be generated using the first prediction image (MV) as the prediction image for the first partition. Furthermore, a second prediction image with a rectangular shape corresponding to the current block can be generated using the second prediction image (MV) as the prediction image for the second partition. The prediction image for the current block can be generated by performing a weighted sum of the first and second prediction images. Note that the portion to be weighted can be a region spanning the boundary between the first and second partitions.
[0589] Figure 52BThis is a conceptual diagram illustrating an example of a first portion of a first partition overlapping with a second partition, and first and second sample sets that can be weighted as part of a correction process. The first portion may be, for example, one-quarter of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to the edge of the first partition, where N is a positive integer, for example, N could be an integer 2. As shown in the figure... Figure 52B The example on the left shows a rectangular partition with a rectangular portion whose width is one-quarter of the width of the first partition, wherein the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples inside the first part. Figure 52B The central example shows a rectangular partition with a rectangular portion whose height is one-quarter of the height of a first partition, wherein a first sample set includes samples outside the first part and samples inside the first part, and a second sample set includes samples inside the first part. Figure 52B The example on the right shows a triangular partition with a polygonal portion, the height of which corresponds to two samples, wherein the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples inside the first part.
[0590] The first part can be the portion of the first partition that overlaps with the adjacent partition. Figure 52C This is a conceptual diagram illustrating a first portion of a first partition, which is the part of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition with an overlapping portion that is spatially adjacent to a rectangular partition is shown. Partitions with other shapes can be used, such as triangular partitions, and the overlapping portion can overlap with spatially or temporally adjacent partitions.
[0591] Furthermore, although an example of generating a predicted image for each of two partitions using inter-frame prediction is given, a predicted image for at least one partition can be generated using intra-frame prediction.
[0592] Figure 53 This is a flowchart illustrating an example of a process in a triangle pattern.
[0593] In the triangle mode, firstly, the inter-frame predictor 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode the partition information (which is information related to the partitioning) into prediction parameters in the stream. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0594] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks surrounding the current block in time or space (step Sx_2). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0595] Inter-frame predictor 126 then selects the MV candidate of the first partition and the MV candidate of the second partition from the plurality of MV candidates obtained in step Sx_1 as the first MV and the second MV, respectively (step Sx_3). At this time, inter-frame predictor 126 encodes MV selection information used to identify the selected MV candidate as prediction parameters in the stream. In other words, inter-frame predictor 126 outputs the MV selection information as prediction parameters to entropy encoder 110 through prediction parameter generator 130.
[0596] Next, the inter-frame predictor 126 performs motion compensation using the selected first MV and the encoded reference image to generate a first predicted image (step Sx_4). Similarly, the inter-frame predictor 126 performs motion compensation using the selected second MV and the encoded reference image to generate a second predicted image (step Sx_5).
[0597] Finally, the inter-frame predictor 126 generates the prediction image for the current block by performing a weighted sum of the first and second prediction images (step Sx_6).
[0598] Note that, although in Figure 52A In the example shown, the first and second partitions are triangles, but the first and second partitions can be trapezoids, or other shapes that are different from each other. Furthermore, although the current block is... Figure 52A and Figure 52C The example shown includes two partitions, but the current block can include three or more partitions.
[0599] Furthermore, the first and second partitions can overlap. In other words, the first and second partitions can include the same pixel region. In this case, the predicted image of the current block can be generated using the predicted image from the first and second partitions.
[0600] Furthermore, although an example of generating a predicted image for each of two partitions using inter-frame prediction has been shown, a predicted image for at least one partition can be generated using intra-frame prediction.
[0601] Note that the MV candidate list used to select the first MV and the MV candidate list used to select the second MV can be different from each other, or the MV candidate list used to select the first MV can also be used as the MV candidate list used to select the second MV.
[0602] Note that partitioning information may include indexes indicating the partitioning direction in which at least the current block is divided into multiple partitions. MV selection information may include indexes indicating the selected first MV and indexes indicating the selected second MV. One index may indicate multiple pieces of information. For example, an index that jointly indicates part or all of the partitioning information and part or all of the MV selection information may be encoded.
[0603] (MV Derivation > ATMVP Pattern)
[0604] Figure 54 This is a conceptual diagram used to illustrate an example of an advanced temporal motion vector prediction (ATMVP) pattern that derives MV on a sub-block basis.
[0605] The ATMVP pattern is a pattern that is classified as a merge pattern. For example, in the ATMVP pattern, the MV candidate for each sub-block is registered in the MV candidate list for use in the regular merge pattern.
[0606] More specifically, in the ATMVP pattern, firstly, as Figure 54 As shown, a temporal MV reference block associated with the current block is identified in the encoding reference image specified by the MV (MV0) of the neighboring block located at the lower left position relative to the current block. Next, in each sub-block within the current block, an MV is identified for encoding the region corresponding to the sub-block in the temporal MV reference block. MVs identified in this way are included in the MV candidate list as MV candidates for sub-blocks within the current block. When selecting an MV candidate for each sub-block from the MV candidate list, the sub-block undergoes motion compensation, where the MV candidate is used as the MV of the sub-block. In this way, a predicted image for each sub-block is generated.
[0607] Although Figure 54 In the example shown, the block located at the bottom left relative to the current block is used as the surrounding MV reference block, but it should be noted that another block can be used. Additionally, the size of the child block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the child block can be toggled against units such as slices, bricks, images, etc.
[0608] (MV derivation > DMVR)
[0609] Figure 55 This is a flowchart illustrating the relationship between the merging mode and the decoder motion vector refinement DMVR.
[0610] Inter-frame predictor 126 derives the motion vector of the current block based on the merging mode (step S1_1). Next, inter-frame predictor 126 determines whether to perform motion vector estimation, i.e., motion estimation (step S1_2). Here, when it is determined that motion estimation should not be performed (no in step S1_2), inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merging mode.
[0611] When it is determined in step S1_1 that motion estimation will be performed (Yes in step S1_2), the inter-frame predictor 126 derives the final motion vector of the current block by estimating the region surrounding the reference image specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to the DMVR.
[0612] Figure 56 This is a conceptual diagram illustrating an example of the DMVR process used to determine the MV.
[0613] First, for example in merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference image (L0), which is an coded image in the L0 list, based on the MV candidate (L0). Similarly, reference pixels are identified from the second reference image (L1), which is an coded image in the L1 list, based on the MV candidate (L1). A template is generated by calculating the average of these reference pixels.
[0614] Next, the template is used to estimate the MV candidate in the surrounding regions of the first reference image (L0) and the second reference image (L1), and the MV that produces the minimum cost is determined as the final MV. Note that the cost can be calculated, for example, using the difference between each pixel value in the template and the corresponding pixel value in the estimated region, the value of the MV candidate, etc.
[0615] It is not always necessary to perform the exact same procedure described here. Other procedures can be used to derive the final MV by estimating the region surrounding the MV candidate.
[0616] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV. Unlike Figure 56 The example of DMVR shown is in Figure 57 In the example shown, the cost is calculated without generating a template.
[0617] First, the inter-frame predictor 126 estimates the surrounding region of the reference block in each reference image contained in the L0 and L1 lists based on the initial MVs as MV candidates obtained from each MV candidate list. For example, as Figure 57 As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets the search position for the reference images in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set (specifically, the initial MV (i.e., InitMV_L0)), the vector difference with the search position is MVd_L0. The inter-frame predictor 126 then determines the estimated position in the reference images in the L1 list. This search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines the search position in each reference image in the L0 and L1 lists as the position symmetrical with respect to the position indicated by the initial MV. The inter-frame predictor 126 calculates the sum of the absolute differences (SAD) between pixel values at the search location in the block as the cost for each search location, and finds the search location that produces the minimum cost.
[0618] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in a DMVR, and Figure 58B This is a flowchart illustrating an example of the motion estimation process.
[0619] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also called the starting point) and eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at a search position other than the starting point is determined to be minimized, the inter-frame predictor 126 changes its objective to obtaining the search position with the minimum cost and executes the process in step 2. When the cost at the starting point is minimized, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.
[0620] In step 2, the inter-frame predictor 126 performs a search similar to that in step 1, taking the search position after the target change as the new starting point based on the result of the process in step 1. Then, the inter-frame predictor 126 determines whether the cost is minimized at each search position other than the starting point. Here, when the cost is minimized at each search position other than the starting point, the inter-frame predictor 126 performs the process in step 4. When the cost is minimized at the starting point, the inter-frame predictor 126 performs the process in step 3.
[0621] In step 4, the inter-frame predictor 126 considers the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0622] In step 3, the inter-frame predictor 126 determines the pixel position with the minimum cost sub-pixel precision based on the costs at four points located above, below, left, and right relative to the starting point in step 1 or step 2, and considers this pixel position as the final search position. The pixel position at the sub-pixel precision is determined by performing a weighted summation on each of the four vectors ((0,1), (0,-1), (-1,0), (1,0)) using the costs at the corresponding search positions among the four search positions as weights. The inter-frame predictor 126 then determines the vector difference between the position indicated by the initial MV and the final search position.
[0623] (Motion compensation > BIO / OBMC / LIC)
[0624] Motion compensation involves modes used to generate and correct predicted images. Examples of modes include bidirectional optical flow (BIO), overlapping block motion compensation (OBMC), local illumination compensation (LIC), etc., which will be described later.
[0625] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.
[0626] Inter-frame predictor 126 generates a predicted image (step Sm_1) and corrects the predicted image, for example, according to any of the modes described above (step Sm_2).
[0627] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.
[0628] Inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, inter-frame predictor 126 generates a predicted image using the motion vector (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, when it is determined that a correction process should be performed (Yes in step Sn_3), inter-frame predictor 126 generates the final predicted image by correcting the predicted image (step Sn_4). Note that in the LIC described later, luminance and chrominance can be corrected in step Sn_4. When it is determined that a correction process should not be performed (No in step Sn_3), inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).
[0629] (Motion compensation > OBMC)
[0630] Note that in addition to the motion information of the current block obtained through motion estimation, motion information from neighboring blocks can also be used to generate inter-frame predicted images. More specifically, an inter-frame predicted image can be generated for each sub-block in the current block by performing a weighted sum of the predicted image based on the motion information obtained through motion estimation (in the reference image) and the predicted image based on the motion information of neighboring blocks (in the current image). This inter-frame prediction (motion compensation) is also known as Overlapping Block Motion Compensation (OBMC) or OBMC mode.
[0631] In OBMC mode, information indicating the size of OBMC sub-blocks (e.g., referred to as OBMC block size) can be signaled at the sequence level. Additionally, information indicating whether OBMC mode is applied (e.g., referred to as OBMC flags) can be signaled at the CU level. Note that signaling for this type of information does not necessarily need to be performed at the sequence and CU levels, and can be performed at another level (e.g., picture level, slice level, brick level, CTU level, or sub-block level).
[0632] The OBMC model will be described in more detail. Figure 61 and Figure 62 These are flowcharts and conceptual diagrams used to illustrate an outline of the predictive image correction process performed by OBMC.
[0633] First, such as Figure 62 As shown, the predicted image (Pred) with regular motion compensation is obtained using the MV assigned to the current block. Figure 62 In the image, the arrow "MV" points to the reference image and indicates what the current block of the current image references in order to obtain the predicted image.
[0634] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) already derived for the coded block to the left of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which points to the reference image from the current block. The first correction of the predicted image is performed by overlaying the two predicted images, Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0635] Similarly, a predicted image (Pred_U) is obtained by applying the MV (MV_U) already derived for the coded block adjacent to the current block (reusing the MV of the current block) to the current block. MV (MV_U) is indicated by the arrow "MV_U", which points to the reference image from the current block. A second correction of the predicted image is performed by overlaying the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) that have already undergone the first correction. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image of the current block.
[0636] Although the example above uses a two-path correction method with the left and top adjacent blocks, note that the correction method can also be a three-path or more-path correction method that also uses the right adjacent block and / or the bottom adjacent block.
[0637] Note that the area where this overlap is performed can be only a part of the area near the block boundary, rather than the entire pixel area of the block.
[0638] Note that the prediction image correction process based on OBMC for obtaining a prediction image Pred from a reference image by overlapping and appending prediction images Pred_L and Pred_U has already been described above. However, when correcting prediction images based on multiple reference images, a similar process can be applied to each of the multiple reference images. In this case, after obtaining the corrected prediction images from the respective reference images by performing OBMC image correction based on multiple reference images, the obtained corrected prediction images are further overlapped to obtain the final prediction image.
[0639] Note that in OBMC, the current block cell can be a PU, or a sub-block cell obtained by further dividing the PU.
[0640] An example of a method for determining whether to apply OBMC is a method using the signal obmc_flag, which indicates whether OBMC is applied. As a concrete example, encoder 100 can determine whether the current block belongs to a region with complex motion. Encoder 100 sets obmc_flag to the value "1" and applies OBMC during encoding when the block belongs to a region with complex motion, and sets obmc_flag to the value "0" and encodes the block without applying OBMC when the block does not belong to a region with complex motion. Decoder 200 switches between OBMC application and non-application by decoding obmc_flag written to the stream.
[0641] (Motion compensation > BIO)
[0642] Next, the derivation method of MV is described. First, the mode for deriving MV based on the assumption of uniform linear motion is described. This mode is also called the bidirectional optical flow (BIO) mode. Furthermore, this bidirectional optical flow can be written as BDOF instead of BIO.
[0643] Figure 63 This is a conceptual diagram used to illustrate a model that assumes uniform linear motion. Figure 63 In the middle, (v x v y Let represent the velocity vector, and τ0 and τ1 represent the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1). (MV) x0 MV y0 ) represents the MV corresponding to the reference image Ref0, and (MV x1 MV y1 ) indicates the MV corresponding to the reference image Ref1.
[0644] Here, in the velocity vector (v) x ,v y Under the assumption that it exhibits uniform linear motion, (MV) x0 ,MV y0 ) and (MV x1 ,MV y1 ) are respectively represented as (v xτ0 ,v yτ0 ) and (-v xτ1 ,-v yτ1 ), and the following optical flow equation (2) is given.
[0645] [Mathematical Expression 3]
[0646]
[0647] Here, I(k) represents the motion-compensated luminance value of the reference image k (k = 0, 1) after motion compensation. The optical flow equation states that the sum of the following terms is zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, an MV candidate list can be corrected pixel by pixel.
[0648] Note that motion vectors can be derived on the decoder side using methods different from those based on models that assume uniform linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0649] Figure 64This is a flowchart illustrating an example of the inter-frame prediction process based on BIO. Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor 126 that can perform inter-frame prediction based on BIO.
[0650] like Figure 65 As shown, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolated image deducer 126b, a gradient image deducer 126c, an optical flow deducer 126d, a correction value deducer 126e, and a predicted image corrector 126f. Note that the memory 126a may be a frame memory 122.
[0651] Inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference images (Ref0, Ref1) different from the image including the current block (Cur Pic). Then, inter-frame predictor 126 uses the two motion vectors (M0, M1) to derive the predicted image for the current block (step Sy_1). Note that motion vector M0 is the motion vector (MV) corresponding to reference image Ref0. x0 MV y0 Furthermore, motion vector M1 is the motion vector (MV) corresponding to the reference image Ref1. x1 MV y1 ).
[0652] Next, the interpolation image derivator 126b uses the motion vector M0 and the reference image L0 via the reference memory 126a to derive the interpolation image I of the current block. 0 Next, the interpolation image derivator 126b uses the motion vector M1 and the reference image L1 via the reference memory 126a to derive the interpolation image I of the current block. 1 (Step Sy_2). Here, the interpolated image I 0 It is the image contained in the reference image Ref0 and is to be derived for the current block, while the interpolated image I... 1 It is the image contained in reference image Ref1 and to be derived for the current block. Interpolated image I 0 and interpolated image I 1 Each of them can be the same size as the current block. Alternatively, the interpolated image I 0 and interpolated image I 1 Each of the elements can be an image larger than the current block. Furthermore, the interpolated image I... 0 and interpolated image I 1 This can include a predicted image obtained by using motion vectors (M0, M1) and a reference image (L0, L1) and applying a motion compensation filter.
[0653] Furthermore, the gradient image deriver 126c obtains the interpolated image I 0and interpolated image I 1 Derive the gradient image of the current block (Ix) 0 , Ix 1 ,Iy 0 ,Iy 1 (Step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 ,Iy 1 The gradient image derivator 126c can derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image can indicate the amount of spatial variation of pixel values along the horizontal direction, along the vertical direction, or both.
[0654] Next, the optical flow deriver 126d uses interpolated images (I 0 I 1 ) and gradient image (Ix 0 , Ix 1 ,Iy 0 ,Iy 1 For each sub-block of the current block, derive the optical flow (vx, vy) as a velocity vector (step Sy_4). The optical flow indicates the coefficients used to correct the spatial pixel movement and can be called a local motion estimate, a corrected motion vector, or a corrected weighted vector. As an example, a sub-block can be a 4×4 pixel sub-CU. Note that the optical flow derivation can be performed on a per-pixel unit basis, rather than on a per-sub-block basis.
[0655] Next, the inter-frame predictor 126 uses optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value derivator 126e uses optical flow (vx, vy) to derive correction values for the pixel values included in the current block (step Sy_5). The predicted image corrector 126f then uses the correction values to correct the predicted image of the current block (step Sy_6). Note that the correction values can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis.
[0656] Note that the BIO process flow is not limited to Figure 64 The process is publicly disclosed. For example, it may be possible to execute only the... Figure 64 It is part of the publicly disclosed process, or different processes can be added or different processes can be used as alternatives, or processes can be executed in different processing orders, etc.
[0657] (Motion compensation > LIC)
[0658] Next, an example of a pattern used to generate a predicted image (prediction) using the Local Illumination Compensation (LIC) process is described.
[0659] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses a brightness correction process performed by a LIC. Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.
[0660] First, the inter-frame predictor 126 derives the MV from the encoded reference image and obtains the reference image corresponding to the current block (step Sz_1).
[0661] Next, the inter-frame predictor 126 extracts information indicating how the luminance values change between the current block and the reference image (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region) in the current image, as well as the luminance pixel values at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 126 uses the information indicating how the luminance values change to calculate luminance correction parameters (step Sz_3).
[0662] The inter-frame predictor 126 generates a predicted image for the current block by performing a brightness correction process in which brightness correction parameters are applied to a reference image in a reference picture specified by the MV (step Sz_4). In other words, the predicted image, which serves as a reference image in a reference picture specified by the MV, is corrected based on the brightness correction parameters. In this correction, brightness can be corrected, or chroma can be corrected, or both. In other words, chroma correction parameters can be calculated using information indicating how chroma changes, and a chroma correction process can be performed.
[0663] Notice, Figure 66A The shape of the surrounding reference area shown is an example; another shape can be used.
[0664] Furthermore, although the process of generating a predicted image from a single reference image has been described here, the same approach can be used to describe the case of generating a predicted image from multiple reference images. The predicted image can be generated after performing a brightness correction process on the reference image obtained from the reference image in the same manner as described above.
[0665] One example of a method for determining whether to apply LIC is to use a lic_flag as a signal indicating whether LIC is applied. As a specific example, encoder 100 determines whether the current block belongs to an area with brightness variations. When the block belongs to an area with brightness variations, encoder 100 sets lic_flag to the value "1" and applies LIC during encoding; when the block does not belong to an area with brightness variations, it sets lic_flag to the value "0" and performs encoding without applying LIC. Decoder 200 can decode the lic_flag written to the stream and decode the current block by switching between LIC application and non-application based on the flag value.
[0666] One example of a different method for determining whether to apply the LIC procedure is based on whether the LIC procedure has already been applied to surrounding blocks. As a specific example, when the current block has already been processed in merge mode, the inter-frame predictor 126 determines whether the surrounding blocks selected for encoding in the MV derivation in merge mode have already been encoded using LIC. The inter-frame predictor 126 performs encoding by switching between LIC application and non-application based on the result. Note that the same procedure is also applied on the decoder 200 side in this example.
[0667] The luminance correction (LIC) process has been referenced Figure 66A and Figure 66B It has been described, and will be further described below.
[0668] First, the inter-frame predictor 126 derives the MV of the reference image corresponding to the current block to be encoded from the reference image used for encoding.
[0669] Next, the inter-frame predictor 126 uses the luminance pixel values of the coded surrounding reference regions adjacent to the left and top of the current block, and the luminance values at the corresponding positions in the reference image specified by MV, to extract information indicating how the luminance values of the reference image change to the luminance values of the current image, and calculates luminance correction parameters. For example, suppose the luminance pixel value of a given pixel in the surrounding reference region of the current image is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region of the reference image is p1. The inter-frame predictor 126 calculates coefficients A and B for optimizing A×p1+B=p0 as luminance correction parameters for multiple pixels in the surrounding reference region.
[0670] Next, the inter-frame predictor 126 performs a brightness correction process using the brightness correction parameters of the reference image in the reference picture specified by MV to generate a predicted image for the current block. For example, suppose the brightness pixel value in the reference image is p2, and the brightness-corrected brightness pixel value in the predicted image is p3. The inter-frame predictor 126 generates the predicted image after undergoing the brightness correction process by calculating A×p2+B=p3 for each pixel in the reference image.
[0671] For example, a region having a defined number of pixels extracted from each of its upper and left adjacent pixels can be used as a surrounding reference region. Furthermore, the surrounding reference region is not limited to regions adjacent to the current block, but can also be regions not adjacent to the current block. Figure 66A In the example shown, the surrounding reference region in the reference image can be a region specified by another MV in the current image from the surrounding reference region in the current image. For example, the other MV can be an MV in the surrounding reference region of the current image.
[0672] Although the operations performed by encoder 100 have been described here, it should be noted that decoder 200 performs similar operations.
[0673] Note that LIC can be applied not only to luminance but also to chrominance. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any of Y, Cb, and Cr.
[0674] Furthermore, the LIC procedure can be applied on a sub-block basis. For example, correction parameters can be derived using the surrounding reference regions in the current sub-block and the surrounding reference regions in the reference sub-block of the reference image specified by the MV of the current sub-block.
[0675] (Predictive Controller)
[0676] Prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from intra-frame predictor 124) and inter-frame prediction signal (the image or signal output from inter-frame predictor 126), and outputs the selected prediction image to subtractor 104 and adder 116 as the prediction signal.
[0677] (Prediction parameter generator)
[0678] Prediction parameter generator 130 can output information related to intra-frame prediction, inter-frame prediction, and the selection of prediction images in prediction controller 128 as prediction parameters to entropy encoder 110. Entropy encoder 110 can generate a stream based on the prediction parameters input from prediction parameter generator 130 and quantized coefficients input from quantizer 108. The prediction parameters can be used in decoder 200. Decoder 200 can receive and decode the stream and perform the same process as the prediction process performed by intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128. Prediction parameters may include, for example, (i) selecting a prediction signal (e.g., MV, prediction type, or prediction mode used by intra-frame predictor 124 or inter-frame predictor 126), or (ii) optional indices, flags, or values indicating the prediction process based on the prediction process performed in each of intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128.
[0679] (Decoder)
[0680] Next, the decoder 200, which is capable of decoding the stream output from the encoder 100 described above, will be described. Figure 67 This is a block diagram illustrating the functional structure of the decoder 200 according to this embodiment. The decoder 200 is an apparatus for decoding a stream of encoded images in blocks.
[0681] like Figure 67 As shown, decoder 200 includes entropy decoder 202, inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, loop filter 212, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, prediction parameter generator 222, and segmentation determiner 224. Note that intra-frame predictor 216 and inter-frame predictor 218 are configured as part of the prediction executor.
[0682] (Decoder installation example)
[0683] Figure 68 This is a functional block diagram illustrating an example installation of decoder 200. Decoder 200 includes a processor b1 and a memory b2. For example, Figure 67 The multiple constituent elements of the decoder 200 shown are installed in Figure 68 The processor b1 and memory b2 are shown.
[0684] Processor b1 is a circuit that performs information processing and is coupled to memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 can be a processor such as a CPU. Furthermore, processor b1 can be an assembly of multiple electronic circuits. Additionally, for example, processor b1 can perform... Figure 67The roles of two or more constituent elements other than the constituent element used to store information in the multiple constituent elements of the decoder 200 shown.
[0685] Memory b2 is a dedicated or general-purpose memory used by processor b1 to store information streams that are decoded. Memory b2 can be an electronic circuit and can be connected to processor b1. Furthermore, memory b2 can be contained within processor b1. Additionally, memory b2 can be an assembly of multiple electronic circuits. Furthermore, memory b2 can be a disk, optical disk, etc., or can be represented as a storage device, recording medium, etc. Furthermore, memory b2 can be non-volatile memory or volatile memory.
[0686] For example, memory b2 can store images or streams. Additionally, memory b2 can store programs used by processor b1 to decode the streams.
[0687] Furthermore, for example, memory b2 can act as Figure 67 The decoder 200 shown in the figure has the roles of two or more constituent elements used for storing information. More specifically, memory b2 can act as... Figure 67 The roles of block memory 210 and frame memory 214 are shown. More specifically, memory b2 can store reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.).
[0688] Note that in decoder 200, not... Figure 67 All elements of the multiple constituent elements shown can be implemented, and not all processes described herein can be executed. Figure 67 Some of the constituent elements shown may be contained in another device, or some of the processes described herein may be performed by another device.
[0689] The following describes the overall flow of the process performed by decoder 200, followed by a description of each constituent element included in decoder 200. Note that some constituent elements included in decoder 200 perform the same processes as some of those performed in encoder 100, and therefore these same processes will not be described in detail again. For example, the inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, and loop filter 212 included in decoder 200 perform processes similar to those performed by inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and loop filter 120 included in decoder 200.
[0690] (Overall flow of the decoding process)
[0691] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by decoder 200.
[0692] First, the segmentation determiner 224 in decoder 200 determines the segmentation pattern (step Sp_1) for each of the multiple fixed-size blocks (e.g., 128×128 pixels) included in the image based on parameters input from entropy decoder 202. This segmentation pattern is the segmentation pattern selected by encoder 100. Decoder 200 then performs steps Sp_2 to Sp_6 for each of the multiple blocks in the segmentation pattern.
[0693] Entropy decoder 202 decodes (specifically, entropy decoding) the encoded quantized coefficients and the prediction parameters for the current block (step Sp_2).
[0694] Next, the inverse quantizer 204 performs inverse quantization on the multiple quantized coefficients, and the inverse transformer 206 performs inverse transformation on the result to recover the prediction residual (i.e., the difference block) (step Sp_3).
[0695] Next, all or some of the prediction executors, including the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220, generate the prediction signal for the current block (step Sp_4).
[0696] Next, adder 208 adds the predicted image to the predicted residual to generate the reconstructed image of the current block (also known as the decoded image block) (step Sp_5).
[0697] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).
[0698] Decoder 200 then determines whether the decoding of the entire image has ended (step Sp_7). If it is determined that the decoding has not ended (No in step Sp_7), decoder 200 repeats the process that started from step Sp_1.
[0699] Note that these steps Sp_1 through Sp_7 can be executed sequentially by decoder 200, or two or more steps can be executed in parallel. The processing order of two or more steps can be modified.
[0700] (Segmentation Determiner)
[0701] Figure 70 This is a conceptual diagram used to illustrate the relationship between the segmentation determiner 224 and other constituent elements in an embodiment. As an example, the segmentation determiner 224 may perform the following process.
[0702] For example, segmentation determiner 224 collects block information from block memory 210 or frame memory 214 and further obtains parameters from entropy decoder 202. Segmentation determiner 224 can then determine the segmentation pattern for fixed-size blocks based on the block information and parameters. Segmentation determiner 224 can then output information indicating the determined segmentation pattern to inverse transformer 206, intra-frame predictor 216, and inter-frame predictor 218. Inverse transformer 206 can perform an inverse transform of the transform coefficients based on the segmentation pattern indicated by the information from segmentation determiner 224. Intra-frame predictor 216 and inter-frame predictor 218 can generate predicted images based on the segmentation pattern indicated by the information from segmentation determiner 224.
[0703] (Entropy Decoder)
[0704] Figure 71 This is a block diagram illustrating an example of the functional configuration of the entropy decoder 202.
[0705] Entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the segmentation mode by entropy decoding the stream. For example, CABAC is used for entropy decoding. More specifically, entropy decoder 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a debinarizer 202c. Binary arithmetic decoder 202a arithmetically decodes the stream into a binary signal using context values derived by context controller 202b. Context controller 202b derives context values based on the characteristics of the syntactic elements or the surrounding state (i.e., the probability of occurrence of the binary signal) in the same manner as the context controller 110b of encoder 100. Debinarizer 202c performs debinarization to transform the binary signal output from binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. This binarization can be performed according to the binarization method described above.
[0706] Thus, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can then process the stream (see...) Figure 1 The prediction parameters included in the prediction are output to the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220. The intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220 are capable of performing the same prediction process as that performed by the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128 on the encoder 100 side.
[0707] Figure 72 This is a conceptual diagram illustrating the flow of an exemplary CABAC process in the entropy decoder 202.
[0708] First, initialization is performed in the CABAC function within the entropy decoder 202. During initialization, initialization and setting of the initial context value are performed in the binary arithmetic decoder 202a. The binary arithmetic decoder 202a and the debinarizer 202c then perform arithmetic decoding and debinarization of, for example, the encoded data of a CTU. Meanwhile, the context controller 202b updates the context value each time arithmetic decoding is performed. The context controller 202b then saves the context value for post-processing. For example, the saved context value is used to initialize the context value for the next CTU.
[0709] (Inverse quantizer)
[0710] Inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block input from entropy decoder 202. More specifically, inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. Inverse quantizer 204 then outputs the inverse-quantized transform coefficients (i.e., transform coefficients) of the current block to inverse transform 206.
[0711] Figure 73 This is a block diagram illustrating an example of the functional configuration of the inverse quantizer 204.
[0712] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage device 204d, and an inverse quantization executor 204e.
[0713] Figure 74 This is a flowchart illustrating an example of the inverse quantization process performed by inverse quantizer 204.
[0714] Inverse quantizer 204 can be based on Figure 74 The illustrated process performs an inverse quantization procedure for each CU as an example. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined that inverse quantization should be performed (yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameters for the current block from the entropy decoder 202 (step Sv_12).
[0715] Next, the predictive quantization parameter generator 204b obtains quantization parameters for a different processing unit from the quantization parameter storage device 204d (step Sv_13). Based on the obtained quantization parameters, the predictive quantization parameter generator 204b generates the predictive quantization parameters for the current block (step Sv_14).
[0716] The quantization parameter generator 204a then generates the quantization parameters for the current block based on the differential quantization parameters of the current block obtained from the entropy decoder 202 and the predicted quantization parameters of the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the differential quantization parameters of the current block obtained from the entropy decoder 202 and the predicted quantization parameters of the current block generated by the predicted quantization parameter generator 204b can be added together to generate the quantization parameters for the current block. Furthermore, the quantization parameter generator 204a stores the quantization parameters of the current block in the quantization parameter storage device 204d (step Sv_16).
[0717] Next, the inverse quantization executor 204e uses the quantization parameters generated in step Sv_15 to inverse quantize the quantized coefficients of the current block into transform coefficients (step Sv_17).
[0718] Note that differential quantization parameters can be decoded at the bit sequence level, image level, slice level, brick level, or CTU level. Additionally, the initial values for the quantization parameters can be decoded at the sequence level, image level, slice level, brick level, or CTU level. In this case, the initial values for the quantization parameters and the differential quantization parameters can be used to generate the quantization parameters.
[0719] Note that the inverse quantizer 204 may include multiple inverse quantizers, and the quantized coefficients may be inverse quantized using an inverse quantization method selected from multiple inverse quantization methods.
[0720] (Inverse Transformer)
[0721] The inverse transformer 206 recovers the prediction residual by inversely transforming the transform coefficients, which are inputs from the inverse quantizer 204.
[0722] For example, when the information parsed from the stream indicates that EMT or AMT will be applied (e.g., when the AMT flag is true), the inverse transformer 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type parsed.
[0723] Furthermore, for example, when the information parsed from the stream indicates that NSST should be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.
[0724] Figure 75 This is a flowchart illustrating an example of the process performed by the inverse converter 206.
[0725] For example, inverse transformer 206 determines whether there is information in the stream indicating that no orthogonal transformation has been performed (step St_11). Here, when it is determined that such information does not exist (no in step St_11) (e.g., there is no indication of whether an orthogonal transformation has been performed; there is an indication that an orthogonal transformation will be performed); inverse transformer 206 obtains information indicating the type of transformation decoded by entropy decoder 202 (step St_12). Next, based on this information, inverse transformer 206 determines the type of transformation used for the orthogonal transformation in encoder 100 (step St_13). Inverse transformer 206 then performs an inverse orthogonal transformation using the determined transformation type (step St_14). Figure 75 As shown, when it is determined that there is information indicating that an orthogonal transformation has not been performed (Yes in step St_11) (for example, there is no explicit instruction to perform an orthogonal transformation; there is no instruction to perform an orthogonal transformation), the orthogonal transformation is not performed.
[0726] Figure 76 This is a flowchart illustrating an example of the process performed by the inverse converter 206.
[0727] For example, inverse transformer 206 determines whether the transform size is less than or equal to a predetermined value (step Su_11). This predetermined value can be pre-determined. Here, when it is determined that the transform size is less than or equal to the predetermined value (yes in step Su_11), inverse transformer 206 obtains information from entropy decoder 202 indicating which transform type the encoder 100 used among at least one transform type included in the first transform type group (step Su_12). Note that this information is decoded by entropy decoder 202 and output to inverse transformer 206.
[0728] Based on this information, the inverse transformer 206 determines the transformation type for the orthogonal transformation in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transformation on the transformation coefficients of the current block using the determined transformation type (step Su_14). When it is determined that the transformation size is not less than or equal to a determined value (No in step Su_11), the inverse transformer 206 performs an inverse transformation on the transformation coefficients of the current block using the second transformation type group (step Su_15).
[0729] Note that, as an example, the inverse orthogonal transform of inverse transform 206 can be based on Figure 75 or Figure 76 The illustrated process is executed for each TU. Furthermore, inverse orthogonal transformations can be performed by using a defined transformation type without decoding information indicating the transformation type used for orthogonal transformations. The defined transformation type can be a predefined transformation type or a default transformation type. Specifically, the transformation type can be DST7, DCT8, etc. In inverse orthogonal transformations, the inverse transformation basis functions corresponding to the transformation type are used.
[0730] (Adder)
[0731] Adder 208 reconstructs the current block by adding the prediction residual, which is input from inverse transformer 206, to the prediction residual, which is input from prediction controller 220. In other words, it generates a reconstructed image of the current block. Adder 208 then outputs the reconstructed image of the current block to block memory 210 and loop filter 212.
[0732] (Block memory)
[0733] Block memory 210 is a storage device for storing blocks contained in the current image and referable in intra-frame prediction. More specifically, block memory 210 stores the reconstructed image output from adder 208.
[0734] (Loop filter)
[0735] The loop filter 212 applies the loop filter to the reconstructed image generated by the adder 208, outputs the filtered reconstructed image to the frame memory 214, and provides the output of the decoder 200, for example, to a display device, etc.
[0736] When the information indicating whether ALF is on or off, which is parsed from the stream, indicates that ALF is on, a filter can be selected from multiple filters, for example, based on the direction and activity of the local gradient, and the selected filter is applied to reconstruct the image.
[0737] Figure 77 This is a block diagram illustrating an example of the functional configuration of loop filter 212. Note that the configuration of loop filter 212 is similar to the configuration of loop filter 120 of encoder 100.
[0738] For example, such as Figure 77 As shown, the loop filter 212 includes a deblocking filter executor 212a, a SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filter process on the reconstructed image. The SAO executor 212b performs a SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include... Figure 77 All constituent elements disclosed herein may be included, but only a portion thereof may be included. Furthermore, the loop filter 212 may be configured to interact with... Figure 77 The above process can be executed in a different order than the one disclosed in the document, or it may be omitted. Figure 77 All the processes shown, etc.
[0739] (Frame Memory)
[0740] Frame memory 214 is, for example, a storage device for storing reference images used in inter-frame prediction, and may also be referred to as a frame buffer. More specifically, frame memory 214 stores the reconstructed image filtered by loop filter 212.
[0741] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))
[0742] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of decoder 200. Note that the prediction executor may include all or some of the following constituent elements: intra-frame predictor 216; inter-frame predictor 218; and prediction controller 220. The prediction executor includes, for example, intra-frame predictor 216 and inter-frame predictor 218.
[0743] The predictor generates a prediction image for the current block (step Sq_1). This prediction image is also called the prediction signal or prediction block. Note that the prediction signal is, for example, an intra-frame prediction signal or an inter-frame prediction signal. More specifically, the predictor generates the prediction image for the current block using a reconstructed image obtained for another block through the generation of the prediction image, the recovery of the prediction residual, and the addition of the prediction images. The predictor of decoder 200 generates the same prediction image as the predictor of encoder 100. In other words, the prediction image is generated according to a common method or a mutually corresponding method between the predictors.
[0744] The reconstructed image can be, for example, an image in a reference image, or an image of a decoded block (i.e., the other blocks mentioned above) in the current image that includes the current block. A decoded block in the current image can be, for example, a neighboring block of the current block.
[0745] Figure 79 This is a flowchart illustrating another example of the process performed by the predictor of decoder 200.
[0746] The predictor determines the method or pattern used to generate the predicted image (step Sr_1). For example, the method or pattern can be determined based on, for example, prediction parameters.
[0747] When the first method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the third method (step Sr_2c).
[0748] The first, second, and third methods can be different from each other for generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0749] Figures 80A to 80C (Collectively referred to as Figure 80) is a flowchart illustrating another example of the process performed by the predictor of the decoder 200.
[0750] The predictor can perform the prediction process according to the flow shown in Figure 80, as an example. Note that the intra-block copying shown in Figure 80 is a mode belonging to inter-frame prediction, and the block contained in the current image is called the reference image or reference block. In other words, in intra-block copying, an image different from the current image is not referenced. Additionally, the PCM mode shown in Figure 80 is a mode belonging to intra-frame prediction, in which no transformation or quantization is performed.
[0751] (Intra-frame predictor)
[0752] Intra-predictor 216 performs intra-prediction based on an intra-prediction pattern parsed from the stream, referencing blocks in the current image stored in block memory 210, to generate a predicted image of the current block (i.e., the intra-prediction block). More specifically, intra-predictor 216 performs intra-prediction by referencing pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-prediction image, which is then output to prediction controller 220.
[0753] Note that when the intra-prediction mode that references the luma block in the intra-prediction of the chroma block is selected, the intra-predictor 216 can predict the chroma component of the current block based on the luma component of the current block.
[0754] Furthermore, when the information parsed from the stream indicates that PDPC will be applied, the intra-predictor 216 corrects the intra-predicted pixel values based on the horizontal / vertical reference pixel gradients.
[0755] Figure 81 This is a diagram illustrating an example of the process performed by the intra-frame predictor 216 of the decoder 200.
[0756] The intra-frame predictor 216 first determines whether to use MPM. For example... Figure 81As shown, the intra predictor 216 determines whether the MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 exists (yes in step Sw_11), the intra predictor 216 obtains information from the entropy decoder 202 indicating the intra prediction mode selected in the encoder 100 within the MPM. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode (step Sw_14), which is included among the multiple intra prediction modes included in the MPM and indicated by the information obtained in step Sw_12.
[0757] When it is determined that the MPM flag indicating 1 does not exist (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information from the entropy decoder 202 indicating an intra prediction mode selected from at least one intra prediction mode in the encoder 100 that is not included in the MPM. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Then, the intra predictor 216 determines the intra prediction mode (step Sw_17) that is not included in the plurality of intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15.
[0758] Intra-predictor 216 generates a predicted image based on the intra-prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0759] (Inter-frame predictor)
[0760] The inter-frame predictor 218 predicts the current block by referring to a reference image stored in the frame memory 214. Prediction is performed on a unit basis: the current block or the current sub-block within the current block. Note that sub-blocks are included within blocks and are smaller than blocks. The size of a sub-block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of a sub-block can be switched in units such as slices, bricks, or images.
[0761] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from a stream (e.g., prediction parameters output from the entropy decoder 202) and outputs the inter-frame prediction image to the prediction controller 220.
[0762] When the information parsed from the stream indicates that the OBMC mode should be applied, the inter-frame predictor 218 uses motion information of neighboring blocks to generate an inter-frame predicted image, in addition to the motion information of the current block obtained through motion estimation.
[0763] Furthermore, when the information parsed from the stream indicates that the FRUC mode should be applied, the inter-frame predictor 218 derives motion information by performing motion estimation based on the mode matching method parsed from the stream (e.g., bilateral matching or template matching). The inter-frame predictor 218 then uses the derived motion information to perform motion compensation (prediction).
[0764] Furthermore, when applying BIO mode, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. Additionally, when information from stream parsing indicates that affine mode should be applied, the inter-frame predictor 218 derives the MV of each sub-block based on the MVs of multiple adjacent blocks.
[0765] (MV Derivation Process)
[0766] Figure 82 This is a flowchart illustrating an example of the MV derivation process in decoder 200.
[0767] For example, inter-frame predictor 218 determines whether to decode motion information (e.g., motion signature, video). For example, inter-frame predictor 218 may make this determination based on prediction modes included in the stream, or based on other information included in the stream. Here, when it is determined that motion information should be decoded, inter-frame predictor 218 derives the motion signature of the current block in a mode where motion information is decoded. When it is determined that motion information should not be decoded, inter-frame predictor 218 derives the motion signature in a mode where motion information is not decoded.
[0768] Here, the MV derivation modes include the regular inter-frame mode, regular merging mode, FRUC mode, affine mode, etc., which will be described later. Modes that decode motion information include the regular inter-frame mode, regular merging mode, and affine mode (specifically, affine inter-frame mode and affine merging mode). Note that motion information may include not only the MV but also the MV predictor selection information described later. Modes that do not decode motion information include the FRUC mode, etc. The inter-frame predictor 218 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0769] Figure 83 This is a flowchart illustrating an example of the MV derivation process in decoder 200.
[0770] For example, the inter-frame predictor 218 can determine whether to decode the MV difference, i.e., based on a prediction mode contained in the stream, or based on other information contained in the stream. Here, when it is determined that the MV difference should be decoded, the inter-frame predictor 218 can derive the MV of the current block in the mode in which the MV difference is decoded. In this case, for example, the MV difference contained in the stream is decoded into prediction parameters.
[0771] When it is determined that no MV difference will be decoded, the inter-frame predictor 218 derives the MV in a mode where no MV difference is decoded. In this case, the encoded MV difference is not included in the stream.
[0772] Here, as described above, the MV derivation modes include the regular inter-frame mode, regular merging mode, FRUC mode, affine mode, etc., as described later. Modes that encode the MV difference include the regular inter-frame mode and the affine mode (specifically, the affine inter-frame mode). Modes that do not encode the MV difference include the FRUC mode, the regular merging mode, and the affine mode (specifically, the affine merging mode). The inter-frame predictor 218 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0773] (MV Derivation > Regular Inter-Frame Mode)
[0774] For example, when the information parsed from the stream indicates that a regular inter-frame mode should be applied, the inter-frame predictor 218 derives the MV based on the information parsed from the stream and uses the MV to perform motion compensation (prediction).
[0775] Figure 84 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using a regular inter-frame mode.
[0776] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks around the current block in time or space (step Sg_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0777] Next, the inter-frame predictor 218 extracts N (2 or larger integers) MV candidates from the multiple MV candidates obtained in step Sg_11 as motion vector predictor candidates (also called MV predictor candidates) according to the determined ranking in the priority order (step Sg_12). Note that the ranking in the priority order can be determined in advance for the corresponding N MV predictor candidates, and the ranking can be predetermined.
[0778] Next, the inter-frame predictor 218 decodes the MV predictor selection information from the input stream and uses the decoded MV predictor selection information to select one MV predictor candidate from N MV predictor candidates as the MV predictor for the current block (step Sg_13).
[0779] Next, the inter-frame predictor 218 decodes the MV difference from the input stream and derives the MV of the current block by adding the difference, which is the decoded MV difference, to the selected MV predictor (step Sg_14).
[0780] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference image (step Sg_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed on each block in all blocks of a slice, the inter-frame prediction for the slice using the regular inter-frame mode ends. Similarly, when the processes in steps Sg_11 to Sg_15 are performed on each block in all blocks of an image, the inter-frame prediction for the image using the regular inter-frame mode ends. Note that not all blocks included in a slice undergo the processes in steps Sg_11 to Sg_15, and the inter-frame prediction for the slice using the regular inter-frame mode can end when only some blocks undergo the process. This also applies to the image in steps Sg_11 to Sg_15. When the process is performed on only some blocks in an image, the inter-frame prediction for the image using the regular inter-frame mode can end.
[0781] (MV derivation > Standard merge mode)
[0782] For example, when the information parsed from the stream indicates that a regular merging mode should be applied, the inter-frame predictor 218 derives the MV and uses the MV to perform motion compensation (prediction).
[0783] Figure 85 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using a regular merging mode.
[0784] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks around the current block in time or space (step Sh_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0785] Next, the inter-frame predictor 218 selects an MV candidate from the multiple MV candidates obtained in step Sh_11 and derives the MV of the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains MV selection information included in the stream as prediction parameters, and selects the MV candidate identified by the MV selection information as the MV of the current block.
[0786] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference image (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed on each block in all blocks of a slice, the inter-frame prediction of the slice using the regular merging mode ends. Similarly, when the processes in steps Sh_11 to Sh_13 are performed on each block in all blocks of an image, the inter-frame prediction of the image using the regular merging mode ends. Note that not all blocks included in a slice undergo the processes in steps Sh_11 to Sh_13, and the inter-frame prediction of the slice using the regular merging mode can end when only some blocks undergo the process. This also applies to the images in steps Sh_11 to Sh_13. When the process is performed on only some blocks in an image, the inter-frame prediction of the image using the regular merging mode can end.
[0787] (MV Derivation > FRUC Pattern)
[0788] For example, when information parsed from the stream indicates that FRUC mode will be applied, the inter-frame predictor 218 derives the motion MV in FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is derived on the decoder 200 side, without being signaled from the encoder 100 side. For example, the decoder 200 can derive motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values from the current block.
[0789] Figure 86 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 via FRUC mode.
[0790] First, the inter-frame predictor 218 generates a list of MVs indicating spatially or temporally adjacent decoded blocks as MV candidates by referencing the MV (this list is an MV candidate list and can also be used, for example, as an MV candidate list for the regular merging mode (step Si_11). Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates an evaluation value for each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV of the current block ( Step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV of the current block. Alternatively, for example, the MV of the current block is derived using pattern matching in the surrounding region of the location in the reference image that corresponds to the selected best MV candidate. In other words, estimation using pattern matching and evaluation values in the reference image can be performed in the surrounding region of the best MV candidate, and when an MV that produces a better evaluation value exists, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV of the current block. In an embodiment, updating the MV that produces a better evaluation value may not be performed.
[0791] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference image (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of a slice, the inter-frame prediction of the slice using FRUC mode ends. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of an image, the inter-frame prediction of the image using FRUC mode ends. Each sub-block can be processed similarly to the case of each block.
[0792] (MV Derivation > FRUC Pattern)
[0793] For example, when the information parsed from the stream indicates that an affine merge mode should be applied, the inter-frame predictor 218 derives the MV in the affine merge mode and uses the MV to perform motion compensation (prediction).
[0794] Figure 87 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using affine merging mode.
[0795] In affine merging mode, firstly, the inter-frame predictor 218 derives the MV at the corresponding control point of the current block (step Sk_11). For example... Figure 46AAs shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 46B As shown, these are the top left corner, top right corner, and bottom left corner of the current block.
[0796] For example, when using Figures 47A to 47C When using the MV derivation method shown, such as Figure 47A As shown, the inter-frame predictor 218 checks the decoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in this order, and identifies the first valid block decoded according to the affine pattern. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47B As shown, the inter-frame predictor 218 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the decoded block including block A. In this way, the MV at each control point is derived.
[0797] Note that, as Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, and as follows: Figure 49B As shown, when block A is identified and when block A has three control points, the MV at two control points can be calculated.
[0798] Furthermore, when MV selection information is included in the stream as a prediction parameter, the inter-frame predictor 218 can use the MV selection information to derive the MV at each control point of the current block.
[0799] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and the decoded reference image to perform motion compensation for the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merging mode of the current block ends. In other words, motion compensation for the current block is performed to generate the predicted image of the current block.
[0800] Note that the MV candidate list described above can be generated in step Sk_11. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown Figure 49A and Figure 49B The MV derivation method shown is any combination of other MV derivation methods.
[0801] Note that, in addition to affine patterns, the MV candidate list can include MV candidates from patterns that perform prediction on a sub-block basis.
[0802] Note that, for example, an MV candidate list can be generated that includes MV candidates from both the affine merge pattern using two control points and the affine merge pattern using three control points. Alternatively, an MV candidate list can be generated separately that includes MV candidates from the affine merge pattern using two control points and that includes MV candidates from the affine merge pattern using three control points. Alternatively, an MV candidate list can be generated that includes MV candidates from either the affine merge pattern using two control points or the affine merge pattern using three control points.
[0803] (MV Derivation > Affine Inter-Frame Mode)
[0804] For example, when the information parsed from the stream indicates that an affine inter-frame mode will be applied, the inter-frame predictor 218 derives the MV in the affine inter-frame mode and uses the MV to perform motion compensation (prediction).
[0805] Figure 88 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using affine inter-frame modes.
[0806] In affine inter-frame mode, firstly, the inter-frame predictor 218 derives the MV predictor (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_11). The control points are the top-left corner, the top-right corner, and the bottom-left corner of the current block, such as... Figure 46A or Figure 46B As shown.
[0807] Inter-frame predictor 218 obtains MV predictor selection information included in the stream as prediction parameters, and uses the MV identified by the MV predictor selection information to derive the MV predictor at each control point of the current block. For example, when using Figure 48A and Figure 48B When the MV derivation method is shown, the inter-frame predictor 218 selects... Figure 48A or Figure 48BThe MV predictor in the coded block near the corresponding control point of the current block is selected by the MV of the block identified by the MV predictor selection information to deduce the motion vector predictor (v0, v1) or (v0, v1, v2) at the control point of the current block.
[0808] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block to the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point of the current block is derived.
[0809] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sj_13). The inter-frame predictor 218 then uses these affine MVs and the decoded reference image to perform motion compensation for the sub-blocks (step Sj_14). When the processes in steps Sj_13 and Sj_14 are performed for each sub-block included in the current block, the inter-frame prediction using the affine merging mode of the current block ends. In other words, motion compensation for the current block is performed to generate the predicted image of the current block.
[0810] Note that the above MV candidate list can be generated in step Sj_11 as in step Sk_11.
[0811] (MV Derivation > Triangle Pattern)
[0812] For example, when the information parsed from the stream indicates that a triangle pattern will be applied, the inter-frame predictor 218 derives the MV in the triangle pattern and uses the MV to perform motion compensation (prediction).
[0813] Figure 89 This is a flowchart illustrating an example of the process of inter-frame prediction using a triangle pattern in decoder 200.
[0814] In the triangle mode, firstly, the inter-frame predictor 218 divides the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as prediction parameters; this partition information is information related to the partitioning. The inter-frame predictor 218 can then divide the current block into a first partition and a second partition based on the partition information.
[0815] Next, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks around the current block in time or space (step Sx_12). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0816] Inter-frame predictor 218 then selects the MV candidate of the first partition and the MV candidate of the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, inter-frame predictor 218 can obtain MV selection information from the stream for identifying each selected MV candidate as a prediction parameter. Inter-frame predictor 218 can then select the first MV and the second MV based on the MV selection information.
[0817] Next, the inter-frame predictor 218 performs motion compensation using the selected first MV and the decoded reference image to generate a first predicted image (step Sx_14). Similarly, the inter-frame predictor 218 performs motion compensation using the selected second MV and the decoded reference image to generate a second predicted image (step Sx_15).
[0818] Finally, the inter-frame predictor 218 generates the prediction image for the current block by performing a weighted summation of the first and second prediction images (step Sx_16).
[0819] (MV estimate > DMVR)
[0820] For example, information from stream parsing indicates that DMVR will be applied, and inter-frame predictor 218 uses DMVR to perform motion estimation.
[0821] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed by the DMVR in the decoder 200.
[0822] Inter-frame predictor 218 derives the MV of the current block based on the merging pattern (step S1_11). Next, inter-frame predictor 218 derives the final MV of the current block by searching the region around the reference image indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined based on the DMVR.
[0823] Figure 91 This is a flowchart illustrating an example of the motion estimation process performed by the DMVR in the decoder 200, and... Figure 58B same.
[0824] First of all, Figure 58AIn step 1 shown, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also called the starting point) and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at one of the search positions other than the starting point is determined to be minimized, the inter-frame predictor 218 changes its objective to obtain the search position with the minimum cost and executes the process in step 2 shown in Figure 58. When the cost at the starting point is minimized, the inter-frame predictor 218 skips... Figure 58A The process in step 2 is shown in the figure, and the process in step 3 is executed.
[0825] In such Figure 58A In step 2, as shown, the inter-frame predictor 218 performs a search similar to that in step 1, treating the search position after the target change as the new starting point based on the result of the process in step 1. Then, the inter-frame predictor 218 determines whether the cost is minimized at each search position other than the starting point. Here, when the cost is minimized at one of the search positions other than the starting point, the inter-frame predictor 218 performs the process in step 4. When the cost is minimized at the starting point, the inter-frame predictor 218 performs the process in step 3.
[0826] In step 4, the inter-frame predictor 218 treats the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0827] exist Figure 58A In step 3 shown, the inter-frame predictor 218 determines the pixel position with the minimum cost subpixel precision based on the cost at four points located at the top, bottom, left, and right positions relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position.
[0828] The pixel position at sub-pixel precision is determined by weighting each of the four vectors ((0,1), (0,-1), (-1,0), (1,0)) in the top, bottom, left, and right directions using the cost at the corresponding search position among the four search positions. The inter-frame predictor 218 then determines the vector difference as the difference between the position indicated by the initial MV and the final search position.
[0829] (Motion compensation > BIO / OBMC / LIC)
[0830] For example, when the information parsed from the stream indicates that correction of the predicted image should be performed, the inter-frame predictor 218 corrects the predicted image based on a correction mode when the predicted image is generated. This mode is, for example, one of the BIO, OBMC, and LIC mentioned above.
[0831] Figure 92This is a flowchart illustrating an example of the process of generating a predicted image in decoder 200.
[0832] Inter-frame predictor 218 generates a predicted image (step Sm_11) and corrects the predicted image according to any of the modes described above (step Sm_12).
[0833] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in decoder 200.
[0834] Inter-frame predictor 218 derives the MV of the current block (step Sn_11). Next, inter-frame predictor 218 generates a predicted image using the MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, inter-frame predictor 218 obtains prediction parameters contained in the stream and determines whether to perform a correction process based on these prediction parameters. For example, these prediction parameters are flags indicating whether to apply one or more of the above-described modes. Here, when it is determined that a correction process should be performed (Yes in step Sn_13), inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). Note that in LIC, luminance and chrominance can be corrected in step Sn_14. When it is determined that a correction process should not be performed (No in step Sn_13), inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).
[0835] (Motion compensation > OBMC)
[0836] For example, when the information parsed from the stream indicates that OBMC should be performed, the inter-frame predictor 218 corrects the predicted image based on OBMC when the predicted image is generated.
[0837] Figure 94 This is a flowchart illustrating an example of the process by which the OBMC corrects the predicted image in decoder 200. Note that... Figure 94 The flowchart in the diagram represents the use of Figure 62 The correction process for the predicted image of the current image and the reference image is shown.
[0838] First, such as Figure 62 As shown, the predicted image (Pred) is obtained using the MV assigned to the current block through regular motion compensation.
[0839] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) already derived for the decoded block to the left of the current block (reusing the motion vector of the current block). Then, the inter-frame predictor 218 performs a first correction of the predicted image by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0840] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) already derived for the decoded block adjacent to the current block (reusing motion vectors for the current block). Then, the inter-frame predictor 218 performs a second correction to the predicted image by overlaying the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) that have already undergone the first correction. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.
[0841] (Motion compensation > BIO)
[0842] For example, when the information parsed from the stream indicates that a BIO should be performed, the inter-frame predictor 218 corrects the predicted image based on the BIO when the predicted image is generated.
[0843] Figure 95 This is a flowchart illustrating an example of the correction process for the predicted image performed by the BIO in the decoder 200.
[0844] like Figure 63 As shown, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference images (Ref0, Ref1) that are different from the image (Cur Pic) including the current block. Then, the inter-frame predictor 218 uses the two motion vectors (M0, M1) to derive the predicted image for the current block (step Sy_11). Note that motion vector M0 is the motion vector (MV) corresponding to the reference image Ref0. x0 MV y0 The motion vector M1 is the motion vector (MV) corresponding to the reference image Ref1. x1 MV y1 ).
[0845] Next, the inter-frame predictor 218 uses the motion vector M0 and the reference image L0 to derive the interpolated image I for the current block. 0 Additionally, the inter-frame predictor 218 uses motion vector M1 and reference image L1 to derive the interpolated image I for the current block. 1 (Step Sy_12). Here, the interpolated image I 0 It is the image contained in the reference image Ref0 and is to be derived for the current block, while the interpolated image I... 1 It is an image contained in reference image Ref1 and derived for the current block. Interpolated image I 0 and interpolated image I 1 Each of them can be the same size as the current block. Alternatively, the interpolated image I0 and interpolated image I 1 Each of the elements can be an image larger than the current block. Furthermore, the interpolated image I... 0 and interpolated image I 1 This can include a predicted image obtained by using motion vectors (M0, M1) and a reference image (L0, L1) and applying a motion compensation filter.
[0846] In addition, the inter-frame predictor 218 obtains data from the interpolated image I. 0 and interpolated image I 1 Derive the gradient image of the current block (Ix) 0 , Ix 1 ,Iy 0 ,Iy 1 (Step Sy_13). Note that the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 ,Iy 1 The inter-frame predictor 218 can derive a gradient image, for example, by applying a gradient filter to the interpolated image. The gradient image can be an image in which each pixel indicates the spatial variation of the pixel value along the horizontal direction or the spatial variation of the pixel value along the vertical direction.
[0847] Next, the inter-frame predictor 218 uses interpolated images (I 0 I 1 ) and gradient image (Ix 0 , Ix 1 ,Iy 0 ,Iy 1 Derive the optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_14). As an example, a sub-block can be a 4×4 pixel sub-CU.
[0848] Next, the inter-frame predictor 218 uses optical flow (vx, vy) to correct the predicted image of the current block. For example, the inter-frame predictor 218 uses optical flow (vx, vy) to derive correction values for the pixel values included in the current block (step Sy_15). The inter-frame predictor 218 then uses the correction values to correct the predicted image of the current block (step Sy_16). Note that the correction values can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis, etc.
[0849] Note that the BIO process flow is not limited to Figure 95 The process is publicly disclosed. It can be executed only. Figure 95 It is part of the publicly disclosed process, or different processes can be added or different processes can be used as alternatives, or the process can be executed in a different processing order.
[0850] (Motion compensation > LIC)
[0851] For example, when the information parsed from the stream indicates that LIC should be performed, the inter-frame predictor 218 corrects the predicted image according to the LIC when the predicted image is generated.
[0852] Figure 96 This is a flowchart illustrating an example of the process of correcting the predicted image performed by the LIC in the decoder 200.
[0853] First, the inter-frame predictor 218 uses MV to obtain the reference image corresponding to the current block from the decoded reference image (step Sz_11).
[0854] Next, the inter-frame predictor 218 extracts information indicating how the luminance values change between the current image and the reference image for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values of the decoded left adjacent reference region (surrounding reference region) and the decoded upper adjacent reference region (surrounding reference region), as well as the luminance pixel values at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 218 uses the information indicating how the luminance values change to calculate the luminance correction parameters (step Sz_13).
[0855] The inter-frame predictor 218 generates a predicted image for the current block by performing a brightness correction process, in which brightness correction parameters are applied to a reference image in the reference picture specified by the MV (step Sz_14). In other words, the predicted image, which serves as a reference image in the reference picture specified by the MV, is corrected based on the brightness correction parameters. In this correction, either brightness or chromaticity can be corrected.
[0856] (Predictive Controller)
[0857] Prediction controller 220 selects an intra-frame predicted image or an inter-frame predicted image and outputs the selected image to adder 208. In general, the configuration, function, and procedure of prediction controller 220, intra-frame predictor 216, and inter-frame predictor 218 on the decoder 200 side can correspond to the configuration, function, and procedure of prediction controller 128, intra-frame predictor 124, and inter-frame predictor 126 on the encoder 100 side.
[0858] (First aspect)
[0859] Figure 97 This is a flowchart illustrating an example of a process flow 1000 for decoding an image using the CCALF (Cross-Component Adaptive Loop Filtering) procedure according to the first aspect. Process flow 1000 can be, for example, derived from... Figure 67 The decoder 200 and others are executed.
[0860] In step S1001, a filtering process is applied to the reconstructed image samples of the first component. For example, the first component may be a luminance component. The luminance component may be represented as a Y component. The reconstructed image samples of the luminance may be the output signal of an ALF process. The output signal of the ALF may be a reconstructed luminance sample generated by a SAO process. In some embodiments, the filtering process performed in step S1001 may be represented as a CCALF process. The number of reconstructed luminance samples may be the same as the number of coefficients of the filter to be used in the CCALF process. In other embodiments, a cropping process may be performed on the filtered reconstructed luminance samples.
[0861] In step S1002, the reconstructed image sample of the second component is modified. The second component may be a chroma component. The chroma component may be represented as Cb and / or Cr components. The reconstructed image sample of the chroma may be the output signal of the ALF process. The output signal of the ALF may be the reconstructed chroma sample generated by the SAO process. The modified reconstructed image sample may be the sum of the reconstructed chroma sample and the filtered reconstructed luminance sample, i.e., the output of step S1001. In other words, the modification process can be performed by adding the filtered value of the reconstructed luminance sample generated by the CCALF process in step S1001 to the filtered value of the reconstructed chroma sample generated by the ALF process. In some embodiments, a cropping process may be performed on the reconstructed chroma sample. The first component and the second component may belong to the same block or may belong to different blocks.
[0862] In step S1003, the values of the reconstructed image samples after modification of the chroma components are cropped. By performing the cropping process, it can be ensured that the sample values are within a certain range. Furthermore, cropping can promote better convergence in processes such as least squares optimization, minimizing the difference between the residual (the difference between the original sample value and the reconstructed sample value) and the filtered value of the chroma sample, so as to facilitate the determination of the filter coefficients.
[0863] In step S1004, the image is decoded using cropped reconstructed image samples of the chroma components. In some embodiments, step S1003 is not required. In this case, the image is decoded using uncropped modified reconstructed chroma samples.
[0864] Figure 98 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment. In this embodiment, a cropping process is applied to the reconstructed image samples with modified chroma components, such as... Figure 97In step S1003. For example, for a 10-bit output, the modified reconstructed image sample can be cropped to the range [0, 1023]. In some embodiments, when cropping the filtered reconstructed image sample of the luminance component generated by the CCALF process, cropping the modified reconstructed image sample of the chrominance component may not be necessary.
[0865] Figure 99 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment. In this embodiment, as... Figure 97 In step S1003, cropping is applied to the modified reconstructed image samples of the chroma components. The cropping process does not apply to the filtered reconstructed luminance samples generated by the CCALF process. The filtered values of the reconstructed chroma samples generated by the ALF process do not need to be cropped, such as... Figure 99 As shown in the “No Crop” section, in other words, the reconstructed image samples to be modified are generated using filtered values (ALF chroma) and interpolation (CCALF Cb / Cr), where no cropping is applied to the output of the generated sample values.
[0866] Figure 100 This is a block diagram illustrating the functional configuration of the encoder and decoder according to an embodiment. In this embodiment, a cropping process is applied to the filtered reconstructed luminance sample (“cropped output sample”) and the modified reconstructed image sample of the chrominance component (“summed cropping”) generated by the CCALF process. The filtered values of the reconstructed chrominance sample generated by the ALF process are not cropped (“uncropped”). For example, the cropping range applied to the filtered reconstructed image sample of the luminance component can be [-2^15, 2^15-1] or [-2^7, 2^7-1].
[0867] Figure 101 Another example is shown where the clipping process is applied to: a filtered reconstructed luminance sample (“clipped output sample”) generated by the CCALF process, a modified reconstructed image sample of the chrominance components (“sum-clipped”), and a filtered reconstructed chrominance sample (“clipped”) generated by the ALF process. In other words, the output values of the CCALF and ALF chrominance processes are clipped separately and then clipped again after they are summed. In this embodiment, the modified reconstructed image sample of the chrominance components does not need to be clipped. For example, the final output of the ALF chrominance process might be clipped to a 10-bit value. For example, the clipping range applied to the filtered reconstructed image sample of the luminance components could be [-2^15, 2^15-1] or [-2^7, 2^7-1]. This range can be fixed or can be adaptively determined. In either case, the range can be signaled in the header information, such as in the SPS (Sequence Parameter Set) or APS (Adaptive Parameter Set). In the case of using a nonlinear ALF, it can be... Figure 101 The "Sum and Clip" function defines the clipping parameters.
[0868] The reconstructed image sample for the luminance component to be filtered by the CCALF process can be a neighboring sample of the currently reconstructed image sample for the chrominance component. That is, the modified current reconstructed image sample can be generated by adding the filtered value of the neighboring image sample of the luminance component located adjacent to the current image sample to the filtered value of the current image sample for the chrominance component. The filtered value of the image sample for the luminance component can be represented as a difference.
[0869] The process disclosed in this aspect can reduce the size of the internal hardware memory required to store filtered image sample values.
[0870] (Second aspect)
[0871] Figure 102 This is a flowchart of an example of a process flow 2000 for decoding an image using the CCALF procedure based on the information defined in the second aspect. Process flow 2000 can be, for example, derived from... Figure 67 The decoder 200 and others are executed.
[0872] In step S2001, trimming parameters are parsed from the bitstream. Trimming parameters can be parsed from VPS, APS, SPS, PPS, and slice headers at the CTU or TU level, such as... Figure 103 As described in [the text]. Figure 103 It is a conceptual diagram representing the location of the clipping parameters. Figure 103 The parameters described can be replaced by different types of trimming parameters, flags, or indices. Two or more trimming parameters can be parsed from two or more parameter sets in the bitstream.
[0873] In step S2002, the cropping parameter, cropping difference, is used. Based on the reconstructed image samples of the first component (e.g., Figure 98-101 The difference (CCALF Cb / Cr) is used to generate the difference. For example, the first component is the luminance component, and the difference is a filtered reconstructed luminance sample generated by the CCALF process. In this case, a clipping process is applied to the filtered reconstructed luminance sample using analytical clipping parameters.
[0874] The clipping parameter limits the value to the desired range. If the desired range is [-3, 3], for example, the operation clip(-3, 3, 5) clips the value 5 to 3. In this example, the value -3 is the lower bound, and the value 3 is the upper bound.
[0875] The clipping parameter can indicate the indices used to derive the lower and upper bounds, such as... Figure 104As shown in (i). In this example, ccalf_luma_clip_idx[] is the index, -range_array[] is the lower bound, and range_array[] is the upper bound. In this example, range_array[] is a defined range array, which may differ from the range array used for ALF. The defined range array can be predetermined.
[0876] Clipping parameters can indicate lower and upper limits, such as Figure 104 As shown in (ii). In this example, -ccalf_luma_clip_low_range[] is the lower bound range, while ccalf_luma_clip_up_range[] is the upper bound range.
[0877] Clipping parameters can indicate the common range of both the lower and upper limits, such as... Figure 104 As shown in (iii). In this example, -ccalf_luma_clip_range is the lower bound, while ccalf_luma_clip_range is the upper bound.
[0878] The difference is generated by multiplying, dividing, adding, or subtracting at least two reconstructed image samples of the first component. For example, the two reconstructed image samples can come from the current and neighboring image samples or from two neighboring image samples. The positions of the current and neighboring image samples can be predetermined.
[0879] In step S2003, the reconstructed image sample of the second component, which is different from the first component, is modified using the clipped value. The clipped value can be the clipped value of the reconstructed image sample of the luminance component. The second component can be the chrominance component. This modification can include operations such as multiplying, dividing, adding, or subtracting the clipped value relative to the reconstructed image sample of the second component.
[0880] In step S2004, the image is decoded using the modified reconstructed image sample.
[0881] In this disclosure, one or more pruning parameters for the cross-component adaptive loop filter are signaled in the bitstream. This signaling allows for the combination of the syntax of the cross-component adaptive loop filter and the syntax of the adaptive loop filter to achieve syntactic simplification. Furthermore, this signaling allows for more flexible design of the cross-component adaptive loop filter to improve coding efficiency.
[0882] Clipping parameters can be defined or predefined for both the encoder and decoder without signaling. Clipping parameters can also be derived using luminance information without signaling. For example, if a strong gradient or edge is detected in the luminance-reconstructed image, a clipping parameter corresponding to a large clipping range can be derived, and if a weak gradient or edge is detected in the luminance-reconstructed image, a clipping parameter corresponding to a short clipping range can be derived.
[0883] (Third aspect)
[0884] Figure 105 This is a flowchart illustrating an example of process flow 3000 for decoding an image using the CCALF procedure with filter coefficients based on the third aspect. Process flow 3000 can be, for example, derived from... Figure 67 The decoder 200, etc., is executed. Filter coefficients are used in the filtering step of the CCALF process to generate filtered reconstructed image samples of the luminance component.
[0885] In step S3001, it is determined whether the filter coefficients are located within a defined symmetrical region of the filter. Optionally, an additional step of determining whether the shape of the filter coefficients is symmetrical can be performed. Information indicating whether the samples of the filter coefficients are symmetrical can be encoded into a bitstream. If the shape is symmetrical, the positions of the coefficients within the symmetrical region can be determined or predetermined.
[0886] In step S3002, if the filter coefficients are within the defined symmetrical region ("Yes" in step S3001), the filter coefficients are copied to the symmetrical position and a set of filter coefficients is generated.
[0887] In step S3003, the reconstructed image samples of the first component are filtered using filter coefficients. The first component may be the luminance component.
[0888] In step S3004, the filtered output is used to modify a reconstructed image sample that is different from the first component. The second component may be a chroma component.
[0889] In step S3005, the image is decoded using the modified reconstructed image sample.
[0890] If the filter coefficients are asymmetric (No in step S3001), all filter coefficients can be encoded from the bitstream, and a set of filter coefficients can be generated without copying.
[0891] This can reduce the amount of information that needs to be encoded into the bitstream. That is, only one of the symmetric filter coefficients may need to be encoded into the bitstream.
[0892] Figure 106 , Figure...
Claims
1. A coding method, comprising: generating first coefficient values by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; Clipping the first coefficient value; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the clipped first coefficient value to the clipped second coefficient value; trimming the third coefficient value; as well as The third reconstructed image samples of the chrominance component are encoded using the clipped third coefficient values.
2. A decoding method, comprising: generating first coefficient values by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; Clipping the first coefficient value; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the clipped first coefficient value to the clipped second coefficient value; trimming the third coefficient value; as well as The third reconstructed image samples of the chrominance component are decoded using the clipped third coefficient values.
3. A non-transitory computer-readable medium storing a bitstream, characterized in that: The bitstream includes filter information, the filter information causing the decoder to perform a filtering process, the filtering process comprising: generating first coefficient values by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; Clipping the first coefficient value; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the clipped first coefficient value to the clipped second coefficient value; clipping the third coefficient value; and The third reconstructed image samples of the chrominance component are decoded using the clipped third coefficient values.