Systems and methods for video coding

By applying the cross component adaptive loop filtering (CCALF) process in video encoding technology, the reconstructed image samples of brightness and chrominance components are encoded and decoded, which solves the problems of low encoding efficiency and poor image quality in the prior art, and realizes a more efficient encoding and decoding process.

CN120017829APending Publication Date: 2025-05-16PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332447.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-08
Filing Date
2020-08-07
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing video encoding technologies process high-resolution and large amounts of digital video data, they have low encoding efficiency, poor image quality, and large circuit scale, making it difficult to meet the needs of various applications.

Method used

Cross-component adaptive loop filtering (CCALF) process is used to encode and decode the reconstructed image samples of brightness and chrominance components, and coefficient values ​​are generated through clipping and addition operations, and images are used for encoding and decoding.

Benefits of technology

It improves encoding efficiency, enhances image quality, reduces circuit scale, and improves encoding and decoding processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017829A_ABST
    Figure CN120017829A_ABST
Patent Text Reader

Abstract

The encoder includes a circuit and a memory coupled to the circuit. The circuitry, in operation, generates a first coefficient value by applying a CCALF (cross component adaptive loop filtering) process to a first reconstructed image sample of a luma component, generates a second coefficient value by applying an ALF (adaptive loop filtering) process to a second reconstructed image sample of a chroma component, and crosses the second coefficient value. The circuit generates a third coefficient value by adding the first coefficient value and the clipped second coefficient value, and clips the third coefficient value. The circuit encodes a third reconstructed image sample of the chroma component using the clipped third coefficient value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent with an application date of August 7, 2020, application number 202080051926.7, and invention name “System and method for video encoding”. Technical Field

[0002] The present disclosure relates to video coding, and more particularly to video encoding and decoding systems, components and methods in video encoding and decoding, such as for performing a CCALF (cross-component adaptive loop filtering) process. Background Art

[0003] With the advancement of video coding technology, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec), there is still a need to continuously improve and optimize video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in the CCALF (Cross-Component Adaptive Loop Filtering) process. Summary of the Invention

[0004] According to one aspect, an encoder is provided that includes circuitry and a memory coupled to the circuitry. The circuitry, in operation, generates a first coefficient value by applying a CCALF (cross-component adaptive loop filtering) process to a first reconstructed image sample of a luma component. The circuitry generates a second coefficient value by applying an ALF (adaptive loop filtering) process to a second reconstructed image sample of a chroma component, and clips the second coefficient value. The circuitry generates a third coefficient value by adding the first coefficient value to the clipped second coefficient value, and clips the third coefficient value. The circuitry encodes a third reconstructed image sample of the chroma component using the clipped third coefficient value.

[0005] According to another aspect, the first reconstructed image samples are located adjacent to the second reconstructed image samples.

[0006] According to another aspect, the circuit is operable to set the first coefficient value to zero in response to the first coefficient value being less than 64.

[0007] According to another aspect, an encoder is provided, comprising: a block partitioner that partitions a first image into a plurality of blocks in operation; an intra-frame predictor that predicts a block included in the first image using a reference block included in the first image in operation; an inter-frame predictor that predicts a block included in the first image using a reference block included in a second image different from the first image in operation; a loop filter that filters the block included in the first image in operation; a transformer that transforms a prediction error between an original signal and a prediction signal generated by the intra-frame predictor or the inter-frame predictor in operation to generate a transform coefficient; a quantizer that quantizes the transform coefficient in operation to generate a quantized coefficient; and an entropy encoder that variably encodes the quantized coefficient in operation to generate a coded bitstream including the encoded quantized coefficient and control information. The loop filter performs the following operations:

[0008] generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luma component;

[0009] generating second coefficient values ​​by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma component;

[0010] clipping the second coefficient value;

[0011] generating a third coefficient value by adding the first coefficient value to the clipped second coefficient value;

[0012] clipping the third coefficient value; and

[0013] The third reconstructed image samples of the chroma component are encoded using the clipped third coefficient values.

[0014] According to another aspect, a decoder is provided that includes circuitry and a memory coupled to the circuitry. The circuitry, in operation, generates a first coefficient value by applying a CCALF (cross-component adaptive loop filtering) process to a first reconstructed image sample of a luma component. The circuitry generates a second coefficient value by applying an ALF (adaptive loop filtering) process to a second reconstructed image sample of a chroma component, and clips the second coefficient value. The circuitry generates a third coefficient value by adding the first coefficient value to the clipped second coefficient value, and clips the third coefficient value. The circuitry decodes a third reconstructed image sample of the chroma component using the clipped third coefficient value.

[0015] According to another aspect, a decoding apparatus is provided, comprising: a decoder that decodes a coded bitstream to output quantized coefficients; an inverse quantizer that inversely quantizes the quantized coefficients to output transform coefficients; an inverse transformer that inversely transforms the transform coefficients to output prediction errors; an intra-frame predictor that uses a reference block included in a first image to predict a block included in the first image; an inter-frame predictor that uses a reference block included in a second image different from the first image to predict a block included in the first image; a loop filter that filters the block included in the first image; and an output terminal that outputs a picture including the first image. The loop filter performs the following operations:

[0016] generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luma component;

[0017] generating second coefficient values ​​by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma component;

[0018] clipping the second coefficient value;

[0019] generating a third coefficient value by adding the first coefficient value to the clipped second coefficient value;

[0020] clipping the third coefficient value; and

[0021] The third reconstructed image samples of the chroma component are decoded using the clipped third coefficient values.

[0022] According to another aspect, there is provided a coding method comprising:

[0023] generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luma component;

[0024] generating second coefficient values ​​by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma component;

[0025] clipping the second coefficient value;

[0026] generating a third coefficient value by adding the first coefficient value to the clipped second coefficient value;

[0027] clipping the third coefficient value; and

[0028] The third reconstructed image samples of the chroma components are encoded using the clipped third coefficient values.

[0029] According to another aspect, there is provided a decoding method comprising:

[0030] generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luma component;

[0031] generating second coefficient values ​​by applying an ALF (Adaptive Loop Filtering) process to the second reconstructed image samples of the chroma component;

[0032] clipping the second coefficient value;

[0033] generating a third coefficient value by adding the first coefficient value to the clipped second coefficient value;

[0034] clipping the third coefficient value; and

[0035] The third reconstructed image samples of the chroma component are decoded using the clipped third coefficient values.

[0036] In video coding technology, new methods are needed to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of the embodiments of the present disclosure, including the constituent elements of the embodiments of the present disclosure considered individually or in various combinations, can promote one or more of the following: improved coding efficiency, enhanced image quality, reduced processing resource utilization associated with encoding / decoding, reduced circuit size, and increased encoding / decoding processing speed.

[0037] In addition, some implementations of the embodiments of the present disclosure, including constituent elements of the embodiments of the present disclosure considered individually or in various combinations, can facilitate the appropriate selection of one or more elements in encoding and decoding, such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations. Note that the present disclosure includes disclosures about configurations and methods that can provide advantages in addition to those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing the increase in processing resource usage.

[0038] Additional benefits and advantages of the disclosed embodiments will become apparent from the description and drawings. Benefits and / or advantages can be obtained individually from the various embodiments and features of the description and drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.

[0039] It should be noted that the general or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] [ Figure 1 ] Figure 1is a schematic diagram showing one example of a functional configuration of a transmission system according to the embodiment.

[0041] [ Figure 2 ] Figure 2 is a conceptual diagram showing one example of a hierarchical structure of data in a stream.

[0042] [ Figure 3 ] Figure 3 is a conceptual diagram illustrating an example of a slice configuration.

[0043] [ Figure 4 ] Figure 4 is a conceptual diagram illustrating an example of a tile configuration.

[0044] [ Figure 5 ] Figure 5 is a conceptual diagram for illustrating an example of a coding structure in scalable coding.

[0045] [ Figure 6 ] Figure 6 is a conceptual diagram for illustrating an example of a coding structure in scalable coding.

[0046] [ Figure 7 ] Figure 7 is a block diagram showing a functional configuration of an encoder according to an embodiment.

[0047] [ Figure 8 ] Figure 8 is a functional block diagram showing an example of installation of an encoder.

[0048] [ Figure 9 ] Figure 9 is a flowchart indicating one example of an overall encoding process performed by an encoder.

[0049] [ Figure 10 ] Figure 10 is a conceptual diagram showing an example of block division.

[0050] [ Figure 11 ] Figure 11 is a block diagram showing one example of a functional configuration of a splitter according to the embodiment.

[0051] [ Figure 12 ] Figure 12 is a conceptual diagram for illustrating an example of a splitting pattern.

[0052] [ Figure 13A ] Figure 13A is a conceptual diagram showing an example of a syntax tree of a segmentation pattern.

[0053] [ Figure 13B ] Figure 13B is a conceptual diagram for illustrating another example of a syntax tree of a partitioning pattern.

[0054] [ Figure 14 ] Figure 14 is a chart indicating example transform basis functions for various transform types.

[0055] [ Figure 15 ] Figure 15 is a conceptual diagram for illustrating an example space-varying transform (SVT).

[0056] [ Figure 16 ] Figure 16 is a flowchart illustrating one example of a process performed by a converter.

[0057] [ Figure 17 ] Figure 17 is a flowchart illustrating another example of a process performed by the converter.

[0058] [ Figure 18 ] Figure 18 is a block diagram showing one example of a functional configuration of a quantizer according to an embodiment.

[0059] [ Figure 19 ] Figure 19 is a flowchart illustrating one example of a quantization process performed by a quantizer.

[0060] [ Figure 20 ] Figure 20 is a block diagram showing one example of a functional configuration of an entropy encoder according to an embodiment.

[0061] [ Figure 21 ] Figure 21 is a conceptual diagram for illustrating an example flow of a context-based adaptive binary arithmetic coding (CABAC) process in an entropy encoder.

[0062] [ Figure 22 ] Figure 22 is a block diagram showing one example of a functional configuration of a loop filter according to the embodiment.

[0063] [ Figure 23A ] Figure 23A is a conceptual diagram for illustrating an example of a filter shape used in an adaptive loop filter (ALF).

[0064] [ Figure 23B ] Figure 23B is a conceptual diagram for illustrating another example of a filter shape used in ALF.

[0065] [ Figure 23C ] Figure 23Cis a conceptual diagram for illustrating another example of a filter shape used in ALF.

[0066] [ Figure 23D ] Figure 23D is a conceptual diagram for illustrating an example flow of cross-component ALF (CC-ALF).

[0067] [ Figure 23E ] Figure 23E is a conceptual diagram for illustrating an example of a filter shape used in CC-ALF.

[0068] [ Figure 23F ] Figure 23F is a conceptual diagram for illustrating an example flow of Joint Chroma CCALF (JC-CCALF).

[0069] [ Figure 23G ] Figure 23G is a table showing example weight index candidates that may be employed in JC-CCALF.

[0070] [ Figure 24 ] Figure 24 : is a block diagram showing one example of a specific configuration of a loop filter used as a deblocking filter (DBF).

[0071] [ Figure 25 ] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having symmetric filtering characteristics with respect to a block boundary.

[0072] [ Figure 26 ] Figure 26 is a conceptual diagram for illustrating block boundaries where a deblocking filtering process is performed.

[0073] [ Figure 27 ] Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.

[0074] [ Figure 28 ] Figure 28 is a flowchart illustrating one example of a process performed by a predictor of an encoder.

[0075] [ Figure 29 ] Figure 29 is a flowchart illustrating another example of a process performed by a predictor of an encoder.

[0076] [ Figure 30 ] Figure 30 is a flowchart illustrating another example of a process performed by a predictor of an encoder.

[0077] [ Figure 31 ] Figure 31: is a conceptual diagram for illustrating sixty-seven intra prediction modes used in intra prediction in the embodiment.

[0078] [ Figure 32 ] Figure 32 is a flowchart illustrating one example of a process performed by an intra predictor.

[0079] [ Figure 33 ] Figure 33 is a conceptual diagram for illustrating an example of a reference picture.

[0080] [ Figure 34 ] Figure 34 is a conceptual diagram illustrating an example of a reference picture list.

[0081] [ Figure 35 ] Figure 35 is a flow chart illustrating an example basic process flow for inter-frame prediction.

[0082] [ Figure 36 ] Figure 36 is a flowchart showing an example of a derivation process of a motion vector.

[0083] [ Figure 37 ] Figure 37 is a flowchart illustrating another example of a derivation process of a motion vector.

[0084] [ Figure 38A ] Figure 38A is a conceptual diagram for illustrating an example representation of a pattern for MV derivation.

[0085] [ Figure 38B ] Figure 38B is a conceptual diagram for illustrating an example representation of a pattern for MV derivation.

[0086] [ Figure 39 ] Figure 39 is a flowchart illustrating an example of an inter prediction process in a conventional inter mode.

[0087] [ Figure 40 ] Figure 40 is a flowchart illustrating an example of an inter prediction process in normal merge mode.

[0088] [ Figure 41 ] Figure 41 is a conceptual diagram for illustrating an example of a motion vector derivation process in merge mode.

[0089] [ Figure 42 ] Figure 42 is a conceptual diagram for illustrating an example of an MV derivation process for a current picture through the HMVP merge mode.

[0090] [ Figure 43 ] Figure 43 is a flow chart illustrating one example of a frame rate up-conversion (FRUC) process.

[0091] [ Figure 44 ] Figure 44 is a conceptual diagram for illustrating an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.

[0092] [ Figure 45 ] Figure 45 is a conceptual diagram for illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.

[0093] [ Figure 46A ] Figure 46A is a conceptual diagram for illustrating an example of deriving a motion vector for each subblock based on motion vectors of a plurality of neighboring blocks.

[0094] [ Figure 46B ] Figure 46B is a conceptual diagram for illustrating an example of deriving a motion vector for each subblock in an affine mode in which three control points are used.

[0095] [ Figure 47A ] Figure 47A is a conceptual diagram illustrating example MV derivation at control points in affine mode.

[0096] [ Figure 47B ] Figure 47B is a conceptual diagram illustrating example MV derivation at control points in affine mode.

[0097] [ Figure 47C ] Figure 47C is a conceptual diagram illustrating example MV derivation at control points in affine mode.

[0098] [ Figure 48A ] Figure 48A is a conceptual diagram for illustrating an affine mode in which two control points are used.

[0099] [ Figure 48B ] Figure 48B is a conceptual diagram for illustrating an affine mode in which three control points are used.

[0100] [ Figure 49A ] Figure 49A is a conceptual diagram for illustrating an example of a method of MV derivation at a control point when the number of control points for a coding block and the number of control points for a current block are different from each other.

[0101] [ Figure 49B ] Figure 49Bis a conceptual diagram for illustrating another example of a method of MV derivation at a control point when the number of control points for a coding block and the number of control points for a current block are different from each other.

[0102] [ Figure 50 ] Figure 50 is a flowchart showing one example of a process in affine merge mode.

[0103] [ Figure 51 ] Figure 51 is a flowchart showing one example of a process in affine inter mode.

[0104] [ Figure 52A ] Figure 52A is a conceptual diagram for illustrating the generation of two triangular prediction images.

[0105] [ Figure 52B ] Figure 52B is a conceptual diagram illustrating an example of a first portion of a first partition overlapping a second partition, and first and second sample sets that may be weighted as part of a correction process.

[0106] [ Figure 52C ] Figure 52C is a conceptual diagram for illustrating a first portion of a first partition, which is a portion of the first partition overlapping with a portion of an adjacent partition.

[0107] [ Figure 53 ] Figure 53 is a flowchart showing one example of a process in triangle mode.

[0108] [ Figure 54 ] Figure 54 is a conceptual diagram for illustrating an example of an Advanced Temporal Motion Vector Prediction (ATMVP) mode in which an MV is derived in units of subblocks.

[0109] [ Figure 55 ] Figure 55 is a flow chart illustrating the relationship between merge mode and dynamic motion vector refresh (DMVR).

[0110] [ Figure 56 ] Figure 56 is a conceptual diagram for illustrating an example of DMVR.

[0111] [ Figure 57 ] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining MV.

[0112] [ Figure 58A ] Figure 58A is a conceptual diagram for illustrating an example of motion estimation in DMVR.

[0113] [ Figure 58B ] Figure 58B is a flowchart illustrating one example of a motion estimation process in DMVR.

[0114] [ Figure 59 ] Figure 59 is a flowchart illustrating an example of a process of generating a predicted image.

[0115] [ Figure 60 ] Figure 60 is a flowchart illustrating another example of a generation process of a predicted image.

[0116] [ Figure 61 ] Figure 61 is a flowchart illustrating one example of a correction process of a predicted image through overlapped block motion compensation (OBMC).

[0117] [ Figure 62 ] Figure 62 is a conceptual diagram for illustrating one example of a predicted image correction process by OBMC.

[0118] [ Figure 63 ] Figure 63 This is a conceptual diagram illustrating a model assuming uniform linear motion.

[0119] [ Figure 64 ] Figure 64 is a flowchart illustrating one example of an inter-frame prediction process according to BIO.

[0120] [ Figure 65 ] Figure 65 is a functional block diagram illustrating one example of a functional configuration of an inter-frame predictor that can perform inter-frame prediction according to BIO.

[0121] [ Figure 66A ] Figure 66A is a conceptual diagram for illustrating one example of a process of a predicted image generation method using a brightness correction process performed by the LIC.

[0122] [ Figure 66B ] Figure 66B is a flowchart showing one example of the process of a predicted image generation method using LIC.

[0123] [ Figure 67 ] Figure 67 is a block diagram showing a functional configuration of a decoder according to an embodiment.

[0124] [ Figure 68 ] Figure 68is a functional block diagram showing an example of installation of a decoder.

[0125] [ Figure 69 ] Figure 69 is a flowchart illustrating one example of the overall decoding process performed by a decoder.

[0126] [ Figure 70 ] Figure 70 is a conceptual diagram for illustrating the relationship between the segmentation determiner and other constituent elements.

[0127] [ Figure 71 ] Figure 71 is a block diagram showing one example of a functional configuration of an entropy decoder.

[0128] [ Figure 72 ] Figure 72 is a conceptual diagram for illustrating an example flow of a CABAC process in an entropy decoder.

[0129] [ Figure 73 ] Figure 73 is a block diagram showing one example of a functional configuration of an inverse quantizer.

[0130] [ Figure 74 ] Figure 74 is a flowchart illustrating one example of an inverse quantization process performed by an inverse quantizer.

[0131] [ Figure 75 ] Figure 75 is a flowchart showing one example of a process performed by an inverse converter.

[0132] [ Figure 76 ] Figure 76 is a flowchart illustrating another example of a process performed by an inverse converter.

[0133] [ Figure 77 ] Figure 77 is a block diagram showing one example of a functional configuration of a loop filter.

[0134] [ Figure 78 ] Figure 78 is a flowchart illustrating one example of a process performed by a predictor of a decoder.

[0135] [ Figure 79 ] Figure 79 is a flow chart illustrating another example of a process performed by a predictor of a decoder.

[0136] [ Figure 80A ] Figure 80A is a flow chart illustrating another example of a process performed by a predictor of a decoder.

[0137] [ Figure 80B ] Figure 80B is a flow chart illustrating another example of a process performed by a predictor of a decoder.

[0138] [ Figure 80C ] Figure 80C is a flow chart illustrating another example of a process performed by a predictor of a decoder.

[0139] [ Figure 81 ] Figure 81 is a diagram showing one example of a process performed by an intra predictor of a decoder.

[0140] [ Figure 82 ] Figure 82 is a flowchart illustrating an example of an MV derivation process in a decoder.

[0141] [ Figure 83 ] Figure 83 is a flowchart illustrating another example of an MV derivation process in a decoder.

[0142] [ Figure 84 ] Figure 84 is a flowchart illustrating an example of a process for inter prediction by conventional inter mode in a decoder.

[0143] [ Figure 85 ] Figure 85 is a flow chart illustrating an example of a process for inter prediction by conventional merge mode in a decoder.

[0144] [ Figure 86 ] Figure 86 is a flowchart illustrating an example of a process of inter-frame prediction by FRUC mode in a decoder.

[0145] [ Figure 87 ] Figure 87 is a flowchart illustrating an example of a process for inter prediction by affine merge mode in a decoder.

[0146] [ Figure 88 ] Figure 88 is a flowchart illustrating an example of a process of inter prediction by affine inter mode in a decoder.

[0147] [ Figure 89 ] Figure 89 is a flow chart illustrating an example of a process for inter-frame prediction by triangular mode in a decoder.

[0148] [ Figure 90 ] Figure 90 is a flow chart illustrating an example of a process of motion estimation by DMVR in a decoder.

[0149] [ Figure 91 ] Figure 91 is a flow chart illustrating an example process of motion estimation by DMVR in a decoder.

[0150] [ Figure 92 ] Figure 92 is a flowchart illustrating an example of a process of generating a predicted image in a decoder.

[0151] [ Figure 93 ] Figure 93 is a flowchart illustrating another example of a process of generating a predicted image in a decoder.

[0152] [ Figure 94 ] Figure 94 is a flowchart illustrating an example of a correction process of a predicted image by OBMC in a decoder.

[0153] [ Figure 95 ] Figure 95 is a flowchart illustrating an example of a correction process of a predicted image by BIO in a decoder.

[0154] [ Figure 96 ] Figure 96 is a flowchart illustrating an example of a correction process of a predicted image by LIC in a decoder.

[0155] [ Figure 97 ] Figure 97 is a flow chart of a sample process flow for decoding an image by applying a CCALF (Cross-Component Adaptive Loop Filter) process according to the first aspect.

[0156] [ Figure 98 ] Figure 98 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.

[0157] [ Figure 99 ] Figure 99 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.

[0158] [ Figure 100 ] Figure 100 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.

[0159] [ Figure 101 ] Figure 101 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment.

[0160] [ Figure 102 ] Figure 102 is a flow chart of a sample process flow for applying the CCALF process to decode an image according to the second aspect.

[0161] [ Figure 103 ] Figure 103 The sample positions of the cropping parameters to be parsed from, for example, a VPS, APS, SPS, PPS, slice header, CTU, or TU of a bitstream are illustrated.

[0162] [ Figure 104 ] Figure 104 An example of clipping parameters is shown.

[0163] [ Figure 105 ] Figure 105 is a flow chart of a sample process flow for decoding an image using filter coefficients applying a CCALF process according to the third aspect.

[0164] [ Figure 106 ] Figure 106 is a conceptual diagram indicating an example of the positions of filter coefficients to be used in the CCALF process.

[0165] [ Figure 107 ] Figure 107 is a conceptual diagram indicating an example of the positions of filter coefficients to be used in the CCALF process.

[0166] [ Figure 108 ] Figure 108 is a conceptual diagram indicating an example of the positions of filter coefficients to be used in the CCALF process.

[0167] [ Figure 109 ] Figure 109 is a conceptual diagram indicating an example of the positions of filter coefficients to be used in the CCALF process.

[0168] [ Figure 110 ] Figure 110 is a conceptual diagram indicating an example of the positions of filter coefficients to be used in the CCALF process.

[0169] [ Figure 111 ] Figure 111 is a conceptual diagram indicating a further example of the positions of filter coefficients to be used in the CCALF process.

[0170] [ Figure 112 ] Figure 112 is a conceptual diagram indicating a further example of the positions of filter coefficients to be used in the CCALF process.

[0171] [ Figure 113 ] Figure 113 is a block diagram illustrating a functional configuration of a CCALF process performed by an encoder and a decoder according to an embodiment.

[0172] [ Figure 114 ] Figure 114is a flow chart of a sample process flow for decoding an image by applying a CCALF process using a filter selected from a plurality of filters according to the fourth aspect.

[0173] [ Figure 115 ] Figure 115 An example of the process flow for selecting a filter is illustrated.

[0174] [ Figure 116 ] Figure 116 An example of a filter is shown.

[0175] [ Figure 117 ] Figure 117 An example of a filter is shown.

[0176] [ Figure 118 ] Figure 118 is a flow chart of a sample process flow for decoding an image by applying a CCALF process using parameters in accordance with the fifth aspect.

[0177] [ Figure 119 ] Figure 119 An example of the number of coefficients to be parsed from the bitstream is illustrated.

[0178] [ Figure 120 ] Figure 120 is a flow chart of a sample process flow for decoding an image by applying a CCALF process using parameters according to the sixth aspect.

[0179] [ Figure 121 ] Figure 121 is a conceptual diagram illustrating an example of generating a CCALF value of a luma component of a current chroma sample by calculating a weighted average of adjacent samples.

[0180] [ Figure 122 ] Figure 122 is a conceptual diagram illustrating an example of generating a CCALF value of a luma component of a current chroma sample by calculating a weighted average of adjacent samples.

[0181] [ Figure 123 ] Figure 123 is a conceptual diagram illustrating an example of generating a CCALF value of a luma component of a current chroma sample by calculating a weighted average of adjacent samples.

[0182] [ Figure 124 ] Figure 124 is a conceptual diagram illustrating an example of generating a CCALF value of a luma component of a current sample by calculating a weighted average of neighboring samples, where positions of the neighboring samples are adaptively determined as chroma types.

[0183] [ Figure 125 ] Figure 125is a conceptual diagram illustrating an example of generating a CCALF value of a luma component of a current sample by calculating a weighted average of neighboring samples, where positions of the neighboring samples are determined adaptively to a chroma type.

[0184] [ Figure 126 ] Figure 126 is a conceptual diagram illustrating an example of generating a CCALF value of a luminance component by applying a bit shift to an output value of a weighted calculation.

[0185] [ Figure 127 ] Figure 127 is a conceptual diagram illustrating an example of generating a CCALF value of a luminance component by applying a bit shift to an output value of a weighted calculation.

[0186] [ Figure 128 ] Figure 128 is a flow chart of a sample process flow for decoding an image by applying a CCALF process using parameters according to the seventh aspect.

[0187] [ Figure 129 ] Figure 129 Illustrate the sample positions of one or more parameters to be parsed from the bitstream.

[0188] [ Figure 130 ] Figure 130 A sample process for retrieving one or more parameters is depicted.

[0189] [ Figure 131 ] Figure 131 Sample values ​​of the second parameter are shown.

[0190] [ Figure 132 ] Figure 132 An example of parsing the second parameter using arithmetic coding is shown.

[0191] [ Figure 133 ] Figure 133 is a conceptual diagram of a variation of the present embodiment applied to rectangular partitions and non-rectangular partitions (eg, triangular partitions).

[0192] [ Figure 134 ] Figure 134 is a flowchart of an example process flow for decoding an image by applying a CCALF process using parameters according to the eighth aspect.

[0193] [ Figure 135 ] Figure 135 is a flow chart of a sample process flow for decoding an image by applying a CCALF process using parameters according to the eighth aspect.

[0194] [ Figure 136 ] Figure 136 Example locations for chroma sample types 0 through 5 are shown.

[0195] [ Figure 137 ] Figure 137 is a conceptual diagram showing sample symmetric filling.

[0196] [ Figure 138 ] Figure 138 is a conceptual diagram showing sample symmetric filling.

[0197] [ Figure 139 ] Figure 139 is a conceptual diagram showing sample symmetric filling.

[0198] [ Figure 140 ] Figure 140 is a conceptual diagram showing sample asymmetric filling.

[0199] [ Figure 141 ] Figure 141 is a conceptual diagram showing sample asymmetric filling.

[0200] [ Figure 142 ] Figure 142 is a conceptual diagram showing sample asymmetric filling.

[0201] [ Figure 143 ] Figure 143 is a conceptual diagram showing sample asymmetric filling.

[0202] [ Figure 144 ] Figure 144 is a conceptual diagram illustrating further sample asymmetric filling.

[0203] [ Figure 145 ] Figure 145 is a conceptual diagram illustrating further sample asymmetric filling.

[0204] [ Figure 146 ] Figure 146 is a conceptual diagram illustrating further sample asymmetric filling.

[0205] [ Figure 147 ] Figure 147 is a conceptual diagram illustrating further sample asymmetric filling.

[0206] [ Figure 148 ] Figure 148 is a conceptual diagram illustrating further sample symmetric filling.

[0207] [ Figure 149 ] Figure 149 is a conceptual diagram illustrating further sample symmetric filling.

[0208] [ Figure 150 ] Figure 150 is a conceptual diagram illustrating further sample symmetric filling.

[0209] [ Figure 151 ] Figure 151 is a conceptual diagram illustrating further sample asymmetric filling.

[0210] [ Figure 152 ] Figure 152 is a conceptual diagram illustrating further sample asymmetric filling.

[0211] [ Figure 153 ] Figure 153 is a conceptual diagram illustrating further sample asymmetric filling.

[0212] [ Figure 154 ] Figure 154 is a conceptual diagram illustrating further sample asymmetric filling.

[0213] [ Figure 155 ] Figure 155 A further example of padding with horizontal and vertical virtual borders is illustrated.

[0214] [ Figure 156 ] Figure 156 is a block diagram illustrating a functional configuration of an encoder and a decoder according to an example, in which symmetric padding is used on a virtual boundary position of an ALF and symmetric or asymmetric padding is used on a virtual boundary position of a CC-ALF.

[0215] [ Figure 157 ] Figure 157 is a block diagram illustrating a functional configuration of an encoder and a decoder according to another example, in which symmetric padding is used on a virtual boundary position of an ALF and unilateral padding is used on a virtual boundary position of a CC-ALF.

[0216] [ Figure 158 ] Figure 158 is a conceptual diagram illustrating an example of single-side padding with a horizontal or vertical virtual border.

[0217] [ Figure 159 ] Figure 159 is a conceptual diagram illustrating an example of single-side padding with horizontal and vertical virtual boundaries.

[0218] [ Figure 160 ] Figure 160 is a diagram showing an example overall configuration of a content providing system for realizing a content distribution service.

[0219] [ Figure 161 ] Figure 161 is a conceptual diagram for illustrating an example of a display screen of a web page.

[0220] [ Figure 162 ] Figure 162 is a conceptual diagram for illustrating an example of a display screen of a web page.

[0221] [ Figure 163 ] Figure 163 is a block diagram illustrating one example of a smartphone.

[0222] [ Figure 164 ] Figure 164 is a block diagram illustrating an example of a functional configuration of a smartphone. DETAILED DESCRIPTION

[0223] In the drawings, like reference numerals denote similar elements unless context dictates otherwise. The sizes and relative positions of elements in the drawings are not necessarily drawn to scale.

[0224] Hereinafter, embodiments will be described with reference to the accompanying drawings. Note that the embodiments described below each illustrate general or specific examples. The numerical values, shapes, materials, components, arrangement and connection of components, steps, relationships between steps, and order, etc. indicated in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0225] Embodiments of encoders and decoders are described below. The embodiments are examples of encoders and decoders, and the processes and / or configurations presented in the description of aspects of the present disclosure are applicable to the encoders and decoders. The processes and / or configurations may also be implemented in encoders and decoders different from the encoders and decoders according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented:

[0226] (1) Any component of an encoder or decoder according to the embodiments presented in the description of aspects of this disclosure may be replaced with or combined with another component presented anywhere in the description of aspects of this disclosure.

[0227] (2) In an encoder or decoder according to an embodiment, any changes may be made to the functions or processes performed by one or more components of the encoder or decoder, such as addition, replacement, removal, etc. of the functions or processes. For example, any function or process may be replaced with or combined with another function or process presented anywhere in the description of aspects of this disclosure.

[0228] (3) In the method implemented by the encoder or decoder according to the embodiment, any changes may be made, such as adding, replacing, and removing one or more processes included in the method. For example, any process in the method may be replaced by or combined with another process presented anywhere in the description of the aspects of the present disclosure.

[0229] (4) One or more components included in an encoder or decoder according to an embodiment may be combined with components presented anywhere in the description of aspects of the present disclosure, may be combined with components including one or more functions presented anywhere in the description of aspects of the present disclosure, and may be combined with components that implement one or more processes implemented by components presented in the description of aspects of the present disclosure.

[0230] (5) A component including one or more functions of an encoder or decoder according to an embodiment, or a component implementing one or more processes of an encoder or decoder according to an embodiment, may be combined with or replaced by a component presented anywhere in the description of an aspect of the present disclosure, combined with or replaced by a component including one or more functions presented anywhere in the description of an aspect of the present disclosure, or combined with or replaced by a component implementing one or more processes presented anywhere in the description of an aspect of the present disclosure.

[0231] (6) In a method implemented by an encoder or decoder according to an embodiment, any process included in the method may be replaced or combined with a process presented anywhere in the description of aspects of this disclosure or with any corresponding or equivalent process.

[0232] (7) One or more processes included in the method implemented by the encoder or decoder according to the embodiment may be combined with processes presented anywhere in the description of various aspects of the present disclosure.

[0233] (8) The implementation of the processes and / or configurations presented in the description of aspects of the present disclosure is not limited to the encoder or decoder according to the embodiments. For example, the processes and / or configurations may be implemented in a device for a purpose different from that of the motion picture encoder or motion picture decoder disclosed in the embodiments.

[0234] (Definition of terms)

[0235] The corresponding terms may be defined as indicated below as examples.

[0236] An image is a data unit configured with a group of pixels, a picture, or includes blocks smaller than pixels. In addition to videos, images also include still images.

[0237] A picture is an image processing unit configured with a group of pixels and may also be referred to as a frame or field. For example, a picture may take the form of an array of luma samples in a monochrome format or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0238] A block is a processing unit, which is a group of a certain number of pixels. Blocks can have any number of different shapes. For example, a block can have a rectangle of M×N (M columns×N rows) pixels, a square of M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, tiles, bricks, CTUs, super blocks, basic partition units, VPDUs, processing partition units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and sub-blocks. A block can take the form of an M×N sample array or an M×N transform coefficient array. For example, a block can be a square or rectangular pixel area that includes a luminance matrix and two chrominance matrices.

[0239] A pixel or sample is the smallest point of an image. Pixels or samples include pixels at integer positions and pixels at sub-pixel positions, such as those generated based on pixels at integer positions.

[0240] A pixel value or sample value is a characteristic value of a pixel and may include one or more of a brightness value, a chrominance value, an RGB grayscale level, a depth value, a binary value of zero or one, and the like.

[0241] Chroma or chrominance is the intensity of a color, typically represented by the symbols Cb and Cr, which specify that the value of an array of samples or a single sample value represents the value of one of two color difference signals associated with a primary color.

[0242] Luma or luminance is the brightness of an image and is typically represented by the symbols or subscripts Y or L, which specify that the value of an array of samples or a single sample value represents a monochromatic signal associated with a primary color.

[0243] A flag comprises one or more bits indicating the value of, for example, a parameter or an index. A flag may be a binary flag, which indicates a binary value of the flag, or it may indicate a non-binary value of a parameter.

[0244] Signals convey information that is symbolized or encoded into them. Signals include discrete digital signals and continuous analog signals.

[0245] A stream or bitstream is a string of digital data that is a stream of digital data. A stream or bitstream can be a single stream or can be configured with multiple streams having multiple hierarchical layers. A stream or bitstream can be transmitted in a serial communication manner using a single transmission path, or can be transmitted in a packet communication manner using multiple transmission paths.

[0246] Difference refers to various mathematical differences, such as simple difference (xy), absolute value of difference (|xy|), square difference (x^2-y^2), square root of difference (√(x–y)), weighted difference (ax-by: a and b are constants), offset difference (x-y+a: a is an offset), etc. In the case of scalars, simple difference is sufficient, and difference calculation is included.

[0247] The sum refers to various mathematical sums, such as simple sum (x+y), absolute value of sum (|x+y|), square sum (x^2+y^2), square root of sum (√(x+y)), weighted difference (ax+by: a and b are constants), offset sum (x+y+a: a is an offset), etc. In the case of scalars, a simple sum is sufficient, and sum calculations are included.

[0248] A frame is a combination of a top field and a bottom field, where sample lines 0, 2, 4, ... are derived from the top field and sample lines 1, 3, 5, ... are derived from the bottom field.

[0249] A slice is an integer number of coding tree units contained in an independent slice segment and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same access unit.

[0250] A tile is a rectangular region of coding tree blocks within a particular tile column and a particular tile row in a picture. A tile can be a rectangular region of a frame that is intended to be independently decodable and coded, although loop filtering across tile edges can still be applied.

[0251] A coding tree unit (CTU) can be a coding tree block of luma samples for a picture with three sample arrays, or two corresponding coding tree blocks of chroma samples. Alternatively, a CTU can be a coding tree block of samples for one of a monochrome picture and a picture coded using three separate color planes and syntax structures for coding samples. A superblock can be a square block of 64×64 pixels consisting of one or two mode information blocks, or recursively divided into four 32×32 blocks, which themselves can be further divided.

[0252] (System Configuration)

[0253] First, a transmission system according to the embodiment will be described. Figure 1 is a schematic diagram showing one example of the configuration of a transmission system 400 according to the embodiment.

[0254] The transmission system 400 is a system for transmitting a stream generated by encoding an image and decoding the transmitted stream. As shown in the figure, the transmission system 400 includes Figure 1 An encoder 100, a network 300 and a decoder 200 are shown.

[0255] An image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image and outputs the stream to the network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by encoding.

[0256] It should be noted that the image before being encoded by the encoder 100 is also referred to as the original image, original signal or original sample. The image can be a video or a still image. The image is a general concept of a sequence, a picture and a block, and therefore is not limited to a spatial area of ​​a specific size and a time area of ​​a specific size unless otherwise specified. An image is an array of pixels or pixel values, and a signal representing an image or pixel values ​​is also referred to as a sample. The stream can be referred to as a bit stream, a coded bit stream, a compressed bit stream or a coded signal. In addition, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, an image encoding method or a video encoding method.

[0257] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 200 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasting, satellite broadcasting, etc. Alternatively, the network 300 may be replaced by a recording medium such as a digital versatile disk (DVD) and a Blu-ray disc (BD) on which the stream is recorded.

[0258] The decoder 200 generates a decoded image as an uncompressed image by, for example, decoding a stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method adopted by the encoder 100.

[0259] It should be noted that the decoder 200 may also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 may also be referred to as a decoding method, an image decoding method, or a video decoding method.

[0260] (Data Structure)

[0261] Figure 2 is a conceptual diagram showing an example of a hierarchical structure of data in a stream. Figure 1 The transmission system 400 is described Figure 2 The stream includes, for example, a video sequence. Figure 2 As shown in (a), a video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplemental enhancement information (SEI) and multiple pictures.

[0262] In a video having multiple layers, the VPS may include encoding parameters shared between some of the multiple layers, and encoding parameters related to some of the multiple layers included in the video or to a single layer.

[0263] The SPS includes parameters for the sequence, i.e., encoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the picture. It should be noted that there may be multiple SPSs.

[0264] The PPS includes parameters for a picture, i.e., encoding parameters that decoder 200 refers to in order to decode each picture in a sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the use of weighted prediction. It should be noted that multiple PPSs may exist. Each of the SPS and PPS may be referred to simply as a parameter set.

[0265] like Figure 2 As shown in (b), a picture may include a picture header and one or more slices. The picture header includes coding parameters that the decoder 200 refers to in order to decode one or more slices.

[0266] like Figure 2 As shown in (c), a slice includes a slice header and one or more bricks. The slice header includes coding parameters that the decoder 200 refers to in order to decode one or more bricks.

[0267] like Figure 2 As shown in (d), a brick includes one or more coding tree units (CTUs).

[0268] Note that a picture may not include any slices and may include slice groups instead of slices. In this case, a slice group includes at least one slice. Furthermore, a tile may include slices.

[0269] CTU is also called super block or basis splitting unit. Figure 2 As shown in (e), a CTU includes a CTU header and at least one coding unit (CU). As shown in the figure, the CTU includes four coding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes coding parameters that the decoder 200 refers to in order to decode at least one CU.

[0270] A CU can be split into multiple smaller CUs. As shown in the figure, CU (10) is not split into smaller coding units; CU (11) is split into four smaller coding units CU (110), CU (111), CU (112) and CU (113); CU (12) is not split into smaller coding units; and CU (13) is split into seven smaller coding units CU (1310), CU (1311), CU (1312), CU (1313), CU (132), CU (133) and CU (134). Figure 2 As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information representing the prediction residual described later. Although the CU is basically the same as the prediction unit (PU) and the transform unit (TU), it should be noted that, for example, the sub-block transform (SBT) to be described later may include multiple TUs smaller than the CU. In addition, the CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. The VPDU is, for example, a fixed unit that can be processed at one stage when pipeline processing is performed in hardware.

[0271] It should be noted that a stream may not include Figure 2 All hierarchical layers shown. The order of the hierarchical layers can be exchanged, or any hierarchical layer can be replaced by another hierarchical layer. Here, a picture that is the target of a process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as a current picture. When the process is an encoding process, the current picture represents the current picture to be encoded, and when the process is a decoding process, the current picture represents the current picture to be decoded. Likewise, for example, a CU or a CU block that is the target of a process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as a current block. When the process is an encoding process, the current block represents the current block to be encoded, and when the process is a decoding process, the current block represents the current block to be decoded.

[0272] (Image structure: slice / slice)

[0273] A picture may be configured with one or more slice units or one or more tiling units to facilitate parallel encoding / decoding of the picture.

[0274] A slice is a basic coding unit included in a picture. A picture may include, for example, one or more slices. In addition, a slice includes one or more coding tree units (CTUs).

[0275] Figure 3 is a conceptual diagram showing an example of a slice configuration. Figure 3In , a picture includes 11×8 CTUs and is divided into four slices (slices 1 to 4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is the shape obtained by horizontally partitioning the picture. The boundary of each slice does not need to coincide with the end of the image and can coincide with any boundary between CTUs in the image. The processing order (coding order or decoding order) of the CTUs in the slice is, for example, a raster scan order. The slice includes a slice header and coded data. The characteristics of the slice can be written in the slice header. The characteristics may include the CTU address of the top CTU in the slice, the slice type, etc.

[0276] A tile is a unit of a rectangular area included in a picture. Tiles of a picture may be assigned numbers called TileIds in a raster scan order.

[0277] Figure 4 is a conceptual diagram showing an example of a sharding configuration. Figure 4 In

[15] , a picture includes 11×8 CTUs and is partitioned into four slices (slices 1 to 4) of rectangular areas. When using slices, the order in which the CTUs are processed may be different from the order in which they are processed when slices are not used. When slices are not used, multiple CTUs in a picture are typically processed in raster scan order. When multiple slices are used, at least one CTU in each of the multiple slices is processed in raster scan order. For example, Figure 4 As shown, the processing order of the CTUs included in slice 1 is from the left end of the first column of slice 1 to the right end of the first column of slice 1, and then continues from the left end of the second column of slice 1 to the right end of the second column of slice 1.

[0278] It should be noted that a shard may include one or more slices, and a slice may include one or more shards.

[0279] It should be noted that a picture can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more tiles. A picture can be configured with one of a tile set, a tile group, and a tile. For example, it is assumed that the order in which multiple tiles are scanned in raster scan order for each tile set is the basic coding order of the tiles. It is assumed that a set of one or more tiles that are consecutive in the basic coding order in each tile set is a tile group. Such a picture can be processed by the segmenter 102 described later (see Figure 7 ) to configure.

[0280] (Scalable Coding)

[0281] Figure 5 and Figure 6is a conceptual diagram showing an example of a scalable stream structure, and for convenience, reference will be made to Figure 1 Provide a description.

[0282] like Figure 5 As shown, encoder 100 can generate a temporally and spatially scalable stream by dividing each of multiple pictures into any of multiple layers and encoding the pictures in the layers. For example, encoder 100 encodes pictures for each layer, thereby achieving scalability when enhancement layers exist above a base layer. This encoding of each picture is also called scalable coding. In this way, decoder 200 can switch the image quality of the image displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as decoder 200's processing power and external factors such as communication bandwidth conditions. As a result, decoder 200 can decode content while freely switching between low and high resolutions. For example, a user of a stream may watch a streamed video on a smartphone on their way home and continue watching the video at home on a device (e.g., an internet-connected TV). It should be noted that each of the smartphone and device described above includes a decoder 200 with the same or different capabilities. In this case, when the device decodes a layer into a higher layer in the stream, the user can watch the video at high quality at home. In this way, the encoder 100 does not need to generate a plurality of streams having different image qualities for the same content, and thus can reduce the processing load.

[0283] In addition, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 may generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the metadata. Super-resolution imaging may include, for example, an improvement in the signal-to-noise ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values ​​in a filtering process, machine learning, or least squares method (used in super-resolution processing).

[0284] In an embodiment, a configuration may be provided in which a picture is divided into, for example, slices according to the meaning of, for example, an object in the picture. In this case, the decoder 200 can decode only a partial area in the picture by selecting the slice to be decoded. In addition, the attributes of the object (person, car, ball, etc.) and the position of the object in the picture (coordinates in the same image) can be stored as metadata. In this case, the decoder 200 is able to identify the position of the desired object based on the metadata and determine the slice that includes the object. For example, Figure 6As shown, a data storage structure different from that of image data can be used to store metadata, such as SEI (Supplementary Enhancement Information) messages in HEVC. The metadata indicates, for example, the position, size or color of the main object.

[0285] The metadata can be stored in units of multiple pictures (e.g., streams, sequences, or random access units). In this way, the decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the picture unit information, it can identify the picture in which the object (person) appears and determine the position of the object in the picture.

[0286] (Encoder)

[0287] An encoder according to an embodiment will be described. Figure 7 1 is a block diagram illustrating a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes a video in units of blocks.

[0288] like Figure 7 As shown, the encoder 100 is a device for encoding an image in units of blocks, and includes a partitioner 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown in the figure, the intra-frame predictor 124 and the inter-frame predictor 126 are part of the prediction controller.

[0289] The encoder 100 is implemented as, for example, a general-purpose processor and memory. In this case, when the software program stored in the memory is executed by the processor, the processor acts as the partitioner 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the partitioner 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128.

[0290] (Encoder installation example)

[0291] Figure 8 1 is a functional block diagram showing an example of an installation of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The various components of the encoder 100 are mounted on Figure 8 As shown, processor a1 and memory a2.

[0292] Processor a1 is a circuit that performs information processing and is coupled to memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes an image. Processor a1 may be a processor such as a CPU. In addition, processor a1 may be a collection of multiple electronic circuits. In addition, for example, processor a1 may be responsible for Figure 7 The roles of two or more constituent elements among the multiple constituent elements of the encoder 100 and the like shown.

[0293] Memory a2 is a dedicated or general-purpose memory for storing information used by processor a1 to encode images. Memory a2 may be an electronic circuit and may be connected to processor a1. Furthermore, memory a2 may be included in processor a1. Furthermore, memory a2 may be a collection of multiple electronic circuits. Furthermore, memory a2 may be a magnetic disk, an optical disk, or the like, or may be represented as a storage device, a recording medium, or the like. Furthermore, memory a2 may be either non-volatile memory or volatile memory.

[0294] For example, the memory a2 may store an image to be encoded or a bit stream corresponding to an encoded image. In addition, the memory a2 may store a program for causing the processor a1 to encode an image.

[0295] In addition, for example, memory a2 can act as Figure 7 The memory a2 may serve as a memory element for storing information. Figure 7 The roles of the block memory 118 and the frame memory 122 are shown. More specifically, the memory a2 can store reconstructed blocks, reconstructed pictures, etc.

[0296] It should be noted that in the encoder 100, it is not necessary to implement Figure 7 All of the multiple constituent elements and the like shown, and not all of the processes described here may be performed. Figure 7 A portion of the illustrated constituent elements and the like may be included in another device, or a portion of the process described herein may be performed by another device.

[0297] Hereinafter, the overall flow of a process performed by the encoder 100 is described, and then each constituent element included in the encoder 100 will be described.

[0298] (Overall flow of the encoding process)

[0299] Figure 9is a flowchart showing one example of the overall encoding process performed by the encoder 100, and for convenience will be referred to as Figure 7 Provide a description.

[0300] First, the encoder 100's segmenter 102 segments each picture included in the input image into a plurality of blocks of a fixed size (e.g., 128×128 pixels) (step Sa_1). The segmenter 102 then selects a segmentation pattern for the fixed-size blocks (also referred to as a block shape) (step Sa_2). In other words, the segmenter 102 further segments the fixed-size blocks into a plurality of blocks that form the selected segmentation pattern. For each of the plurality of blocks, the encoder 100 performs steps Sa_3 through Sa_9 for that block (i.e., the current block to be encoded).

[0301] The prediction controller 128 and the prediction executor (which includes the intra predictor 124 and the inter predictor 126) generate a prediction image for the current block (step Sa-3). The prediction image may also be referred to as a prediction signal, a prediction block, or a prediction sample.

[0302] Next, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Sa_4). The prediction residual may also be called a prediction error.

[0303] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients may sometimes be referred to as a coefficient block.

[0304] Next, the entropy encoder 110 encodes (specifically, entropy encodes) the plurality of quantized coefficients and prediction parameters related to the generation of the predicted image to generate a stream (step Sa_6). This stream may sometimes be referred to as an encoded bit stream or a compressed bit stream.

[0305] Next, the inverse quantizer 112 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 114 performs inverse transformation on the result to restore the prediction residual (step Sa_7).

[0306] Next, the adder 116 adds the predicted image and the restored prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image can also be called a reconstructed block or a decoded image block.

[0307] When the reconstructed image is generated, the loop filter 120 performs filtering on the reconstructed image as necessary (step Sa_9 ).

[0308] The encoder 100 then determines whether encoding of the entire picture has been completed (step Sa_10). When it is determined that encoding has not been completed (No in step Sa_10), the process starting from step Sa_2 is repeatedly performed on the next block of the picture.

[0309] Although in the above example, the encoder 100 selects a partitioning mode for a fixed-size block and encodes each block according to the partitioning mode, it should be noted that each block may be encoded according to a corresponding partitioning mode from among a plurality of partitioning modes. In this case, the encoder 100 may evaluate the cost of each of the plurality of partitioning modes and, for example, may select a stream that can be obtained by encoding according to the partitioning mode that produces the smallest cost as the output stream.

[0310] As shown in the figure, the processes in steps Sa_1 to Sa_10 are sequentially performed by the encoder 100. Alternatively, two or more processes may be performed in parallel, the processes may be reordered, and so on.

[0311] The encoding process employed by encoder 100 is a hybrid encoding process using predictive coding and transform coding. Predictive coding is performed by an encoding loop configured with a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, a block memory 118, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. In other words, the prediction execution unit configured with intra-frame predictor 124 and inter-frame predictor 126 is part of the encoding loop.

[0312] (Splitters)

[0313] The splitter 102 splits each picture contained in the original image into a plurality of blocks and outputs each block to the subtractor 104. For example, the splitter 102 first splits the picture into blocks of a fixed size (e.g., 128×128 pixels). Other fixed block sizes may be used. Fixed-size blocks are also referred to as coding tree units (CTUs). The splitter 102 then splits each fixed-size block into blocks of a variable size (e.g., 64×64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a splitting mode. Variable-size blocks may also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in a picture may be processed in units of CUs, PUs, or TUs.

[0314] Figure 10 : is a conceptual diagram for illustrating an example of block segmentation according to an embodiment. Figure 10, solid lines represent block boundaries of blocks partitioned by quadtree block partitioning, and dotted lines represent block boundaries of blocks partitioned by binarytree block partitioning.

[0315] Here, the block 10 is a square block having 128×128 pixels (128×128 block). This 128×128 block 10 is first partitioned into four square blocks of 64×64 pixels (quadtree block partitioning).

[0316] The 64×64 pixel block in the upper left corner is further vertically divided into two rectangular 32×64 pixel blocks, and the 32×64 pixel block on the left is further vertically divided into two rectangular 16×64 pixel blocks (binary tree block division). As a result, the 64×64 pixel block in the upper left corner is divided into two 16×64 pixel blocks 11 and 12 and one 32×64 pixel block 13.

[0317] The upper right 64×64 pixel block is horizontally partitioned into two rectangular 64×32 pixel blocks 14 and 15 (binary tree block partitioning).

[0318] The 64×64 pixel block in the lower left corner is first split into four square 32×32 pixel blocks (quadtree block splitting). The upper left block and the lower right block of the four square 32×32 pixel blocks are further split. The square 32×32 pixel block in the upper left corner is vertically split into two rectangular 16×32 pixel blocks, and the right 16×32 pixel block is further horizontally split into two 16×16 pixel blocks (binary tree block splitting). The 32×32 pixel block in the lower right corner is horizontally split into two 32×16 pixel blocks (binary tree block splitting). The square 32×32 pixel block in the upper right corner is horizontally split into two rectangular 32×16 pixel blocks (binary tree block splitting). As a result, the square 64×64 pixel block in the lower left corner is divided into a rectangular 16×32 pixel block 16, two square 16×16 pixel blocks 17 and 18, two square 32×32 pixel blocks 19 and 20, and two rectangular 32×16 pixel blocks 21 and 22.

[0319] The 64×64 pixel block 23 in the lower right corner is not segmented.

[0320] As mentioned above, in Figure 10 In FIG, block 10 is partitioned into 13 variable-sized blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. This type of partitioning is also called quadtree plus binary tree (QTBT) partitioning.

[0321] It should be noted that in Figure 10 In the example above, a block is partitioned into four or two blocks (quadtree or binary tree block partitioning), but the partitioning is not limited to these examples. For example, a block can be partitioned into three blocks (ternary block partitioning). Partitioning that includes this ternary block partitioning is also called multi-type tree (MBT) partitioning.

[0322] Figure 11 1 is a block diagram showing an example of the functional configuration of the splitter 102 according to one embodiment. Figure 11 As shown, the partitioner 102 may include a block partition determiner 102a. As an example, the block partition determiner 102a may perform the following process.

[0323] For example, the block segmentation determiner 102a may obtain or retrieve block information from the block memory 118 and / or the frame memory 122 and determine a segmentation mode (e.g., the above-described segmentation mode) based on the block information. The segmenter 102 segments the original image according to the segmentation mode and outputs at least one block obtained by the segmentation to the subtractor 104.

[0324] Furthermore, for example, the block partition determiner 102a outputs one or more parameters indicating the determined partition mode (e.g., the partition mode described above) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 may transform the prediction residual based on the one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 may generate a predicted image based on the one or more parameters. Furthermore, the entropy encoder 110 may perform entropy encoding on the one or more parameters.

[0325] Parameters related to the split mode can be written in the stream, as an example, as shown below.

[0326] Figure 12 is a conceptual diagram illustrating examples of partitioning modes. Examples of partitioning modes include: partitioning into four regions (QT), in which a block is partitioned into two regions both horizontally and vertically; partitioning into three regions (HT or VT), in which a block is partitioned in the same direction at a ratio of 1:2:1; partitioning into two regions (HB or VB), in which a block is partitioned in the same direction at a ratio of 1:1; and no partitioning (NS).

[0327] It should be noted that the division mode does not have a block division direction when the division is into four regions or when no division is performed, and has division direction information when the division is into two regions or three regions.

[0328] Figure 13A is a conceptual diagram showing an example of a syntax tree of a segmentation pattern.

[0329] Figure 13B is a conceptual diagram for illustrating another example of a syntax tree of a partitioning pattern.

[0330] Figure 13A and Figure 13Bis a conceptual diagram showing an example of a syntax tree of a segmentation pattern. Figure 13A In the example, first, there is information indicating whether to perform segmentation (S: segmentation flag), and next there is information indicating whether to perform segmentation into 4 areas (QT: QT flag). Next there is information indicating which of three areas and two areas to perform segmentation (TT: TT flag or BT: BT flag), and then there is information indicating the direction of division (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of at least one block obtained by segmentation according to such a segmentation pattern can be further repeatedly segmented in a similar process. In other words, as an example, whether to perform segmentation, whether to perform segmentation into four areas, which of the horizontal direction and the vertical direction is the direction in which the segmentation method is to be performed, which of the segmentation into three areas and the segmentation into two areas to be performed can be recursively determined, and can be determined according to Figure 13A The encoding order disclosed by the syntax tree shown determines the encoding of the result in the stream.

[0331] In addition, although the information items indicating S, QT, TT and Ver are Figure 13A The information items indicated in the syntax tree are arranged in the order listed, and the information items indicating S, QT, Ver and BT respectively can also be arranged in the order listed. Figure 13B In the example of , first, there is information indicating whether to perform splitting (S: Split Flag), followed by information indicating whether to perform splitting into four regions (QT: QT Flag). Next, there is information indicating the splitting direction (Ver: Vertical Flag, or Hor: Horizontal Flag), and next, there is information indicating whether to perform splitting into two regions or splitting into three regions (BT: BT Flag or TT: TT Flag).

[0332] It should be noted that the above-described division patterns are examples, and a division pattern other than the described division pattern may be used, or a part of the described division pattern may be used.

[0333] (Subtractor)

[0334] The subtractor 104 subtracts the predicted image (prediction samples input from the prediction controller 128 indicated below) from the original image in units of blocks, which is input from the divider 102 and divided by the divider 102. In other words, the subtractor 104 calculates the prediction residual (also referred to as the error) of the current block. The subtractor 104 then outputs the calculated prediction residual to the transformer 106.

[0335] The original image may be an image that has been input to the encoder 100 as a signal (for example, a luminance signal and two chrominance signals) representing an image of each picture included in a video. The signal representing the image may also be referred to as a sample.

[0336] (Converter)

[0337] The transformer 106 transforms the prediction residual in the spatial domain into a transform coefficient in the frequency domain and outputs the transform coefficient to the quantizer 108. More specifically, the transformer 106 applies, for example, a defined discrete cosine transform (DCT) or discrete sine transform (DST) to the prediction residual in the spatial domain. The defined DCT or DST may be predefined.

[0338] It should be noted that the transformer 106 can adaptively select a transform type from a plurality of transform types and transform the prediction residual into a transform coefficient by using a transform basis function corresponding to the selected transform type. This transform is also called an explicit multi-kernel transform (EMT) or an adaptive multi-kernel transform (AMT). The transform basis function may also be referred to as a basis.

[0339] Transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transform types can be expressed as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 is a diagram of example transform basis functions indicating example transform types. Figure 14 In , N represents the number of input pixels. For example, selecting a transform type from a plurality of transform types may depend on a prediction type (one of intra prediction and inter prediction) and may depend on an intra prediction mode.

[0340] Information indicating whether such EMT or AMT is applied (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and may also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0341] In addition, the transformer 106 can re-transform the transform coefficients (which are the transform results). This re-transformation is also called adaptive secondary transform (AST) or non-separable secondary transform (NSST). For example, the transformer 106 performs re-transformation in units of sub-blocks (e.g., 4×4 pixel sub-blocks) included in the transform coefficient block corresponding to the intra-frame prediction residual. Information indicating whether NSST is applied and information about the transform matrix used in NSST are typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level, and can also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0342] Transformer 106 can employ both separable and non-separable transforms. A separable transform is a method in which a transform is performed multiple times by performing a transform separately for each of a plurality of directions according to the dimensionality of the input. A non-separable transform is a method in which a collective transform is performed in which two or more dimensions in a multi-dimensional input are collectively treated as a single dimension.

[0343] In one example of a non-separable transform, when the input is a 4x4 pixel block, the 4x4 pixel block is considered to be a single array containing 16 elements, and the transform applies a 16x16 transformation matrix to that array.

[0344] In another example of a non-separable transform, an input block of 4x4 pixels is treated as a single array containing 16 elements, and a transform that applies a number of given rotations to the array (a hypercube given transform) can then be performed.

[0345] In the transform in the transformer 106, the transform type of the transform basis function to be transformed into the frequency domain can be switched according to the region in the CU. Examples include spatially varying transform (SVT).

[0346] Figure 15 is a conceptual diagram for illustrating an example of SVT.

[0347] In SVT, Figure 15As shown, the CU is split into two equal regions horizontally or vertically, and only one of the regions is transformed into the frequency domain. The transform base type can be set for each region. For example, DST7 and DST8 are used. For example, in the two regions obtained by splitting the CU vertically into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, in the two regions, DST7 can be used for the region at position 1. Similarly, in the two regions obtained by splitting the CU horizontally into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Although in Figure 15 In the example shown, one of the two regions in the CU is transformed and the other is not transformed, but each of the two regions can be transformed. In addition, the partitioning method can include not only partitioning into two regions but also partitioning into four regions. In addition, the partitioning method can be more flexible. For example, information indicating the partitioning method can be encoded and can be signaled in the same way as CU partitioning. It should be noted that SVT can also be called sub-block transform (SBT).

[0348] The AMT and EMT described above may be referred to as MTS (Multiple Transform Selection). When MTS is applied, transform types such as DST7, DCT8, etc. may be selected, and information indicating the selected transform type may be encoded as index information for each CU. There is another process called IMTS (Implicit MTS) as a process for selecting a transform type to be used for an orthogonal transform performed without encoding index information. When IMTS is applied, for example, when the CU has a rectangular shape, the orthogonal transform of the rectangular shape may be performed using DST7 (for the short side) and DST2 (for the long side). In addition, for example, when the CU has a square shape, the orthogonal transform of the rectangular shape may be performed by using DCT2 when MTS is valid in the sequence and using DST7 when MTS is invalid in the sequence. DCT2 and DST7 are merely examples. Other transform types may be used, and the combination of transform types used may also be changed to a different transform type combination. IMTS may be used only for intra-prediction blocks, or may be used for both intra-prediction blocks and inter-prediction blocks.

[0349] The three processes of MTS, SBT and IMTS have been described above as selection processes for selectively switching the transform type used for orthogonal transform. However, all three selection processes may be adopted, or only some of the selection processes may be selectively adopted. For example, whether to adopt one or more selection processes may be identified based on flag information in a header such as SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching the transform type may be a selection process different from the above three selection processes, or each of the three selection processes may be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is performed. Function [1] is a function for performing orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing orthogonal transform of the entire CU and determining the transform type based on a determined rule without encoding information indicating the transform type. Function [3] is a function for performing orthogonal transform of a partial area of ​​the CU and encoding information indicating the transform type used in the transform. Function [4] is a function for performing orthogonal transform on a partial region of a CU and determining a transform type based on a determined rule without encoding information indicating a transform type used in the transform. The determined rule may be predetermined.

[0350] It should be noted that whether to apply MTS, IMTS, and / or SBT may be determined for each processing unit, for example, for each sequence, picture, tile, slice, CTU, or CU.

[0351] It should be noted that the tool for selectively switching transform types in the present invention can be described as a method for selectively selecting a basis used in a transform process, a selection process, or a process for selecting a basis. In addition, the tool for selectively switching transform types can be described as a mode for adaptively selecting transform types.

[0352] Figure 16 is a flowchart showing one example of a process performed by the converter 106, and for convenience will be referred to Figure 7 Provide a description.

[0353] For example, the transformer 106 determines whether to perform an orthogonal transform (step St_1). Here, when it is determined that an orthogonal transform is to be performed (yes in step St_1), the transformer 106 selects a transform type for the orthogonal transform from a plurality of transform types (step St_2). Next, the transformer 106 performs the orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transform type to the entropy encoder 110, allowing the entropy encoder 110 to encode the information (step St_4). On the other hand, when it is determined that an orthogonal transform is not to be performed (no in step St_1), the transformer 106 outputs information indicating that an orthogonal transform is not to be performed, allowing the entropy encoder 110 to encode the information (step St_5). It should be noted that whether or not to perform an orthogonal transform in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transform can be performed using a defined transform type without encoding information indicating the transform type used in the orthogonal transform. The defined transform type can be predefined.

[0354] Figure 17 is a flowchart showing one example of a process performed by the converter 106, and for convenience will be referred to Figure 7 It should be noted that Figure 17 The example shown in FIG is a case where the transform type used in the orthogonal transform is selectively switched (as in Figure 16 An example of an orthogonal transform in the case of the example shown).

[0355] As an example, the first transform type group may include DCT2, DCT7, and DCT8. As another example, the second transform type group may include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group may partially overlap with each other, or may be completely different from each other.

[0356] The transformer 106 determines whether the transform size is less than or equal to the determined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the determined value (yes in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform type included in the first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one transform type included in the first transform type group to the entropy encoder 110, so as to allow the entropy encoder 110 to encode the information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to the predetermined value (no in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The determined value may be a threshold value and may be a predetermined value.

[0357] In step Su_3, the information indicating the transform type used in the orthogonal transform may be information indicating a combination of a transform type to be applied vertically in the current block and a transform type to be applied horizontally in the current block. The first type group may include only one transform type, and information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and information indicating the transform type used for the orthogonal transform from one or more transform types included in the second transform type group may be encoded.

[0358] Alternatively, the transform type may be indicated based on the transform size without encoding the information indicating the transform type. It should be noted that such determination is not limited to determining whether the transform size is less than or equal to the determined value, and other processes for determining the transform type used in the orthogonal transform based on the transform size are also possible.

[0359] (Quantizer)

[0360] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in a determined scanning order and quantizes the scanned transform coefficients based on a quantization parameter (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients of the current block (hereinafter also referred to as quantized coefficients) to the entropy encoder 110 and the inverse quantizer 112. The determined scanning order may be predetermined.

[0361] The determined scanning order is the order used to quantize / inverse quantize transform coefficients. For example, the determined scanning order can be defined as an ascending order of frequency (from low frequency to high frequency) or a descending order of frequency (from high frequency to low frequency).

[0362] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, as the value of the quantization parameter increases, the quantization step size also increases. In other words, as the value of the quantization parameter increases, the error in the quantized coefficients (quantization error) increases.

[0363] In addition, a quantization matrix can be used for quantization. For example, a variety of quantization matrices can be used corresponding to frequency transform size (e.g., 4×4, 8×8), prediction mode (e.g., intra-frame prediction, inter-frame prediction), and pixel components (e.g., luminance, chrominance pixel components). It should be noted that quantization means digitizing values ​​sampled at intervals determined corresponding to a determined level. In the art, quantization can be referred to using other expressions, such as rounding and scaling, and rounding and scaling can be employed. The determined interval and the determined level can be predetermined.

[0364] Methods for using a quantization matrix may include: a method of using a quantization matrix directly set on the encoder 100 side, and a method of using a quantization matrix that has been set as a default (default matrix). On the encoder 100 side, a quantization matrix suitable for image features can be set by directly setting the quantization matrix. However, this case may have the disadvantage of increasing the amount of code used to encode the quantization matrix. It should be noted that the quantization matrix for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.

[0365] There is a method for quantizing high-frequency coefficients and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be considered equivalent to a method using a quantization matrix (flat matrix) whose coefficients have the same value.

[0366] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level. The quantization matrix can be specified using, for example, a sequence parameter set (SPS) or a picture parameter set (PPS). The SPS includes parameters for a sequence, and the PPS includes parameters for a picture. Each of the SPS and PPS can be simply referred to as a parameter set.

[0367] When a quantization matrix is ​​used, the quantizer 108 uses the values ​​of the quantization matrix to scale the quantization width, which may be calculated based on, for example, a quantization parameter, for each transform coefficient. A quantization process performed without using a quantization matrix may be a process for quantizing the transform coefficients according to the quantization width calculated based on, for example, a quantization parameter. It should be noted that in a quantization process performed without using any quantization matrix, the quantization width may be multiplied by a predetermined value common to all transform coefficients in a block. The predetermined value may be predetermined.

[0368] Figure 181 is a block diagram showing an example of a functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a difference quantization parameter generator 108a, a predicted quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization performer 108e.

[0369] Figure 19 is a flowchart showing one example of a quantization process performed by the quantizer 108, and for convenience will be referred to as Figure 7 and 18 Provide a description.

[0370] As an example, the quantizer 108 may be based on Figure 19 The flowchart shown in FIG. 10A shows that quantization is performed for each CU. More specifically, the quantization parameter generator 108 c determines whether to perform quantization (step Sv_1). If it is determined that quantization is to be performed (yes in step Sv_1), the quantization parameter generator 108 c generates a quantization parameter for the current block (step Sv_2) and stores the quantization parameter in the quantization parameter storage device 108 d (step Sv_3).

[0371] Next, the quantization performer 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). The predicted quantization parameter generator 108b then obtains the quantization parameters of a processing unit different from the current block from the quantization parameter storage device 108d (step Sv_5). The predicted quantization parameter generator 108b generates a predicted quantization parameter for the current block based on the obtained quantization parameters (step Sv_6). The difference quantization parameter generator 108a calculates the difference between the quantization parameter of the current block generated by the quantization parameter generator 108c and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 108b (step Sv_7). The difference quantization parameter can be generated by calculating the difference. The difference quantization parameter generator 108a outputs the difference quantization parameter to the entropy encoder 110 to allow the entropy encoder 110 to encode the difference quantization parameter (step Sv_8).

[0372] It should be noted that the difference quantization parameter can be encoded at, for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level. Furthermore, the initial value of the quantization parameter can be encoded at the sequence level, the picture level, the slice level, the tile level, or the CTU level. During initialization, the quantization parameter can be generated using the initial value of the quantization parameter and the difference quantization parameter.

[0373] It should be noted that the quantizer 108 may include a plurality of quantizers and may apply dependent quantization in which a transform coefficient is quantized using a quantization method selected from a plurality of quantization methods.

[0374] (Entropy Encoder)

[0375] Figure 20 is a block diagram showing one example of the functional configuration of the entropy encoder 110 according to the embodiment, and for convenience will be referred to as Figure 7 Described. The entropy encoder 110 generates a stream by entropy encoding the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as entropy coding. More specifically, the entropy encoder 110 shown in the figure includes a binarizer 110a, a context controller 110b and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, in which multi-level signals such as quantized coefficients and prediction parameters are converted into binary signals. Examples of binarization methods include truncated Rice binarization, exponential Golomb code and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntactic element or the surrounding state (i.e., the probability of occurrence of the binary signal). Examples of methods for deriving context values ​​include bypassing, referring to syntactic elements, referring to upper and left adjacent blocks, referring to hierarchical information, etc. The binary arithmetic encoder 110c uses the derived context to perform arithmetic encoding on the binary signal.

[0376] Figure 21 1 is a conceptual diagram illustrating an example flow of a CABAC process in the entropy encoder 110. First, initialization is performed in the entropy encoder 110 using CABAC. During initialization, initialization and setting of an initial context value in the binary arithmetic encoder 110c are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c may sequentially perform binarization and arithmetic coding of multiple quantized coefficients in a CTU. Each time arithmetic coding is performed, the context controller 110b may update the context value. The context controller 110b may then save the context value as post-processing. For example, the saved context value may be used to initialize the context value of the next CTU.

[0377] (Inverse Quantizer)

[0378] The inverse quantizer 112 inversely quantizes the quantized coefficients input from the quantizer 108. More specifically, the inverse quantizer 112 inversely quantizes the quantized coefficients of the current block in a determined scanning order. The inverse quantizer 112 then outputs the inversely quantized transform coefficients of the current block to the inverse transformer 114. The determined scanning order may be predetermined.

[0379] (Inverse Converter)

[0380] The inverse transformer 114 restores the prediction residual by inversely transforming the transform coefficients input from the inverse quantizer 112. More specifically, the inverse transformer 114 restores the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the restored prediction residual to the adder 116.

[0381] It should be noted that since information is typically lost in quantization, the recovered prediction residual does not match the prediction residual calculated by the subtractor 104. In other words, the recovered prediction residual typically includes quantization errors.

[0382] (Adder)

[0383] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transformer 114 and the predicted image input from the prediction controller 128. Subsequently, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block may also be referred to as a local decoded block.

[0384] (Block Storage)

[0385] The block memory 118 is a storage device for storing blocks in the current picture used for intra prediction, for example. More specifically, the block memory 118 stores the reconstructed image output from the adder 116.

[0386] (Frame Memory)

[0387] The frame memory 122 is a storage device for storing reference pictures used in inter-frame prediction, for example, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.

[0388] (Loop Filter)

[0389] The loop filter 120 applies a loop filter to the reconstructed image output by the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter used in the encoding loop (in-loop filter). Examples of the loop filter include an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, and the like.

[0390] Figure 22 1 is a block diagram showing an example of the functional configuration of the loop filter 120 according to the embodiment. Figure 22As shown, the loop filter 120 includes a deblocking filter executor 120a, an SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs a deblocking filter process on the reconstructed image. The SAO executor 120b performs an SAO process on the reconstructed image after the deblocking filter process. The ALF executor 120c performs an ALF process on the reconstructed image after the SAO process. The ALF and deblocking filters will be described in detail later. The SAO process is a process for improving image quality by reducing ringing (a phenomenon in which pixel values ​​are distorted like waves around edges) and correcting deviations in pixel values. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that in some embodiments, the loop filter 120 may not include Figure 22 All the constituent elements disclosed in, and may include some constituent elements, and may include additional elements. In addition, the loop filter 120 may be configured to Figure 22 The above processes may be performed in a different processing order than that disclosed in the , and not all processes may be performed, etc.

[0391] (Loop Filter > Adaptive Loop Filter)

[0392] In ALF, a least square error filter is applied to remove compression artifacts. For example, a filter selected from multiple filters based on the direction and activity of the local gradient is applied to each 2×2 pixel sub-block in the current block.

[0393] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of a plurality of classes (e.g., fifteen or twenty-five classes). The classification of the sub-block can be based on, for example, gradient directionality and activity. In an example, a class index C (e.g., C=5D+A) is calculated or determined based on the gradient directionality D (e.g., 0 to 2 or 0 to 4) and the gradient activity A (e.g., 0 to 4). Then, based on the classification index C, each sub-block is classified into one of the plurality of classes.

[0394] For example, the gradient directionality D is calculated by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the gradient activity A is calculated by adding the gradients in multiple directions and quantizing the addition result.

[0395] A filter to be used for each subblock may be determined from among a plurality of filters based on such classification results.

[0396] The filter shape to be used in the ALF is, for example, a circularly symmetric filter shape. Figures 23A to 23C is a conceptual diagram for illustrating an example of a filter shape used in ALF. Figure 23A The figure shows a 5×5 diamond filter, Figure 23B A 7×7 diamond filter is shown, and Figure 23C A 9×9 diamond filter is shown. Information indicating the filter shape is typically signaled at the picture level. It should be noted that signaling of such information indicating the filter shape does not necessarily need to be performed at the picture level and can be performed at another level (e.g., at the sequence level, slice level, tile level, CTU level, or CU level).

[0397] For example, whether ALF is turned on or off may be determined at the picture level or the CU level. For example, a decision may be made at the CU level whether to apply ALF to luma, and a decision may be made at the picture level whether to apply ALF to chroma. Information indicating whether ALF is turned on or off is typically signaled at the image level or the CU level. It should be noted that the signaling of information indicating whether ALF is turned on or off does not necessarily need to be performed at the picture level or the CU level, and may be performed at another level (e.g., at the sequence level, slice level, tile level, or CTU level).

[0398] In addition, as described above, a filter is selected from a plurality of filters and the ALF process of the sub-block is performed. A coefficient set for each of the plurality of filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that the signaling of the coefficient set does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0399] (Loop Filter > Cross-Component Adaptive Loop Filter)

[0400] Figure 23D is a conceptual diagram for illustrating an example flow of cross-component ALF (CC-ALF). Figure 23E is a conceptual diagram for illustrating an example of a filter shape used in CC-ALF, such as Figure 23D CC-ALF. Figure 23D and Figure 23E An example CC-ALF operates by applying a linear diamond filter to the luma channel of each chroma component. For example, the filter coefficients can be transmitted in APS, scaled by a factor of 2^10, and rounded for fixed-point representation. For example, in Figure 23D In , the Y sample (first component) is used for CCALF for Cb and CCALF for Cr (a component different from the first component).

[0401] The application of the filter can be controlled over a variable block size and signaled via a context-coded flag received for each sample block. The block size, along with a CC-ALF enable flag, can be received at the slice level for each chroma component. CC-ALF can support various block sizes, such as (in chroma samples) 16×16 pixels, 32×32 pixels, 64×64 pixels, 128×128 pixels.

[0402] (Loop Filter > Joint Chroma Cross-Component Adaptive Loop Filter)

[0403] An example of joint chroma-CCALF is given in Figure 23F and Figure 23G Shown in. Figure 23F is a conceptual diagram illustrating an example flow of joint chroma CCALF. Figure 23G is a table showing example weight index candidates. As shown, a CCALF filter is used to generate a CCALF filtered output as a chroma refinement signal for one color component, while applying a weighted version of the same chroma refinement signal to another color component. In this way, the complexity of the existing CCALF is reduced by about half. The weight value can be encoded as a sign flag and a weight index. The weight index (denoted as weight_index) can be encoded into 3 bits and specifies the size of the JC-CCALF weight JcCcWeight, which is a non-zero size. For example, the size of JcCcWeight can be determined as follows:

[0404] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index>>2;

[0405] Otherwise, JcCcWeight is equal to 4 / (weight_index–4).

[0406] The block-level on / off control of the ALF filters for Cb and Cr can be separated. This is the same as in CCALF, and two separate sets of block-level on / off control flags can be encoded. Unlike CCALF, the Cb and Cr on / off control block sizes are the same, so only one block size variable can be encoded.

[0407] (Loop filter > Deblocking filter)

[0408] In the deblocking filtering process, the loop filter 120 performs a filtering process on block boundaries in the reconstructed image in order to reduce distortion occurring at the block boundaries.

[0409] Figure 24 The loop filter 120 (see FIG. Figure 7 and Figure 22) is a block diagram of an example of a specific configuration of the deblocking filter executor 120a.

[0410] The deblocking filter performer 120 a includes: a boundary determiner 1201 ; a filter determiner 1203 ; a filter performer 1205 ; a process determiner 1208 ; a filter characteristic determiner 1207 ; and switches 1202 , 1204 , and 1206 .

[0411] The boundary determiner 1201 determines whether a pixel to be deblocking filtered (ie, a current pixel) exists around a block boundary. The boundary determiner 1201 then outputs the determination result to the switch 1202 and the processing determiner 1208.

[0412] In the case where the boundary determiner 1201 determines that the current pixel exists around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1204. In the opposite case (where the boundary determiner 1201 determines that the current pixel does not exist around the block boundary), the switch 1202 outputs the unfiltered image to the switch 1206. Note that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.

[0413] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel value of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.

[0414] In the case where the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained by the switch 1202 to the filter executor 1205. In the opposite case (where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel), the switch 1204 outputs the unfiltered image obtained by the switch 1202 to the switch 1206.

[0415] When an unfiltered image is obtained through switches 1202 and 1204 , filter executor 1205 performs deblocking filtering on the current pixel with the filter characteristics determined by filter characteristic determiner 1207 . Filter executor 1205 then outputs the filtered pixel to switch 1206 .

[0416] Under the control of the processing determiner 1208 , the switch 1206 selectively outputs one of the pixels that have not been deblocking filtered and the pixels that have been deblocking filtered by the filtering executor 1205 .

[0417] The processing determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary and when the filter determiner 1203 has determined that deblocking filtering of the current pixel is to be performed, the processing determiner 1208 causes the switch 1207 to output the pixel that has been subjected to deblocking filtering. In addition, in addition to the above-mentioned case, the processing determiner 1208 causes the switch 1206 to output the pixel that has not been subjected to deblocking filtering. By repeating the output of the pixels in this manner, the filtered image is output from the switch 1206. It should be noted that Figure 24 The configuration shown in is one example of the configuration in the deblocking filter performer 120a. The deblocking filter performer 120a may have various configurations.

[0418] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having symmetric filtering characteristics with respect to a block boundary.

[0419] In the deblocking filtering process, pixel values ​​and quantization parameters can be used to select one of two deblocking filters (i.e., a strong filter and a weak filter) with different characteristics. In the case of a strong filter, when pixels p0 to p2 and pixels q0 to q2 exist across block boundaries, such as Figure 25 As shown, by performing calculation according to the following expressions, for example, the pixel values ​​of the respective pixels q0 to q2 are changed to pixel values ​​q'0 to q'2.

[0420] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8

[0421] q'1=(p0+q0+q1+q2+2) / 4

[0422] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0423] Note that in the above expressions, p0 to p2 and q0 to q2 are the pixel values ​​of the corresponding pixels p0 to p2 and q0 to q2. Furthermore, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of pixel q2 relative to the block boundary. Furthermore, on the right side of each expression, the coefficients multiplied by the corresponding pixel values ​​of the pixels to be used for deblocking filtering are filter coefficients.

[0424] Furthermore, during deblocking filtering, clipping can be performed so that the calculated pixel value does not vary by more than a threshold. For example, during clipping, the pixel value calculated according to the above expression can be clipped to a value obtained by "calculated pixel value ± 2 × threshold" (using a threshold determined based on a quantization parameter). This prevents oversmoothing.

[0425] Figure 26 is a conceptual diagram for illustrating block boundaries on which a deblocking filtering process is performed. Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.

[0426] The block boundary on which the deblocking filtering process is performed is, for example, a boundary between CUs, Pus, or TUs having 8×8 pixel blocks, such as Figure 26 As shown. The deblocking filtering process can be performed in units of four rows or four columns, for example. First, as Figure 27 As shown for block P and block Q ( Figure 26 ) to determine the boundary strength (Bs) value.

[0427] according to Figure 27 The Bs value in the image can determine whether to perform a deblocking filtering process with different intensities on block boundaries belonging to the same image. When the Bs value is 2, a deblocking filtering process for a chroma signal is performed. When the Bs value is 1 or greater and a determined condition is satisfied, a deblocking filtering process for a luminance signal is performed. The determined condition may be predetermined. Note that the condition for determining the Bs value is not limited to Figure 27 , and the Bs value can be determined based on another parameter.

[0428] (Predictor (Intra Predictor, Inter Predictor, Prediction Controller))

[0429] Figure 28 1 is a flowchart showing one example of a process performed by the predictor of the encoder 100. Note that the predictor includes all or part of the following constituent elements: an intra predictor 124; an inter predictor 126; and a prediction controller 128. The prediction controller includes, for example, the intra predictor 124 and the inter predictor 126.

[0430] The predictor generates a predicted image for the current block (step Sb_1). This predicted image may also be referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra-frame predicted image (image prediction signal) or an inter-frame predicted image (inter-frame prediction signal). The predictor generates a predicted image for the current block using a reconstructed image obtained by generating a predicted image, generating a prediction residual, generating quantized coefficients, restoring the prediction residual, and adding it to the predicted image for another block.

[0431] The reconstructed image may be, for example, an image in a reference picture, or an image of a coding block (ie, the other blocks mentioned above) in a current picture, where the current picture is a picture including the current block. The coding block in the current picture may be, for example, a neighboring block of the current block.

[0432] Figure 29 is a flowchart illustrating another example of a process performed by the predictor of the encoder 100 .

[0433] The predictor generates a predicted image using a first method (step Sc_1a), generates a predicted image using a second method (step Sc_1b), and generates a predicted image using a third method (step Sc_1c). The first method, the second method, and the third method may be different methods for generating predicted images. Each of the first to third methods may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.

[0434] Next, the prediction processor evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates a cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the predicted images by comparing the costs C of the predicted images. It should be noted that the cost C can be calculated, for example, according to an expression of the RD optimization model, such as C=D+λ×R. In this expression, D represents a compression artifact of the predicted image, and is expressed as, for example, the sum of the absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. In addition, R represents the bit rate of the stream. In addition, λ represents a multiplier, for example, according to the Lagrangian method multiplier.

[0435] The predictor then selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects a method or mode for obtaining the final predicted image. For example, the predictor selects a predicted image with a minimum cost C based on the cost C calculated for the predicted image. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be performed based on parameters used in the encoding process. The encoder 100 may transform information for identifying the selected predicted image, method, or mode into a stream. The information may be, for example, a flag, etc. In this way, the decoder 200 is able to generate a predicted image according to the method or mode selected by the encoder 100 based on the information. It should be noted that in Figure 29 In the example shown, the predictor selects any predicted image after generating the predicted image using the corresponding method. However, the predictor may select a method or mode based on the parameters used in the encoding process before generating the predicted image, and may generate the predicted image according to the selected method or mode.

[0436] For example, the first method and the second method may be intra prediction and inter prediction, respectively, and the predictor may select a final predicted image of the current block from predicted images generated according to the prediction methods.

[0437] Figure 30 is a flowchart illustrating another example of a process performed by the predictor of the encoder 100 .

[0438] First, the predictor generates a predicted image using intra-frame prediction (step Sd_1a) and generates a predicted image using inter-frame prediction (step Sd_1b). It should be noted that the predicted image generated by intra-frame prediction is also called an intra-frame predicted image, and the predicted image generated by inter-frame prediction is also called an inter-frame predicted image.

[0439] Next, the predictor evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). The above-mentioned cost C can be used in the evaluation. The predictor can then select the prediction image for which the minimum cost C has been calculated from the intra-frame prediction image and the inter-frame prediction image as the final prediction image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the prediction image for the current block is selected.

[0440] The prediction processor then selects the prediction image for which the minimum cost C has been calculated among the intra-frame prediction image and the inter-frame prediction image as the final prediction image of the current block (step Sd_3). In other words, the prediction method or mode for generating the prediction image of the current block is selected.

[0441] (Intra-frame predictor)

[0442] The intra predictor 124 generates a prediction signal (i.e., an intra-prediction image) by performing intra prediction (also referred to as intra prediction) of the current block with reference to one or more blocks in the current picture and stored in the block memory 118. More specifically, the intra predictor 124 generates an intra-prediction image by performing intra prediction with reference to pixel values ​​(e.g., luminance and / or chrominance values) i of one or more blocks adjacent to the current block, and then outputs the intra-prediction image to the prediction controller 128.

[0443] For example, the intra-frame predictor 124 performs intra-frame prediction by using one of a plurality of defined intra-frame prediction modes. The intra-frame prediction modes generally include one or more non-directional prediction modes and a plurality of directional prediction modes. The defined modes may be predefined.

[0444] The one or more non-directional prediction modes include, for example, planar prediction mode and DC prediction mode as defined in the H.265 / High Efficiency Video Coding (HEVC) standard.

[0445] The plurality of directional prediction modes include, for example, thirty-three directional prediction modes defined in the H.265 / HEVC standard. It should be noted that, in addition to the thirty-three directional prediction modes, the plurality of directional prediction modes may also include thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31: is a conceptual diagram showing a total of sixty-seven intra prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra prediction. Solid arrows represent thirty-three directions defined in the H.265 / HEVC standard, and dotted arrows represent additional thirty-two directions ( Figure 31 The two non-directional prediction modes are not shown).

[0446] In various processing examples, a luma block may be referenced in intra prediction of a chroma block. In other words, the chroma component of a current block may be predicted based on the luma component of the current block. This intra prediction is also known as cross-component linear model (CCLM) prediction. An intra prediction mode for a chroma block that references this luma block (also known as a CCLM mode, for example) may be added as one of the intra prediction modes for a chroma block.

[0447] The intra predictor 124 can correct the pixel values ​​of the intra prediction based on the horizontal / vertical reference pixel gradients. Intra prediction with such correction is also called position-dependent intra prediction combination (PDPC). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0448] Figure 32 is a flowchart illustrating one example of a process performed by the intra predictor 124 .

[0449] The intra-frame predictor 124 selects an intra-frame prediction mode from a plurality of intra-frame prediction modes (step Sw_1). The intra-frame predictor 124 then generates a predicted image based on the selected intra-frame prediction mode (step Sw_2). Next, the intra-frame predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra-frame prediction modes. For example, two of the six intra-frame prediction modes may be a planar mode and a DC prediction mode, and the other four modes may be directional prediction modes. The intra-frame predictor 124 determines whether the intra-frame prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0450] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the intra prediction mode selected in these MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110.

[0451] When it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flag. The intra predictor 124 then generates information indicating the intra prediction mode selected from at least one intra prediction mode not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110. The information indicating the intra prediction mode indicates, for example, any one of 0 to 60.

[0452] (Inter-frame predictor)

[0453] The inter-frame predictor 126 generates a predicted image (inter-frame prediction image) by performing inter-frame prediction (also called inter-frame prediction) of the current block with reference to one or more blocks in a reference picture that is different from the current picture and stored in the frame memory 122. Inter-frame prediction is performed in units of the current block or the current sub-block (e.g., 4×4 block) in the current block. A sub-block is included in a block and is a unit smaller than a block. The size of the sub-block can be in the form of a slice, a brick, a picture, etc.

[0454] For example, the inter-frame predictor 126 performs motion estimation in the reference picture of the current block or current sub-block and finds the reference block or reference sub-block that best matches the current block or current sub-block. The inter-frame predictor 126 then obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or reference sub-block to the current block or sub-block. The inter-frame predictor 126 generates an inter-frame predicted image for the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter-frame predictor 126 outputs the generated inter-frame predicted image to the prediction controller 128.

[0455] The motion information used in motion compensation can be signaled in various forms as an inter-frame prediction signal. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a motion vector predictor can be signaled.

[0456] (Reference picture list)

[0457] Figure 33 is a conceptual diagram for illustrating an example of a reference picture. Figure 34 1 is a conceptual diagram for illustrating an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. Figure 33In the figure, each rectangle represents a picture, each arrow represents a picture reference relationship, the horizontal axis represents time, I, P and B in the rectangle represent intra-frame prediction pictures, single-prediction pictures and double-prediction pictures respectively, and the numbers in the rectangle represent the decoding order. Figure 33 As shown in , the decoding order of pictures is in the order of I0, P1, B2, B3, B4, and the display order of pictures is in the order of I0, B3, B2, B4, P1. Figure 34 As shown, a reference picture list is a list representing reference picture candidates. For example, a picture (or slice) may include at least one reference picture list. For example, when the current picture is a uni-predicted picture, one reference picture list is used, and when the current picture is a bi-predicted picture, two reference picture lists are used. Figure 33 and Figure 34 In the example of , picture B3, which is the current picture currPic, has two reference picture lists, namely, the L0 list and the L1 list. When the current picture currPic is picture B3, the reference picture candidates of the current picture currPic are I0, P1, and B2, and the reference picture lists (i.e., the L0 list and the L1 list) indicate these pictures. The inter-frame predictor 126 or the prediction controller 128 specifies which picture in each reference picture list to actually reference in the form of a reference picture index refidxLx. Figure 34 , reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.

[0458] Such a reference picture list can be generated for each unit such as a sequence, picture, slice, block, CTU, or CU. In addition, among the reference pictures indicated in the reference picture list, a reference picture index indicating a reference picture to be referenced in inter prediction can be signaled at the sequence level, picture level, slice level, tile level, CTU level, or CU level. In addition, a common reference picture list can be used in multiple inter prediction modes.

[0459] (Basic process of inter-frame prediction)

[0460] Figure 35 is a flowchart illustrating an example basic processing flow of an inter-frame prediction process.

[0461] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4).

[0462] Here, when generating a predicted image, the inter-frame predictor 126 determines the motion vector (MV) of the current block (steps Se_1 and Se_2) and performs motion compensation (step Se_3). Furthermore, when determining the MV, the inter-frame predictor 126 determines the MV by selecting motion vector candidates (MV candidates) (step Se_1) and deriving the MV (step Se_2). MV candidate selection is performed, for example, by the inter-frame predictor 126 generating an MV candidate list and selecting at least one MV candidate from the MV candidate list. It should be noted that previously derived MVs can be added to the MV candidate list. Alternatively, when deriving the MV, the inter-frame predictor 126 can select at least one MV candidate from the at least one MV candidate and determine the selected at least one MV candidate as the MV for the current block. Alternatively, the inter-frame predictor 126 can determine the MV for the current block by performing estimation in a reference picture region specified by each of the at least one selected MV candidates. It should be noted that estimation in a reference picture region can be referred to as motion estimation.

[0463] In addition, although steps Se_1 to Se_3 are performed by the inter predictor 126 in the above-described example, processes such as step Se_1 , step Se_2 , and the like may be performed by another constituent element included in the encoder 100 .

[0464] It should be noted that the MV candidate list can be generated for each process in the inter-frame prediction mode, or a common MV candidate list can be used in multiple inter-frame prediction modes. The processes in steps Se_3 and Se_4 correspond to Figure 9 The process in step Se_3 corresponds to Figure 30 The process in step Sd_1b in .

[0465] (Motion vector derivation process)

[0466] Figure 36 is a flowchart showing an example of a derivation process of a motion vector.

[0467] The inter-frame predictor 126 can derive the MV of the current block in a mode for encoding motion information (e.g., MV). In this case, for example, the motion information can be encoded as a prediction parameter and can be signaled. In other words, the encoded motion information is included in the stream.

[0468] Alternatively, the inter predictor 126 may derive the MV in a mode where motion information is not encoded. In this case, motion information is not included in the stream.

[0469] Here, the MV derivation mode may include the conventional inter mode, conventional merge mode, FRUC mode, affine mode, etc., which will be described later. The modes for encoding motion information include the conventional inter mode, conventional merge mode, affine mode (specifically, affine inter mode and affine merge mode), etc. It should be noted that the motion information may include not only the MV but also the motion vector predictor selection information described later. Modes that do not encode motion information include the FRUC mode, etc. The inter predictor 126 selects a mode for deriving the MV of the current block from a plurality of modes and derives the MV of the current block using the selected mode.

[0470] Figure 37 is a flowchart illustrating another example of derivation of motion vectors.

[0471] The inter-frame predictor 126 can derive the MV of the current block in a mode that encodes the MV difference. In this case, for example, the MV difference can be encoded as a prediction parameter and can be signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. It should be noted that the MV predictor is a motion vector predictor.

[0472] Alternatively, the inter-frame predictor 126 may derive the MV in a mode where the MV difference is not encoded. In this case, the encoded MV difference is not included in the stream.

[0473] Here, as described above, the MV derivation mode includes the normal inter mode, normal merge mode, FRUC mode, affine mode, and the like described later. Modes in which the MV difference is encoded include the normal inter mode, affine mode (specifically, affine inter mode), and the like. Modes in which the MV difference is not encoded include the FRUC mode, normal merge mode, affine mode (specifically, affine merge mode), and the like. The inter predictor 126 selects a mode for deriving the MV of the current block from the plurality of modes, and derives the MV of the current block using the selected mode.

[0474] (Motion vector derivation mode)

[0475] Figure 38A and Figure 38B is a conceptual diagram for illustrating an example classification of patterns for MV derivation. Figure 38A As shown in FIG, the MV derivation mode is roughly classified into three modes according to whether motion information is encoded and whether MV difference is encoded. The three modes are inter-frame mode, merge mode, and frame rate up conversion (FRUC) mode. Inter-frame mode is a mode that performs motion estimation and encodes motion information and MV difference. For example, Figure 38BAs shown, inter-frame mode includes affine inter-frame mode and normal inter-frame mode. Merge mode is a mode in which motion estimation is not performed and MV is selected from the coded surrounding blocks and the MV of the current block is derived using the MV. Merge mode is a mode in which motion information is basically encoded without encoding the MV difference. For example, Figure 38B As shown, the merge mode includes a regular merge mode (also called a regular merge mode or a normal merge mode), a merge with motion vector difference (MMVD) mode, a combined inter merge / intra prediction (CIIP) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, in the MMVD mode among the modes included in the merge mode, the MV difference is encoded as an exception. It should be noted that the affine merge mode and the affine inter mode are modes included in the affine mode. The affine mode is a mode for deriving the MV of each of the multiple sub-blocks included in the current block as the MV of the current block assuming an affine transformation. The FRUC mode is a mode for deriving the MV of the current block by performing estimation between coding regions, and neither encoding motion information nor encoding any MV difference. It should be noted that the corresponding modes will be described in more detail later.

[0476] It should be noted that Figure 38A and Figure 38B The classification of modes shown in is an example, and the classification is not limited thereto. For example, when the MV difference is encoded in the CIIP mode, the CIIP mode is classified as the inter-frame mode.

[0477] (MV derivation > conventional inter-frame mode)

[0478] The normal inter mode is an inter prediction mode for deriving the MV of the current block from a reference picture region specified by an MV candidate based on a block similar to the image of the current block. In this normal inter mode, the MV difference is encoded.

[0479] Figure 39 is a flowchart illustrating an example of an inter prediction process in a conventional inter mode.

[0480] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as MVs of multiple coding blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter-frame predictor 126 generates an MV candidate list.

[0481] Next, the inter-frame predictor 126 extracts N (an integer of 2 or greater) MV candidates from the plurality of MV candidates obtained in step Sg_1 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the determined priority order (step Sg_2). It should be noted that the priority order may be predetermined for each of the N MV candidates.

[0482] Next, the inter-frame predictor 126 selects a motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor (also called MV predictor) for the current block (step Sg_3). At this time, the inter-frame predictor 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor in the stream. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.

[0483] Next, the inter-frame predictor 126 derives the MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as an MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130. It should be noted that the encoded reference picture is a picture that includes multiple blocks that have been reconstructed after being encoded.

[0484] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image for the current block (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in the slice, inter-frame prediction for the slice using the conventional inter-frame mode is completed. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in the picture, inter-frame prediction for the picture using the conventional inter-frame mode is completed. It should be noted that in steps Sg_1 to Sg_5, not all blocks included in the slice may undergo these processes, and when some blocks undergo the processes, inter-frame prediction for the slice using the conventional inter-frame mode may be completed. This also applies to the processes in steps Sg_1 to Sg_5. When the processes are performed for some blocks in the picture, inter-frame prediction for the picture using the conventional inter-frame mode may be completed.

[0485] It should be noted that the predicted image is the inter prediction signal as described above. In addition, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the predicted image is encoded as a prediction parameter in the encoded signal, for example.

[0486] It should be noted that the MV candidate list can also be used as a list for use in another mode. In addition, the processes related to the MV candidate list can be applied to the processes related to the list for use in another mode. The processes related to the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering MV candidates, or deleting MV candidates.

[0487] (MV derivation > conventional merge mode)

[0488] Normal merge mode is an inter-frame prediction mode used to select an MV candidate from an MV candidate list as the MV of the current block, thereby deriving the MV. It should be noted that normal merge mode is a type of merge mode and can be simply referred to as merge mode. In this embodiment, normal merge mode and merge mode are distinguished, and merge mode is used in a broader sense.

[0489] Figure 40 is a flowchart illustrating an example of inter prediction in normal merge mode.

[0490] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as MVs of multiple coding blocks temporally or spatially surrounding the current block (step Sh_1). In other words, the inter-frame predictor 126 generates an MV candidate list.

[0491] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter-frame predictor 126 encodes MV selection information used to identify the selected MV candidate in the stream. In other words, the inter-frame predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.

[0492] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image for the current block (step Sh_3). For example, the process in steps Sh_1 to Sh_3 is performed on each block. For example, when the process in steps Sh_1 to Sh_3 is performed on all blocks in the slice, the inter-frame prediction of the slice using the normal merge mode is completed. In addition, when the process in steps Sh_1 to Sh_3 is performed on all blocks in the picture, the inter-frame prediction of the picture using the normal merge mode is completed. It should be noted that not all blocks included in the slice can undergo the process of steps Sh_1 to Sh_3, and when some blocks undergo the process, the inter-frame prediction of the slice using the normal merge mode can be completed. This also applies to the process in steps Sh_1 to Sh_3. When the process is performed on some blocks in the picture, the inter-frame prediction of the picture using the normal merge mode can be completed.

[0493] In addition, information indicating the inter prediction mode (normal merge mode in the above example) contained in the encoded signal and used to generate the predicted image is encoded as a prediction parameter in the stream, for example.

[0494] Figure 41 is a conceptual diagram for illustrating an example of a motion vector derivation process of a current picture through a normal merge mode.

[0495] First, the inter-frame predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of multiple coding blocks located spatially around the current block; temporally adjacent MV candidates, which are MVs of surrounding blocks onto which the position of the current block in the encoded reference picture is projected; combined MV candidates, which are MVs generated by combining the MV values ​​of spatially adjacent MV predictors and the MV values ​​of temporally adjacent MV predictors; and zero MV candidates, which are MVs having a value of zero.

[0496] Next, the inter predictor 126 selects one MV candidate from the plurality of MV candidates registered in the MV candidate list, and determines the MV candidate as the MV of the current block.

[0497] Furthermore, the entropy encoder 110 writes and encodes merge_idx, which is a signal indicating which MV candidate has been selected, in the stream.

[0498] It should be noted that registration Figure 41 The MV candidates in the MV candidate list described in the figure are examples. The number of MV candidates may be different from the number of MV candidates in the figure, and the MV candidate list may be configured in such a way that some types of MV candidates in the figure are not included, or one or more types of MV candidates other than the types of MV candidates in the figure are included.

[0499] The final MV can be determined by performing dynamic motion vector refresh (DMVR) to be described later using the MV of the current block derived by the normal merge mode. It should be noted that in the normal merge mode, motion information is encoded and no MV difference is encoded. In the MMVD mode, an MV candidate is selected from the MV candidate list and the MV difference is encoded just like in the case of the normal merge mode. Figure 38B As shown, MMVD can be classified as a merge mode along with the regular merge mode. It should be noted that the MV difference in MMVD mode does not always need to be the same as the MV difference used for inter mode. For example, MV difference derivation in MMVD mode can be a process that requires less processing than the MV difference derivation in inter mode.

[0500] In addition, a combined inter-frame merging / intra-frame prediction (CIIP) mode may be performed. This mode is used to overlap a prediction image generated in inter-frame prediction and a prediction image generated in intra-frame prediction to generate a prediction image of a current block.

[0501] It should be noted that the MV candidate list can be referred to as a candidate list. In addition, merge_idx is MV selection information.

[0502] (MV derivation > HMVP mode)

[0503] Figure 42 is a conceptual diagram for illustrating an example of an MV derivation process for a current picture using the HMVP merge mode.

[0504] In conventional merge mode, the MV of the CU, for example, the current block, is determined by selecting an MV candidate from the MV list generated from the reference coding block (e.g., CU). Here, another MV candidate can be registered in the MV candidate list. The mode of registering such another MV candidate is called HMVP mode.

[0505] In HMVP mode, the MV candidates are managed using the HMVP's First-In-First-Out (FIFO) server, separate from the MV candidate list of the regular merge mode.

[0506] In the FIFO buffer, the latest motion information such as the MV of the block processed in the past is stored first. In the management FIFO buffer, each time a block is processed, the MV of the latest block (i.e., the CU processed immediately before) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the earliest processed CU) is deleted from the FIFO buffer. Figure 42 In the example shown, HMVP1 is the MV of the newest block, and HMVP5 is the MV of the oldest MV.

[0507] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is an MV different from all MV candidates already registered in the MV candidate list for the normal merge mode starting from HMVP1. If it is determined that the MV is different from all MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for the normal merge mode as an MV candidate. At this time, one or more MV candidates in the FIFO buffer can be registered (added to the MV candidate list).

[0508] By using the HMVP mode in this way, not only the MVs of blocks adjacent to the current block in space or time can be added, but also the MVs of blocks processed in the past can be added. As a result, the variety of MV candidates for the conventional merge mode is expanded, which increases the possibility of improving coding efficiency.

[0509] Note that MV may be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer may include not only MV values ​​but also reference picture information, reference direction, number of pictures, etc. In addition, a block may be, for example, a CU.

[0510] Notice, Figure 42 The MV candidate list and FIFO buffer shown in are examples. The size of the MV candidate list and FIFO buffer can be Figure 42 or can be configured to be different from Figure 42 Furthermore, the process described herein may be common between the encoder 100 and the decoder 200 .

[0511] Note that the HMVP mode can be applied to modes other than the regular merge mode. For example, motion information such as the MV of a block processed in the past in the affine mode can also be stored recently and used as an MV candidate, which can promote better efficiency. The mode obtained by applying the HMVP mode to the affine mode can be called the historical affine mode.

[0512] (MV derivation > FRUC mode)

[0513] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, motion estimation is performed on the decoder side without using any pixel values ​​in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values ​​in the current block include frame rate up conversion (FRUC) mode, pattern matching motion vector derivation (PMMVD) mode, and the like.

[0514] Figure 43 An example of a FRUC process in the form of a flowchart is shown in FIG. First, MVs of coding blocks each of which is spatially or temporally adjacent to the current block are indicated by reference to the MVs as a list of MV candidates (this list can be an MV candidate list and can also be used as an MV candidate list for a regular merge mode) (step Si_1).

[0515] Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation values ​​of the corresponding MV candidates included in the MV candidate list are calculated, and an MV candidate is selected based on the evaluation value. Based on the selected motion vector candidate, the motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector of the current block. In addition, for example, the motion vector of the current block can be derived using pattern matching in the surrounding area of ​​the position in the reference picture, where the position in the reference picture corresponds to the selected motion vector candidate. In other words, estimation using pattern matching and evaluation values ​​can be performed in the surrounding area of ​​the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, updating of the motion vector that produces a better evaluation value may not be performed.

[0516] Finally, by performing motion compensation for the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image for the current block (step Si_5). For example, the process in steps Si_1 to Si_5 is performed on each block. For example, when the process in steps Si_1 to Si_5 is performed on all blocks in the slice, the inter-frame prediction of the slice using the FRUC mode is completed. For example, when the process in steps Si_1 to Si_5 is performed on all blocks in the picture, the inter-frame prediction of the picture using the FRUC mode is completed. Note that not all blocks included in the slice undergo the process in steps Si_1 to Si_5, and when some blocks undergo the process, the inter-frame prediction of the slice using the FRUC mode can be completed. When the process in steps Si_1 to Si_5 is performed on some blocks included in the picture in a similar manner, the inter-frame prediction of the picture using the FRUC mode can be completed.

[0517] A similar process may be performed in units of sub-blocks.

[0518] The evaluation value can be calculated using various methods. For example, a comparison can be made between a reconstructed image in a region in a reference picture corresponding to a motion vector and a reconstructed image in a determined region (such as a region in another reference picture or a region in an adjacent block of the current picture, as described below). The determined region can be predetermined.

[0519] The difference between the pixel values ​​of the two reconstructed images can be used for the evaluation value of the motion vector. Note that the evaluation value can be calculated using information other than the difference value.

[0520] Next, an example of pattern matching is described in detail. First, a MV candidate included in a MV candidate list (e.g., a merged list) is selected as the starting point for estimation through pattern matching. For example, first pattern matching or second pattern matching can be used as pattern matching. First pattern matching and second pattern matching can be referred to as bilateral matching and template matching, respectively.

[0521] (MV derivation > FRUC > bilateral matching)

[0522] In the first pattern matching, pattern matching is performed between two blocks located along the motion trajectory of the current block and included in two different reference pictures. Therefore, in the first pattern matching, a region in another reference picture along the motion trajectory of the current block is used as a specific region for calculating the candidate evaluation value. This specific region may be predetermined.

[0523] Figure 44 : is a conceptual diagram for illustrating an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. Figure 44 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the best matching pair among the pairs of two blocks included in two different reference pictures (Ref0, Ref1) and located along the motion trajectory of the current block (Cur block). More specifically, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the MV candidate and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by the symmetric MV obtained by scaling the MV candidate at the display time interval is derived for the current block, and an evaluation value is calculated using the obtained difference. The MV candidate that produces the best evaluation value and is likely to produce good results can be selected as the final MV from among multiple MV candidates.

[0524] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) of the two reference blocks are specified to be proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, a mirror-symmetric bidirectional motion vector is derived in the first pattern matching.

[0525] (MV derivation > FRUC > template matching)

[0526] In the second pattern matching (template matching), pattern matching is performed between a block in a reference picture and a template in a current picture (a template is a block in the current picture that is adjacent to the current block (an adjacent block is, for example, an upper and / or left adjacent block)). Therefore, in the second pattern matching, the adjacent blocks of the current block in the current picture are used as a determined area for calculating the evaluation value of the above-mentioned MV candidate.

[0527] Figure 45 is a conceptual diagram for illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 45 As shown, in the second mode matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference picture (Ref0) that best matches the neighboring block of the current block in the current picture (Cur Pic). More specifically, the difference between the reconstructed image in the coding area adjacent to the left and above or to the left or above and the reconstructed image in the corresponding area in the coding reference picture (Ref0) and specified by the MV candidate is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.

[0528] Such information indicating whether the FRUC mode is applied (e.g., referred to as a FRUC flag) may be signaled at the CU level. Furthermore, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (e.g., first pattern matching or second pattern matching) may be signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and may be performed at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0529] (MV Derivation > Affine Mode)

[0530] Affine mode uses affine transformation to generate MVs. For example, MVs can be derived in sub-block units based on the motion vectors of multiple adjacent blocks. This mode is also called affine motion compensation prediction mode.

[0531] Figure 46A : is a conceptual diagram for illustrating an example of MV derivation in sub-block units based on motion vectors of multiple adjacent blocks. Figure 46A In the example, the current block includes sixteen 4×4 sub-blocks. Here, the motion vector V0 at the upper left control point of the current block is derived based on the motion vectors of the adjacent blocks, and similarly, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of the adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (vx ,v y ).

[0532] [Mathematical expression 1]

[0533]

[0534] Such information indicating the affine mode (e.g., called an affine flag) can be signaled at the CU level. Note that the signaling of the information indicating the affine mode does not necessarily need to be performed at the CU level and can be performed at another level (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0535] In addition, the affine mode may include several modes for different methods of deriving motion vectors at the upper left and upper right control points. For example, the affine mode includes two modes: an affine inter mode (also called an affine regular inter mode) and an affine merge mode.

[0536] (MV Derivation > Affine Mode)

[0537] Figure 46B : is a conceptual diagram for illustrating an example of MV derivation in units of sub-blocks in an affine mode in which three control points are used. Figure 46B In the example, the current block includes sixteen 4×4 blocks. Here, the motion vector V0 at the upper left control point in the current block is derived based on the motion vectors of the adjacent blocks. Here, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of the adjacent blocks, and similarly, the motion vector V2 at the lower left control point of the current block is derived based on the motion vectors of the adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v x ,v y ).

[0538] [Mathematical expression 2]

[0539]

[0540] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w and h may be weighting coefficients, which may be predetermined weighting coefficients. In an embodiment, w may represent the width of the current block, and h may represent the height of the current block.

[0541] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0542] Furthermore, the affine mode using three control points may include different methods for deriving motion vectors at the upper left, upper right, and lower left control points. For example, as in the case of the affine mode using two control points, the affine mode using three control points may include two modes: an affine inter mode and an affine merge mode.

[0543] Note that in the affine mode, the size of each sub-block included in the current block may not be limited to 4×4 pixels, and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.

[0544] (MV Derivation > Affine Mode > Control Points)

[0545] Figure 47A 、 Figure 47B and Figure 47C is a conceptual diagram for illustrating an example of MV derivation at a control point in an affine mode.

[0546] like Figure 47A As shown, in affine mode, for example, based on a plurality of motion vectors corresponding to blocks encoded according to the affine mode among coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block, a motion vector predictor at a corresponding control point of the current block is calculated. More specifically, coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in the order listed, and the first valid block encoded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on the plurality of motion vectors corresponding to the identified blocks.

[0547] For example, Figure 47B As shown in FIG, when block A adjacent to the left side of the current block has been encoded according to the affine mode using two control points, motion vectors v3 and v4 projected at the upper left corner position and the upper right corner position of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated.

[0548] For example, Figure 47CAs shown in FIG, when block A adjacent to the left side of the current block has been encoded according to the affine mode using three control points, motion vectors v3, v4, and v5 projected at the upper left corner position, the upper right corner position, and the lower left corner position of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, motion vector v0 at the upper left control point of the current block, motion vector v1 at the upper right control point of the current block, and motion vector v2 at the lower left control point of the current block are calculated.

[0549] Figures 47A to 47C The MV derivation method shown can be found in Figure 50 The MV of each control point of the current block in step Sk_1 shown in FIG. 1 is used for derivation, or can be used for the MV of each control point of the current block described later. Figure 51 The MV predictor at each control point of the current block is derived in step Sj_1 as shown.

[0550] Figure 48A and Figure 48B is a conceptual diagram for illustrating an example of MV derivation at a control point in an affine mode.

[0551] Figure 48A is a conceptual diagram for illustrating an exemplary affine pattern using two control points.

[0552] In affine mode, such as Figure 48A As shown in FIG, an MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, an MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block.

[0553] Figure 48B is a conceptual diagram for illustrating an exemplary affine pattern using three control points.

[0554] In affine mode, such as Figure 48B As shown in FIG, the MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, the MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block. Furthermore, the MV selected from the MVs of the coded blocks F and G adjacent to the current block is used as the motion vector v2 at the lower left corner control point of the current block.

[0555] Notice, Figure 48A and Figure 48B The MV derivation method shown can be used for the later described Figure 50 In the MV derivation at each control point of the current block in step Sk_1 shown in FIG, or can be used for the MV derivation described later Figure 51 The MV predictor at each control point of the current block is derived in step Sj_1 shown in .

[0556] Here, when affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level, the number of control points of the coding block and the number of control points of the current block can be different from each other.

[0557] Figure 49A and Figure 49B is a conceptual diagram for illustrating an example of a method of MV derivation at a control point when the number of control points of a coding block and the number of control points of a current block are different from each other.

[0558] For example, Figure 49A As shown in the figure, the current block has three control points at the upper left, upper right, and lower left corners, and block A, which is adjacent to the left side of the current block, has been encoded according to an affine mode using two control points. In this case, motion vectors v3 and v4 projected at the upper left and upper right corners of the coded block including block A are derived. Then, motion vector v0 at the upper left control point and motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3 and v4. In addition, motion vector v2 at the lower left control point is calculated based on the derived motion vectors v0 and v1.

[0559] For example, Figure 49B As shown in FIG. 1 , the current block has two control points at the upper left and upper right corners, and block A adjacent to the left of the current block has been encoded according to an affine mode using three control points. In this case, motion vectors v3, v4, and v5 projected at the upper left corner position in the coded block including block A, the upper right corner position in the coded block, and the lower left corner position in the coded block are derived. Then, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.

[0560] Notice, Figure 49A and Figure 49B The MV derivation method shown can be used in the later described Figure 50 In the MV derivation at each control point of the current block in step Sk_1 shown in , or can be used for the later described Figure 51 The MV predictor at each control point of the current block in step Sj_1 is shown to be derived.

[0561] (MV derivation > affine mode > affine merge mode)

[0562] Figure 50 is a flowchart showing one example of a process in affine merge mode.

[0563] In the affine merge mode as shown in the figure, first, the inter-frame predictor 126 derives the MV at the corresponding control point of the current block (step Sk_1). Figure 46A As shown, the control points are the upper left corner point of the current block and the upper right corner point of the current block, or as Figure 46B Shown are the top left corner point of the current block, the top right corner point of the current block, and the bottom left corner point of the current block.The inter-frame predictor 126 may encode MV selection information for identifying two or three derived MVs in the stream.

[0564] For example, when using Figures 47A to 47C When the MV derivation method shown is used, Figure 47A As shown, the inter-frame predictor 126 examines the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in the order listed and identifies the first valid block encoded according to the affine mode.

[0565] The inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine mode to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as Figure 47B As shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left corner and the upper right corner of the coding block including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block by projecting the motion vectors v3 and v4 at the upper left corner and the upper right corner of the coding block onto the current block.

[0566] Alternatively, when block A is identified and block A has three control points, as Figure 47C As shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block, the motion vector v1 at the upper right control point of the current block, and the motion vector v2 at the lower left control point of the current block based on the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the coding block including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block, the motion vector v1 at the upper right control point of the current block, and the motion vector v2 at the lower left control point of the current block by projecting the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the coding block onto the current block.

[0567] Note that, as mentioned above, Figure 49A As shown in , when block A is identified and block A has two control points, the MV at three control points can be calculated, and as described above in Figure 49B As shown in , when block A is identified and block A has three control points, MVs at two control points can be calculated.

[0568] Next, the inter-frame predictor 126 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the MV of each of the multiple sub-blocks as an affine MV using, for example, two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-blocks (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all sub-blocks included in the current block, the process of generating a predicted image using the affine merge mode of the current block is completed. In other words, motion compensation of the current block is performed to generate a predicted image for the current block.

[0569] Note that the above MV candidate list can be generated in step Sk_1. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B Any combination of the MV derivation methods shown in and other MV derivation methods.

[0570] Note that, in addition to the affine mode, the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks.

[0571] Note that, for example, an MV candidate list including MV candidates in the affine merge mode using two control points and the affine merge mode using three control points may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points may be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points may be generated. The MV candidates may be, for example, MVs of block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left) used for encoding, or MVs of valid blocks in the block.

[0572] Note that an index indicating one of the MVs in the MV candidate list may be transmitted as the MV selection information.

[0573] (MV derivation > affine mode > affine inter-frame mode)

[0574] Figure 51 is a flowchart showing one example of a process in affine inter mode.

[0575] In the affine inter mode, first, the inter predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the upper left corner point of the current block, the upper right corner point of the current block, and the upper right corner point of the current block, such as Figure 46A or Figure 46B shown.

[0576] For example, when using Figure 48A and Figure 48B When the MV derivation method shown in FIG. 1 is used, the inter-frame predictor 126 selects Figure 48A or Figure 48B MV predictors (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block are derived by using the MVs of any blocks in the coding blocks near the corresponding control points of the current block shown in FIG. At this time, the inter-frame predictor 126 encodes MV predictor selection information for identifying the selected two or three MV predictors in the stream.

[0577] For example, the inter-frame predictor 126 may determine a block from which to select the MV as the MV predictor at the control point from among the coding blocks adjacent to the current block using cost estimation or the like, and may write a flag indicating which MV predictor has been selected into the bitstream. In other words, the inter-frame predictor 126 outputs MV predictor selection information such as a flag as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.

[0578] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 calculates the MV of each subblock corresponding to the updated MV predictor as an affine MV using the above expression (1A) or expression (1B) (step Sj_3). The inter-frame predictor 126 then performs motion compensation for the subblocks using these affine MVs and the encoded reference picture (step Sj_4). When the MV predictor is updated in step Sj_2, the process in steps Sj_3 and Sj_4 is performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor as an MV difference in the stream. In other words, the inter predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130 .

[0579] Finally, the inter predictor 126 generates a predicted image of the current block by performing motion compensation of the current block using the determined MV and the encoded reference picture (step Sj_6).

[0580] Note that the above MV candidate list can be generated in step Sj_1. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B Any combination of the MV derivation methods shown in and other MV derivation methods.

[0581] Note that, in addition to the affine mode, the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks.

[0582] Note that, for example, an MV candidate list including MV candidates in an affine inter mode using two control points and an affine inter mode using three control points may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in an affine inter mode using two control points and an MV candidate list including MV candidates in an affine inter mode using three control points may be generated separately. Alternatively, an MV candidate list including MV candidates in one of an affine inter mode using two control points and an affine inter mode using three control points may be generated. The MV candidate may be, for example, the MV of block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left) used for encoding, or the MV of a valid block in the block.

[0583] Note that an index indicating one of the MV candidates in the MV candidate list may be transmitted as MV predictor selection information.

[0584] (MV derivation > triangle pattern)

[0585] In the above example, the inter-frame predictor 126 generates a single rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 may generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and may combine the multiple prediction images to generate a final rectangular prediction image. A shape other than a rectangle may be, for example, a triangle.

[0586] Figure 52A It is a conceptual diagram for illustrating the generation of two triangular prediction images.

[0587] The inter-frame predictor 126 performs motion compensation on the first partition having a triangular shape in the current block using the first MV of the first partition to generate a triangular predicted image. Similarly, the inter-frame predictor 126 performs motion compensation on the second partition having a triangular shape in the current block using the second MV of the second partition to generate a triangular predicted image. Then, the inter-frame predictor 126 combines these predicted images to generate a predicted image having the same rectangular shape as the current block.

[0588] Note that a first predicted image having a rectangular shape corresponding to the current block can be generated using the first MV as the predicted image for the first partition. Furthermore, a second predicted image having a rectangular shape corresponding to the current block can be generated using the second MV as the predicted image for the second partition. The predicted image for the current block can be generated by performing a weighted addition of the first predicted image and the second predicted image. Note that the portion where the weighted addition is performed may be a partial area across the boundary between the first partition and the second partition.

[0589] Figure 52Bis a conceptual diagram illustrating an example of a first portion of a first partition that overlaps with a second partition and first and second sample sets that may be weighted as part of a correction process. The first portion may be, for example, one-quarter the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to an edge of the first partition, where N is an integer greater than zero, for example, N may be an integer of 2. As shown in the figure, Figure 52B The left example of shows a rectangular partition with a rectangular portion whose width is one quarter of the width of the first partition, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion. Figure 52B The central example of shows a rectangular partition with a rectangular portion whose height is one quarter of the height of the first partition, where the first sample set includes samples outside the first portion and samples inside the first portion, and the second sample set includes samples inside the first portion. Figure 52B The right example of shows a triangular partition with a polygonal portion whose height corresponds to two samples, where a first set of samples includes samples outside the first portion and samples inside the first portion, and a second set of samples includes samples inside the first portion.

[0590] The first portion may be a portion of the first partition that overlaps with an adjacent partition. Figure 52C This is a conceptual diagram illustrating a first portion of a first partition, which overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition is shown with an overlapping portion with a spatially adjacent rectangular partition. Partitions of other shapes, such as triangular partitions, can be used, and the overlapping portion can overlap with spatially or temporally adjacent partitions.

[0591] Furthermore, although an example has been given of generating a predicted image for each of two partitions using inter prediction, a predicted image may be generated for at least one partition using intra prediction.

[0592] Figure 53 is a flowchart showing one example of a process in triangle mode.

[0593] In triangular mode, first, the inter-frame predictor 126 partitions the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode partition information (which is information related to the partitioning) as prediction parameters in the stream. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.

[0594] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as MVs of multiple coding blocks temporally or spatially surrounding the current block (step Sx_2). In other words, the inter-frame predictor 126 generates an MV candidate list.

[0595] The inter-frame predictor 126 then selects the MV candidate for the first partition and the MV candidate for the second partition from the multiple MV candidates obtained in step Sx_1 as the first MV and the second MV, respectively (step Sx_3). At this time, the inter-frame predictor 126 encodes MV selection information used to identify the selected MV candidate in the stream as a prediction parameter. In other words, the inter-frame predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.

[0596] Next, the inter-frame predictor 126 generates a first predicted image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Similarly, the inter-frame predictor 126 generates a second predicted image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).

[0597] Finally, the inter predictor 126 generates a predicted image of the current block by performing weighted addition of the first predicted image and the second predicted image (step Sx_6).

[0598] Note that although Figure 52A In the example shown, the first partition and the second partition are triangles, but the first partition and the second partition may be trapezoidal or other shapes different from each other. Figure 52A and Figure 52C The example shown includes two partitions, but the current block may include three or more partitions.

[0599] In addition, the first partition and the second partition may overlap each other. In other words, the first partition and the second partition may include the same pixel area. In this case, the predicted image in the first partition and the predicted image in the second partition can be used to generate the predicted image of the current block.

[0600] Furthermore, although an example has been shown in which a predicted image is generated for each of two partitions using inter prediction, a predicted image may be generated for at least one partition using intra prediction.

[0601] Note that the MV candidate list used to select the first MV and the MV candidate list used to select the second MV may be different from each other, or the MV candidate list used to select the first MV may be used as the MV candidate list used to select the second MV.

[0602] Note that the partition information may include an index indicating the direction in which at least the current block is partitioned into multiple partitions. The MV selection information may include an index indicating the selected first MV and an index indicating the selected second MV. One index may indicate multiple pieces of information. For example, one index may be encoded that collectively indicates part or all of the partition information and part or all of the MV selection information.

[0603] (MV derivation > ATMVP mode)

[0604] Figure 54 is a conceptual diagram for illustrating an example of an advanced temporal motion vector prediction (ATMVP) mode for deriving an MV in units of subblocks.

[0605] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, an MV candidate of each subblock is registered in an MV candidate list for use in a normal merge mode.

[0606] More specifically, in the ATMVP mode, first, as Figure 54 As shown, a temporal MV reference block associated with the current block is identified in the coding reference picture specified by the MV (MV0) of the adjacent block located at the lower left position relative to the current block. Next, in each subblock in the current block, an MV is identified for encoding the area corresponding to the subblock in the temporal MV reference block. The MVs identified in this manner are included in the MV candidate list as MV candidates for the subblocks in the current block. When an MV candidate for each subblock is selected from the MV candidate list, the subblock undergoes motion compensation, where the MV candidate is used as the MV of the subblock. In this way, a predicted image is generated for each subblock.

[0607] Although Figure 54 In the example shown, a block located at the lower left position relative to the current block is used as a surrounding MV reference block, but it should be noted that another block can be used. In addition, the size of the sub-block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub-block can be switched for units such as slices, tiles, pictures, etc.

[0608] (MV derivation > DMVR)

[0609] Figure 55 is a flow chart illustrating the relationship between merge mode and decoder motion vector refinement (DMVR).

[0610] The inter-frame predictor 126 derives the motion vector of the current block according to the merge mode (step S1_1). Next, the inter-frame predictor 126 determines whether to perform motion vector estimation, that is, motion estimation (step S1_2). Here, when it is determined not to perform motion estimation (No in step S1_2), the inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merge mode.

[0611] When it is determined in step S1_1 that motion estimation is to be performed (Yes in step S1_2), the inter-frame predictor 126 derives the final motion vector of the current block by estimating the surrounding area of ​​the reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to DMVR.

[0612] Figure 56 is a conceptual diagram for illustrating one example of a DMVR process for determining an MV.

[0613] First, for example, in merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference picture (L0), which is a coded picture in the L0 list, based on the MV candidate (L0). Similarly, reference pixels are identified from the second reference picture (L1), which is a coded picture in the L1 list, based on the MV candidate (L1). A template is generated by calculating the average of these reference pixels.

[0614] Next, the template is used to estimate each of the surrounding areas of the MV candidates for the first reference picture (L0) and the second reference picture (L1), and the MV that produces the minimum cost is determined as the final MV. Note that the cost can be calculated using, for example, the difference between each pixel value in the template and the corresponding pixel value in the estimated area, the value of the MV candidate, etc.

[0615] It is not always necessary to perform exactly the same procedure described here. Other procedures for achieving the derivation of the final MV by estimation in the surrounding area of ​​the MV candidate may be used.

[0616] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining MV. Figure 56 An example of DMVR is shown in Figure 57 In the example shown, the costs are calculated without generating a template.

[0617] First, the inter-frame predictor 126 estimates the surrounding area of ​​the reference block in each reference picture included in the L0 list and the L1 list based on the initial MV as the MV candidate obtained from each MV candidate list. Figure 57 As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets a search position for the reference picture in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set (specifically, the initial MV (i.e., InitMV_L0)), the vector difference from the search position is MVd_L0. The inter-frame predictor 126 then determines an estimated position in the reference picture in the L1 list. The search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference as MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines a position symmetrical with respect to the position indicated by the initial MV as the search position in each reference picture in the L0 list and the L1 list. The inter predictor 126 calculates the sum of absolute differences (SAD) between pixel values ​​at the search positions in the block as a cost for each search position, and finds a search position that generates the minimum cost.

[0618] Figure 58A is a conceptual diagram for illustrating an example of motion estimation in DMVR, and Figure 58B is a flowchart illustrating one example of a process of motion estimation.

[0619] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and the eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, if the cost at a search position other than the starting point is determined to be the smallest, the inter-frame predictor 126 changes the target to the search position that achieves the minimum cost and executes the process in step 2. If the cost at the starting point is the smallest, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.

[0620] In step 2, the inter-frame predictor 126 performs a search similar to the process in step 1, treating the search position after the target change as a new starting point based on the result of the process in step 1. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, if the cost at the search position other than the starting point is determined to be minimized, the inter-frame predictor 126 performs the process in step 4. If the cost at the starting point is minimized, the inter-frame predictor 126 performs the process in step 3.

[0621] In step 4, the inter predictor 126 regards the search position at the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a vector difference.

[0622] In step 3, the inter-frame predictor 126 determines the pixel position with sub-pixel accuracy that achieves the minimum cost based on the costs at four points located above, below, left, and right relative to the starting point in step 1 or step 2, and regards this pixel position as the final search position. The pixel position with sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) above, below, left, and right, using the costs at the corresponding search positions of the four search positions as weights. The inter-frame predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as a vector difference.

[0623] (Motion Compensation > BIO / OBMC / LIC)

[0624] Motion compensation involves a mode for generating a predicted image and correcting the predicted image. Modes such as bidirectional optical flow (BIO), overlapped block motion compensation (OBMC), and local illumination compensation (LIC) are described later.

[0625] Figure 59 is a flowchart illustrating an example of a process of generating a predicted image.

[0626] The inter predictor 126 generates a predicted image (step Sm_1 ), and corrects the predicted image according to, for example, any of the above-described modes (step Sm_2 ).

[0627] Figure 60 is a flowchart illustrating another example of a generation process of a predicted image.

[0628] The inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame predictor 126 generates a predicted image using the motion vector (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, if it is determined that the correction process is to be performed (Yes in step Sn_3), the inter-frame predictor 126 generates a final predicted image by correcting the predicted image (step Sn_4). Note that in the LIC described later, both luminance and chrominance can be corrected in step Sn_4. If it is determined that the correction process is not to be performed (No in step Sn_3), the inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).

[0629] (Motion Compensation > OBMC)

[0630] Note that in addition to the motion information of the current block obtained through motion estimation, the motion information of neighboring blocks can also be used to generate an inter-frame prediction image. More specifically, by performing a weighted addition of a prediction image based on the motion information obtained through motion estimation (in the reference picture) and a prediction image based on the motion information of neighboring blocks (in the current picture), an inter-frame prediction image can be generated for each subblock in the current block. This inter-frame prediction (motion compensation) is also called overlapped block motion compensation (OBMC) or OBMC mode.

[0631] In OBMC mode, information indicating the sub-block size of OBMC (e.g., referred to as OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether OBMC mode is applied (e.g., referred to as OBMC flag) may be signaled at the CU level. Note that signaling of such information does not necessarily need to be performed at the sequence level and the CU level, and may be performed at another level (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0632] The OBMC mode will be described in more detail. Figure 61 and Figure 62 1 is a flowchart and a conceptual diagram for illustrating an outline of a predicted image correction process performed by OBMC.

[0633] First, if Figure 62 As shown, the predicted image (Pred) is obtained by conventional motion compensation using the MV assigned to the current block. Figure 62 In , the arrow “MV” points to the reference picture and indicates what the current block of the current picture refers to in order to obtain the predicted image.

[0634] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) already derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which indicates the reference picture from the current block. The first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.

[0635] Similarly, a predicted image (Pred_U) is obtained by applying the MV (MV_U) derived for the coding block adjacent to the current block above to the current block (reusing the MV of the current block). The MV (MV_U) is indicated by the arrow "MV_U", which indicates the reference picture from the current block. A second correction of the predicted image is performed by overlapping the predicted image Pred_U with the predicted image (e.g., Pred and Pred_L) on which the first correction has been performed. This provides an effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image of the current block.

[0636] Although the above example is a two-pass correction method using left and upper neighboring blocks, it is noted that the correction method may also be a three-pass or more correction method also using right and / or lower neighboring blocks.

[0637] Note that the area where such overlap is performed may be only a portion of the area near the block boundary, rather than the entire pixel area of ​​the block.

[0638] Note that the above description describes a predicted image correction process based on OBMC, which uses overlapping additional predicted images Pred_L and Pred_U to obtain a single predicted image Pred from a single reference picture. However, when correcting a predicted image based on multiple reference pictures, a similar process can be applied to each of the multiple reference pictures. In this case, after performing OBMC image correction based on multiple reference pictures and obtaining corrected predicted images from the corresponding reference pictures, the corrected predicted images are further overlapped to obtain the final predicted image.

[0639] Note that in OBMC, the current block unit may be a PU, or a sub-block unit obtained by further splitting the PU.

[0640] One example of a method for determining whether to apply OBMC is to use a signal, obmc_flag, indicating whether OBMC is applied. As a specific example, encoder 100 may determine whether the current block belongs to an area with complex motion. If the block belongs to an area with complex motion, encoder 100 sets obmc_flag to a value of "1" and applies OBMC during encoding. If the block does not belong to an area with complex motion, encoder 100 sets obmc_flag to a value of "0" and encodes the block without applying OBMC. Decoder 200 switches between applying and not applying OBMC by decoding obmc_flag written into the stream.

[0641] (Motion Compensation > BIO)

[0642] Next, we'll describe the MV derivation method. First, we'll describe a method for derivation based on a model that assumes uniform linear motion. This method is also called the bidirectional optical flow (BIO) method. Alternatively, this bidirectional optical flow can be written as BDOF instead of BIO.

[0643] Figure 63 This is a conceptual diagram showing a model assuming uniform linear motion. Figure 63 In (v x , v y ) represents a velocity vector, and τ0 and τ1 represent the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) represents the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) represents the MV corresponding to the reference picture Ref1.

[0644] Here, in the velocity vector (v x ,v y ) shows uniform linear motion, (MV x0 ,MV y0 ) and (MV x1 ,MV y1 ) are respectively expressed as (v xτ0 ,v yτ0 ) and (-v xτ1 ,-v yτ1 ), and gives the following optical flow equation (2).

[0645] [Mathematical expression 3]

[0646]

[0647] Here, I(k) represents the motion-compensated luminance value of reference picture k (k = 0, 1) after motion compensation. This optical flow equation indicates that the sum of the following is zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected on a pixel-by-pixel basis.

[0648] Note that a method other than deriving a motion vector based on a model assuming uniform linear motion may be used to derive a motion vector at the decoder side 200. For example, a motion vector may be derived in units of sub-blocks based on motion vectors of multiple neighboring blocks.

[0649] Figure 64is a flowchart illustrating one example of an inter-frame prediction process according to BIO. Figure 65 is a functional block diagram showing one example of a functional configuration of the inter-frame predictor 126 that can perform inter-frame prediction according to the BIO.

[0650] like Figure 65 As shown, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolated image deriver 126b, a gradient image deriver 126c, an optical flow deriver 126d, a correction value deriver 126e, and a predicted image corrector 126f. Note that the memory 126a may be the frame memory 122.

[0651] The inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter-frame predictor 126 derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ), and the motion vector M1 is the motion vector (MV x1 , MV y1 ).

[0652] Next, the interpolated image deriver 126b derivates the interpolated image I of the current block using the motion vector M0 and the reference picture L0 through the reference memory 126a. 0 Next, the interpolated image deriver 126b derivates the interpolated image I of the current block using the motion vector M1 and the reference picture L1 through the reference memory 126a. 1 (Step Sy_2). Here, the interpolated image I 0 is an image contained in the reference picture Ref0 and to be derived for the current block, and the interpolated image I 1 Is the image contained in the reference picture Ref1 and to be derived for the current block. 0 and interpolated image I 1 Each of the interpolated images I may be the same size as the current block. 0 and interpolated image I 1 Each of the interpolated images I may be an image larger than the current block. 0 and interpolated image I 1 A predicted image obtained by using a motion vector (M0, M1) and a reference picture (L0, L1) and applying a motion compensation filter may be included.

[0653] In addition, the gradient image deriver 126c obtains the interpolated image I 0and interpolated image I 1 Derive the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1 )(Step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), and the vertical gradient image is (Iy 0 , Iy 1 ). The gradient image deriver 126c may derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may indicate the amount of spatial variation of pixel values ​​along the horizontal direction, along the vertical direction, or both.

[0654] Next, the optical flow deriver 126d uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) derives an optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_4). The optical flow indicates a coefficient for correcting the amount of spatial pixel movement and may be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block may be a 4×4 pixel sub-CU. Note that optical flow derivation may be performed for each pixel unit, etc., rather than for each sub-block.

[0655] Next, the inter-frame predictor 126 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the correction value deriver 126e uses the optical flow (vx, vy) to derive correction values ​​for the pixel values ​​included in the current block (step Sy_5). The predicted image corrector 126f then uses the correction values ​​to correct the predicted image of the current block (step Sy_6). Note that the correction values ​​may be derived in units of pixels, or may be derived in units of multiple pixels or in units of sub-blocks.

[0656] Note that the BIO process flow is not limited to Figure 64 For example, you can just execute Figure 64 A part of the process disclosed in the specification may be used, or a different process may be added or used instead, or the processes may be performed in a different processing order, etc.

[0657] (Motion Compensation > LIC)

[0658] Next, one example of a mode for generating a predicted image (Prediction) using a Local Illumination Compensation (LIC) process is described.

[0659] Figure 66A is a conceptual diagram for illustrating one example of a process of a predicted image generation method using a brightness correction process performed by the LIC. Figure 66B is a flowchart showing one example of the process of a predicted image generation method using LIC.

[0660] First, the inter-frame predictor 126 derives MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).

[0661] Next, the inter-frame predictor 126 extracts information indicating how the luminance value of the current block has changed between the current block and the reference picture (step Sz_2). This extraction is performed based on the luminance pixel values ​​of the encoded left-neighboring reference region (surrounding reference region) and the encoded upper-neighboring reference region (surrounding reference region) in the current picture, as well as the luminance pixel values ​​at corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 126 uses this information indicating how the luminance value has changed to calculate a luminance correction parameter (step Sz_3).

[0662] The inter-frame predictor 126 generates a predicted image for the current block by performing a luma correction process in which a luma correction parameter is applied to a reference image in the reference picture specified by the MV (step Sz_4). In other words, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luma correction parameter. In this correction, luma, chroma, or both can be corrected. In other words, chroma correction parameters can be calculated using information indicating how chroma has changed, and a chroma correction process can be performed.

[0663] Notice, Figure 66A The shape of the surrounding reference area shown in is an example; another shape may be used.

[0664] In addition, although the process of generating a predicted image based on a single reference picture is described here, the case of generating a predicted image based on multiple reference pictures can be described in the same manner. The predicted image can be generated after performing a brightness correction process on the reference image obtained from the reference picture in the same manner as described above.

[0665] One example of a method for determining whether to apply LIC is to use a lic_flag, which serves as a signal indicating whether LIC is applied. As a specific example, encoder 100 determines whether the current block belongs to an area with luminance variations. If the block belongs to an area with luminance variations, encoder 100 sets lic_flag to a value of "1" and applies LIC during encoding. If the block does not belong to an area with luminance variations, encoder 100 sets lic_flag to a value of "0" and performs encoding without applying LIC. Decoder 200 can decode lic_flag written into the stream and decode the current block by switching between application and non-application of LIC according to the flag value.

[0666] One example of a different method for determining whether to apply the LIC process is based on whether the LIC process has already been applied to surrounding blocks. As a specific example, when the current block is already processed in merge mode, the inter-frame predictor 126 determines whether the surrounding blocks selected for encoding in MV derivation in merge mode have already been encoded using the LIC process. The inter-frame predictor 126 performs encoding by switching between applying and not applying the LIC process based on the result. Note that in this example, the same process is also applied to the decoder 200.

[0667] The brightness correction (LIC) process has been referenced Figure 66A and Figure 66B are described and are further described below.

[0668] First, the inter predictor 126 derives an MV for obtaining a reference image corresponding to a current block to be encoded from a reference picture that is a picture to be encoded.

[0669] Next, the inter-frame predictor 126 uses the luminance pixel values ​​of the encoded surrounding reference areas adjacent to the left and above the current block and the luminance values ​​in the corresponding positions of the reference picture specified by the MV to extract information indicating how the luminance value of the reference picture changes to the luminance value of the current picture and calculate the luminance correction parameters. For example, assume that the luminance pixel value of a given pixel in the surrounding reference area of ​​the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference area of ​​the reference picture is p1. The inter-frame predictor 126 calculates coefficients A and B for optimizing A×p1+B=p0 as the luminance correction parameters for multiple pixels in the surrounding reference area.

[0670] Next, the inter-frame predictor 126 performs a luminance correction process using the luminance correction parameters of the reference image in the reference picture specified by MV to generate a predicted image for the current block. For example, assume that the luminance pixel value in the reference image is p2, and the luminance pixel value of the predicted image after luminance correction is p3. The inter-frame predictor 126 generates a predicted image after the luminance correction process by calculating A×p2+B=p3 for each pixel in the reference image.

[0671] For example, a region having a determined number of pixels extracted from each of the upper adjacent pixels and the left adjacent pixels may be used as a surrounding reference region. In addition, the surrounding reference region is not limited to a region adjacent to the current block and may be a region not adjacent to the current block. Figure 66A In the example shown, the surrounding reference region in the reference picture may be a region specified by another MV in the current picture from among the surrounding reference regions in the current picture. For example, the other MV may be an MV in the surrounding reference region in the current picture.

[0672] Although operations performed by the encoder 100 have been described herein, it is noted that the decoder 200 performs similar operations.

[0673] Note that LIC can be applied not only to luminance but also to chrominance. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for any one of Y, Cb, and Cr.

[0674] In addition, the LIC process can be applied in units of subblocks. For example, correction parameters can be derived using the surrounding reference areas in the current subblock and the surrounding reference areas in the reference subblocks in the reference picture specified by the MV of the current subblock.

[0675] (Predictive Controller)

[0676] The prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from the intra-frame predictor 124) and the inter-frame prediction signal (the image or signal output from the inter-frame predictor 126), and outputs the selected prediction image to the subtractor 104 and the adder 116 as a prediction signal.

[0677] (Prediction Parameter Generator)

[0678] The prediction parameter generator 130 may output information related to intra prediction, inter prediction, selection of a prediction image in the prediction controller 128, and the like as prediction parameters to the entropy encoder 110. The entropy encoder 110 may generate a stream based on the prediction parameters input from the prediction parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameters may be used in the decoder 200. The decoder 200 may receive and decode the stream and perform the same prediction process as that performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameters may include, for example, (i) a selection prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) an optional index, flag, or value based on the prediction process performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128 or indicating the prediction process.

[0679] (Decoder)

[0680] Next, a decoder 200 capable of decoding a stream output from the above-described encoder 100 will be described. Figure 67 2 is a block diagram showing the functional structure of the decoder 200 according to the present embodiment. The decoder 200 is a device that decodes a stream, which is a coded image, in units of blocks.

[0681] like Figure 67 As shown, the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformer 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a prediction parameter generator 222, and a partition determiner 224. Note that the intra predictor 216 and the inter predictor 218 are configured as part of a prediction performer.

[0682] (Decoder installation example)

[0683] Figure 68 2 is a functional block diagram showing an example of an installation of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, Figure 67 The multiple components of the decoder 200 are installed in Figure 68 As shown, processor b1 and memory b2.

[0684] Processor b1 is a circuit that performs information processing and is coupled to memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 may be a processor such as a CPU. In addition, processor b1 may be a collection of multiple electronic circuits. In addition, for example, processor b1 may be responsible for Figure 67The roles of two or more constituent elements other than the constituent elements for storing information among the multiple constituent elements of the decoder 200 shown, etc.

[0685] Memory b2 is a dedicated or general-purpose memory for storing information used by processor b1 to decode the stream. Memory b2 may be an electronic circuit and may be connected to processor b1. Furthermore, memory b2 may be included in processor b1. Furthermore, memory b2 may be a collection of multiple electronic circuits. Furthermore, memory b2 may be a magnetic disk, optical disk, or the like, or may be represented as a storage device, recording medium, or the like. Furthermore, memory b2 may be a non-volatile memory or a volatile memory.

[0686] For example, the memory b2 may store images or streams. In addition, the memory b2 may store a program for causing the processor b1 to decode the streams.

[0687] In addition, for example, memory b2 can act as Figure 67 The memory b2 may serve as a memory for storing information. Figure 67 The roles of the block memory 210 and the frame memory 214 are shown. More specifically, the memory b2 can store reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.).

[0688] Note that in decoder 200, instead of Figure 67 All of the multiple constituent elements shown in the drawings may be implemented, and not all of the processes described herein may be performed. Figure 67 A portion of the illustrated constituent elements may be included in another device, or a portion of the processes described herein may be performed by another device.

[0689] Hereinafter, the overall flow of the process performed by the decoder 200 will be described, and then each constituent element included in the decoder 200 will be described. Note that some constituent elements included in the decoder 200 perform the same processes as those performed by some of the encoder 100, and therefore the same processes will not be described in detail again. For example, the inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, and loop filter 212 included in the decoder 200 respectively perform processes similar to those performed by the inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and loop filter 120 included in the decoder 200.

[0690] (Overall flow of the decoding process)

[0691] Figure 69 is a flowchart illustrating one example of the overall decoding process performed by the decoder 200 .

[0692] First, the partition determiner 224 in the decoder 200 determines a partition mode for each of a plurality of fixed-size blocks (e.g., 128×128 pixels) included in a picture based on the parameters input from the entropy decoder 202 (step Sp_1). This partition mode is the partition mode selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks in the partition mode.

[0693] The entropy decoder 202 decodes (specifically, entropy decodes) the encoded quantized coefficients and prediction parameters of the current block (step Sp_2 ).

[0694] Next, the inverse quantizer 204 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 206 performs inverse transform on the result to restore the prediction residual (ie, difference block) (step Sp_3).

[0695] Next, the prediction performer including all or part of the intra predictor 216, the inter predictor 218, and the prediction controller 220 generates a prediction signal for the current block (step Sp_4).

[0696] Next, the adder 208 adds the predicted image and the prediction residual to generate a reconstructed image of the current block (also referred to as a decoded image block) (step Sp_5).

[0697] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).

[0698] The decoder 200 then determines whether decoding of the entire picture has ended (step Sp_7). When it is determined that decoding has not ended (No in step Sp_7), the decoder 200 repeats the process starting from step Sp_1.

[0699] Note that the processes of these steps Sp_1 to Sp_7 may be sequentially performed by the decoder 200, or two or more processes may be performed in parallel. The processing order of the two or more processes may be modified.

[0700] (Segment Determiner)

[0701] Figure 70 2 is a conceptual diagram for illustrating the relationship between the segmentation determiner 224 and other constituent elements in the embodiment. As an example, the segmentation determiner 224 may perform the following process.

[0702] For example, the partition determiner 224 collects block information from the block memory 210 or the frame memory 214 and further obtains parameters from the entropy decoder 202. The partition determiner 224 may then determine a partition mode for the fixed-size block based on the block information and the parameters. The partition determiner 224 may then output information indicating the determined partition mode to the inverse transformer 206, the intra-frame predictor 216, and the inter-frame predictor 218. The inverse transformer 206 may perform an inverse transform of the transform coefficients based on the partition mode indicated by the information from the partition determiner 224. The intra-frame predictor 216 and the inter-frame predictor 218 may generate a predicted image based on the partition mode indicated by the information from the partition determiner 224.

[0703] (Entropy Decoder)

[0704] Figure 71 is a block diagram showing one example of the functional configuration of the entropy decoder 202.

[0705] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the segmentation mode by entropy decoding the stream. For example, CABAC is used for entropy decoding. More specifically, the entropy decoding 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a de-binarizer 202c. The binary arithmetic decoder 202a uses the context value derived by the context controller 202b to arithmetically decode the stream into a binary signal. The context controller 202b derives the context value based on the characteristics of the syntactic element or the surrounding state (i.e., the probability of occurrence of the binary signal) in the same manner as the context controller 110b of the encoder 100. The de-binarizer 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. The binarization can be performed according to the above-mentioned binarization method.

[0706] Thus, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can convert the stream (see Figure 1 ) are output to the intra predictor 216, the inter predictor 218, and the prediction controller 220. The intra predictor 216, the inter predictor 218, and the prediction controller 220 can perform the same prediction process as that performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.

[0707] Figure 72 is a conceptual diagram for illustrating the flow of an exemplary CABAC process in the entropy decoder 202.

[0708] Initialization is first performed in CABAC within the entropy decoder 202. During initialization, the binary arithmetic decoder 202a is initialized and an initial context value is set. The binary arithmetic decoder 202a and the debinarizer 202c then perform arithmetic decoding and debinarization of the coded data, for example, a CTU. At this point, the context controller 202b updates the context value each time arithmetic decoding is performed. The context controller 202b then saves the context value for post-processing. For example, the saved context value is used to initialize the context value for the next CTU.

[0709] (Inverse Quantizer)

[0710] The inverse quantizer 204 inversely quantizes the quantized coefficients of the current block input from the entropy decoder 202. More specifically, the inverse quantizer 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inversely quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.

[0711] Figure 73 204 is a block diagram showing one example of the functional configuration of the inverse quantizer 204 .

[0712] The inverse quantizer 204 includes, for example, a quantization parameter generator 204 a , a predicted quantization parameter generator 204 b , a quantization parameter storage 204 d , and an inverse quantization performer 204 e .

[0713] Figure 74 is a flowchart illustrating one example of an inverse quantization process performed by the inverse quantizer 204 .

[0714] The inverse quantizer 204 may be based on Figure 74 The flow shown is an example of performing an inverse quantization process for each CU. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined that inverse quantization is to be performed (yes in step Sv_11), the quantization parameter generator 204a obtains the difference quantization parameter of the current block from the entropy decoder 202 (step Sv_12).

[0715] Next, the predicted quantization parameter generator 204b obtains a quantization parameter of a processing unit different from the current block from the quantization parameter storage device 204d (step Sv_13). The predicted quantization parameter generator 204b generates a predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).

[0716] The quantization parameter generator 204a then generates a quantization parameter for the current block based on the difference quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the difference quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b may be added together to generate the quantization parameter of the current block. In addition, the quantization parameter generator 204a stores the quantization parameter of the current block in the quantization parameter storage device 204d (step Sv_16).

[0717] Next, the inverse quantization performer 204e inversely quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameters generated in step Sv_15 (step Sv_17).

[0718] Note that the difference quantization parameter can be decoded at the bit sequence level, picture level, slice level, brick level, or CTU level. In addition, the initial value of the quantization parameter can be decoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the initial value of the quantization parameter and the difference quantization parameter can be used to generate the quantization parameter.

[0719] Note that the inverse quantizer 204 may include a plurality of inverse quantizers and may inverse quantize the quantized coefficient using an inverse quantization method selected from a plurality of inverse quantization methods.

[0720] (Inverse Converter)

[0721] The inverse transformer 206 restores a prediction residual by inversely transforming the transform coefficients that are input from the inverse quantizer 204 .

[0722] For example, when information parsed from the stream indicates that EMT or AMT is to be applied (eg, when the AMT flag is true), the inverse transformer 206 inversely transforms transform coefficients of the current block based on information indicating the parsed transform type.

[0723] Furthermore, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.

[0724] Figure 75 is a flowchart illustrating one example of a process performed by the inverse converter 206 .

[0725] For example, the inverse transformer 206 determines whether there is information in the stream indicating that an orthogonal transform is not performed (step St_11). Here, when it is determined that there is no such information (No in step St_11) (for example: there is no indication as to whether an orthogonal transform is performed; there is an indication that an orthogonal transform is to be performed); the inverse transformer 206 obtains information indicating the transform type decoded by the entropy decoder 202 (step St_12). Next, based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step St_13). The inverse transformer 206 then performs an inverse orthogonal transform using the determined transform type (step St_14). Figure 75 As shown, when it is determined that there is information indicating that the orthogonal transform is not performed (Yes in step St_11) (for example, there is no explicit instruction to perform the orthogonal transform; there is no instruction to perform the orthogonal transform), the orthogonal transform is not performed.

[0726] Figure 76 is a flowchart illustrating one example of a process performed by the inverse converter 206 .

[0727] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a determined value (step Su_11). The determined value may be predetermined. Here, when it is determined that the transform size is less than or equal to the determined value (yes in step Su_11), the inverse transformer 206 obtains information indicating which transform type the encoder 100 has used from the at least one transform type included in the first transform type group from the entropy decoder 202 (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transformer 206.

[0728] Based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). If it is determined that the transform size is not less than or equal to the determined value (No in step Su_11), the inverse transformer 206 performs an inverse transform on the transform coefficients of the current block using the second transform type group (step Su_15).

[0729] Note that, as an example, the inverse orthogonal transform of the inverse transformer 206 may be based on Figure 75 or Figure 76 The process shown is performed for each TU. Furthermore, an inverse orthogonal transform can be performed by using a defined transform type without decoding information indicating the transform type used for the orthogonal transform. The defined transform type can be a predefined transform type or a default transform type. Specifically, the transform type can be DST7, DCT8, or the like. In the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.

[0730] (Adder)

[0731] The adder 208 reconstructs the current block by adding the prediction residual as input from the inverse transformer 206 and the prediction residual as input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.

[0732] (Block Storage)

[0733] The block memory 210 is a storage device for storing blocks included in the current picture and that can be referenced in intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.

[0734] (Loop Filter)

[0735] The loop filter 212 applies the loop filter to the reconstructed image generated by the adder 208 and outputs the filtered reconstructed image to the frame memory 214 and provides an output of the decoder 200, for example, to a display device or the like.

[0736] When the information indicating on or off of ALF parsed from the stream indicates that ALF is on, a filter may be selected from a plurality of filters based on, for example, the direction and activity of a local gradient, and the selected filter is applied to the reconstructed image.

[0737] Figure 77 2 is a block diagram showing one example of the functional configuration of the loop filter 212. Note that the configuration of the loop filter 212 is similar to that of the loop filter 120 of the encoder 100.

[0738] For example, Figure 77 As shown, the loop filter 212 includes a deblocking filter executor 212a, an SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filter process on the reconstructed image. The SAO executor 212b performs an SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include Figure 77 In addition, the loop filter 212 may be configured to Figure 77 The above process may be performed in a different processing order than the processing order disclosed in Figure 77 All the processes shown in , etc.

[0739] (Frame Memory)

[0740] The frame memory 214 is, for example, a storage device for storing reference pictures used in inter-frame prediction and may also be referred to as a frame buffer. More specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter 212.

[0741] (Predictor (Intra Predictor, Inter Predictor, Prediction Controller))

[0742] Figure 78 is a flowchart showing an example of a process performed by the predictor of the decoder 200. Note that the prediction performer may include all or part of the following constituent elements: the intra predictor 216; the inter predictor 218; and the prediction controller 220. The prediction performer includes, for example, the intra predictor 216 and the inter predictor 218.

[0743] The predictor generates a predicted image for the current block (step Sq_1). This predicted image is also referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra-frame prediction signal or an inter-frame prediction signal. More specifically, the predictor generates a predicted image for the current block using a reconstructed image obtained for another block by generating a predicted image, restoring a prediction residual, and adding the predicted image. The predictor of the decoder 200 generates the same predicted image as the predicted image generated by the predictor of the encoder 100. In other words, the predicted image is generated according to a method that is common between the predictors or a method that corresponds to each other.

[0744] The reconstructed image may be, for example, an image in a reference picture, or an image of a decoded block (ie, the aforementioned other block) in a current picture that includes the current block. The decoded block in the current picture may be, for example, a neighboring block of the current block.

[0745] Figure 79 is a flowchart illustrating another example of a process performed by the predictor of the decoder 200 .

[0746] The predictor determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode may be determined based on prediction parameters or the like.

[0747] When the first method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the third method (step Sr_2c).

[0748] The first method, the second method, and the third method may be different methods for generating a predicted image. Each of the first method to the third method may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image may be used in these prediction methods.

[0749] Figures 80A to 80C (collectively, FIG. 80 ) is a flowchart illustrating another example of a process performed by the predictor of the decoder 200 .

[0750] The predictor can perform a prediction process according to the flow shown in FIG80 as an example. Note that the intra block copy shown in FIG80 is a mode belonging to inter-frame prediction, and the block included in the current picture is called a reference image or reference block. In other words, intra block copy does not refer to a picture different from the current picture. In addition, the PCM mode shown in FIG80 is a mode belonging to intra-frame prediction, and in which transformation and quantization are not performed.

[0751] (Intra-frame predictor)

[0752] The intra-frame predictor 216 performs intra-frame prediction based on the intra-frame prediction mode parsed from the stream by referring to the blocks in the current picture stored in the block memory 210 to generate a predicted image of the current block (i.e., the intra-frame prediction block). More specifically, the intra-frame predictor 216 performs intra-frame prediction by referring to the pixel values ​​(e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-frame prediction image, and then outputs the intra-frame prediction image to the prediction controller 220.

[0753] Note that when an intra prediction mode in which a luma block is referenced in intra prediction of a chroma block is selected, the intra predictor 216 may predict the chroma components of the current block based on the luma components of the current block.

[0754] Furthermore, when information parsed from the stream indicates that PDPC is to be applied, the intra predictor 216 corrects the intra-predicted pixel value based on horizontal / vertical reference pixel gradients.

[0755] Figure 81 is a diagram showing one example of a process performed by the intra predictor 216 of the decoder 200 .

[0756] The intra predictor 216 first determines whether to use MPM. Figure 81As shown, the intra-frame predictor 216 determines whether the MPM flag indicating 1 is present in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 is present (yes in step Sw_11), the intra-frame predictor 216 obtains information indicating the intra-frame prediction mode selected in the encoder 100 from the entropy decoder 202. Note that this information is decoded by the entropy decoder 202 and output to the intra-frame predictor 216. Next, the intra-frame predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra-frame prediction modes. The intra-frame predictor 216 then determines an intra-frame prediction mode (step Sw_14), which is included in the multiple intra-frame prediction modes included in the MPM and indicated by the information obtained in step Sw_12.

[0757] When it is determined that there is no MPM flag indicating 1 (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 from at least one intra prediction mode not included in the MPM from the entropy decoder 202. Note that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. Then, the intra predictor 216 determines the intra prediction mode that is not included in the plurality of intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15 (step Sw_17).

[0758] The intra predictor 216 generates a predicted image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18 ).

[0759] (Inter-frame predictor)

[0760] The inter-frame predictor 218 predicts the current block by referencing the reference picture stored in the frame memory 214. Prediction is performed in units of the current block or the current subblock in the current block. Note that a subblock is included in a block and is a unit smaller than a block. The size of the subblock can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the subblock can be switched in units such as slices, tiles, and pictures.

[0761] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from the stream (e.g., prediction parameters output from the entropy decoder 202), and outputs the inter-frame prediction image to the prediction controller 220.

[0762] When the information parsed from the stream indicates that the OBMC mode is to be applied, the inter predictor 218 generates an inter predicted image using motion information of neighboring blocks in addition to motion information of the current block obtained through motion estimation.

[0763] Furthermore, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives motion information by performing motion estimation according to a pattern matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter-frame predictor 218 then performs motion compensation (prediction) using the derived motion information.

[0764] In addition, when the BIO mode is to be applied, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. In addition, when the information parsed from the stream indicates that the affine mode is to be applied, the inter-frame predictor 218 derives the MV of each subblock based on the MVs of multiple neighboring blocks.

[0765] (MV derivation process)

[0766] Figure 82 is a flowchart illustrating one example of an MV derivation process in the decoder 200 .

[0767] For example, the inter-frame predictor 218 determines whether to decode motion information (e.g., MV). For example, the inter-frame predictor 218 may make this determination based on a prediction mode included in the stream, or may make this determination based on other information included in the stream. Here, when it is determined that motion information is to be decoded, the inter-frame predictor 218 derives the MV of the current block in a mode that decodes motion information. When it is determined that motion information is not to be decoded, the inter-frame predictor 218 derives the MV in a mode that does not decode motion information.

[0768] Here, MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode, which will be described later. Among the modes, modes for decoding motion information include the normal inter mode, normal merge mode, and affine mode (specifically, affine inter mode and affine merge mode). Note that motion information may include not only the MV but also the MV predictor selection information described later. Modes that do not decode motion information include the FRUC mode, etc. The inter predictor 218 selects a mode for deriving the MV of the current block from a plurality of modes and derives the MV of the current block using the selected mode.

[0769] Figure 83 is a flowchart illustrating one example of a process of MV derivation in the decoder 200 .

[0770] For example, the inter-frame predictor 218 may determine whether to decode the MV difference, i.e., the determination may be made based on a prediction mode included in the stream, or based on other information included in the stream. Here, when it is determined that the MV difference is to be decoded, the inter-frame predictor 218 may derive the MV of the current block in a mode that decodes the MV difference. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.

[0771] When it is determined that no MV difference is to be decoded, the inter-frame predictor 218 derives the MV in a mode in which the MV difference is not to be decoded. In this case, the encoded MV difference is not included in the stream.

[0772] Here, as described above, the MV derivation modes include the normal inter mode, normal merge mode, FRUC mode, and affine mode described later. Modes that encode MV differences include the normal inter mode and affine mode (specifically, affine inter mode). Modes that do not encode MV differences include the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode). The inter predictor 218 selects a mode for deriving the MV of the current block from the multiple modes and derives the MV of the current block using the selected mode.

[0773] (MV derivation > conventional inter-frame mode)

[0774] For example, when the information parsed from the stream indicates that the normal inter mode is to be applied, the inter predictor 218 derives an MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.

[0775] Figure 84 is a flowchart illustrating an example of a process of performing inter prediction by a conventional inter mode in the decoder 200 .

[0776] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks that temporally or spatially surround the current block (step Sg_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.

[0777] Next, the inter-frame predictor 218 extracts N (an integer of 2 or more) MV candidates from the plurality of MV candidates obtained in step Sg_11 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the determined ranking in the priority order (step Sg_12). Note that the ranking in the priority order may be determined in advance for the respective N MV predictor candidates, and the ranking may be predetermined.

[0778] Next, the inter predictor 218 decodes the MV predictor selection information from the input stream and selects one MV predictor candidate from the N MV predictor candidates as the MV predictor of the current block using the decoded MV predictor selection information (step Sg_13).

[0779] Next, the inter predictor 218 decodes the MV difference from the input stream, and derives the MV of the current block by adding the difference value that is the decoded MV difference to the selected MV predictor (step Sg_14).

[0780] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference picture (step Sg_15). The process in steps Sg_11 to Sg_15 is performed for each block. For example, when the process in steps Sg_11 to Sg_15 is performed for each of all blocks in the slice, inter-frame prediction for the slice using the conventional inter-frame mode is completed. For example, when the process in steps Sg_11 to Sg_15 is performed for each of all blocks in the picture, inter-frame prediction for the picture using the conventional inter-frame mode is completed. Note that not all blocks included in the slice undergo the process in steps Sg_11 to Sg_15, and inter-frame prediction for the slice using the conventional inter-frame mode may be completed when some blocks undergo the process. This also applies to the picture in steps Sg_11 to Sg_15. When the process is performed for some blocks in the picture, inter-frame prediction for the picture using the conventional inter-frame mode may be completed.

[0781] (MV derivation > conventional merge mode)

[0782] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter predictor 218 derives an MV and performs motion compensation (prediction) using the MV.

[0783] Figure 85 is a flowchart illustrating an example of a process of inter prediction by conventional merge mode in the decoder 200 .

[0784] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sh_11). In other words, the inter-frame predictor 218 generates an MV candidate list.

[0785] Next, the inter-frame predictor 218 selects an MV candidate from the multiple MV candidates obtained in step Sh_11 and derives the MV of the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains MV selection information included in the stream as a prediction parameter and selects the MV candidate identified by the MV selection information as the MV of the current block.

[0786] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). For example, the process in steps Sh_11 to Sh_13 is performed for each block. For example, when the process in steps Sh_11 to Sh_13 is performed for each of all blocks in the slice, inter-frame prediction for the slice using the normal merge mode is completed. In addition, when the process in steps Sh_11 to Sh_13 is performed for each of all blocks in the picture, inter-frame prediction for the picture using the normal merge mode is completed. Note that not all blocks included in the slice undergo the process in steps Sh_11 to Sh_13, and inter-frame prediction for the slice using the normal merge mode may be completed when some blocks undergo the process. This also applies to the picture in steps Sh_11 to Sh_13. When the process is performed on some blocks in the picture, inter-frame prediction for the picture using the normal merge mode may be completed.

[0787] (MV derivation > FRUC mode)

[0788] For example, when information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, motion information is derived on the decoder 200 side without being signaled from the encoder 100 side. For example, the decoder 200 can derive motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values ​​in the current block.

[0789] Figure 86 is a flowchart illustrating an example of a process of performing inter-frame prediction by the FRUC mode in the decoder 200 .

[0790] First, the inter-frame predictor 218 generates a list indicating MVs of decoded blocks spatially or temporally adjacent to the current block as MV candidates by referring to the MVs (this list is an MV candidate list and can also be used as an MV candidate list of a normal merge mode, for example (step Si_11). Next, the best MV candidate is selected from a plurality of MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates an evaluation value of each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV of the current block ( Step S1_14). More specifically, for example, the selected best candidate MV is directly derived as the MV of the current block. Alternatively, for example, the MV of the current block is derived using pattern matching in the surrounding area of ​​the position corresponding to the selected best MV candidate and included in the reference picture. In other words, an estimation using pattern matching and evaluation values ​​in the reference picture can be performed in the surrounding area of ​​the best MV candidate. If an MV with a better evaluation value exists, the best MV candidate can be updated to the MV with the better evaluation value, and the updated MV can be determined as the final MV of the current block. In embodiments, updating the MV with the better evaluation value may not be performed.

[0791] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference picture (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each of all blocks in the slice, inter-frame prediction for the slice using the FRUC mode is completed. For example, when the processes in steps Si_11 to Si_15 are performed for each of all blocks in the picture, inter-frame prediction for the picture using the FRUC mode is completed. Each sub-block can be processed similarly to the case of each block.

[0792] (MV derivation > FRUC mode)

[0793] For example, when the information parsed from the stream indicates that the affine merge mode is to be applied, the inter predictor 218 derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.

[0794] Figure 87 is a flowchart illustrating an example of a process of inter prediction by affine merge mode in the decoder 200 .

[0795] In the affine merge mode, first, the inter-frame predictor 218 derives the MV at the corresponding control point of the current block (step Sk_11). Figure 46AAs shown, the control points are the upper left corner point of the current block and the upper right corner point of the current block, or as Figure 46B As shown, these are the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block.

[0796] For example, when using Figures 47A to 47C When the MV derivation method shown is used, Figure 47A As shown, the inter-frame predictor 218 examines the decoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in this order and identifies the first valid block decoded according to the affine mode. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine mode to derive the MV at the control point. For example, when block A is identified and block A has two control points, as shown in FIG. Figure 47B As shown, the inter-frame predictor 218 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left and upper right corners of the decoded block including block A. In this way, the MV at each control point is derived.

[0797] Note that Figure 49A As shown, when block A is identified and block A has two control points, the MV at three control points can be calculated, and as Figure 49B As shown, when block A is identified and when block A has three control points, MVs at two control points can be calculated.

[0798] Furthermore, when MV selection information is included in the stream as a prediction parameter, the inter predictor 218 may use the MV selection information to derive the MV at each control point of the current block.

[0799] Next, the inter-frame predictor 218 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and the decoded reference picture to perform motion compensation for the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate a predicted image for the current block.

[0800] Note that the above MV candidate list may be generated in step Sk_11. The MV candidate list may be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown, Figure 49A and Figure 49B Any combination of the MV derivation methods shown in and other MV derivation methods.

[0801] Note that, in addition to the affine mode, the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks.

[0802] Note that, for example, an MV candidate list including MV candidates in the affine merge mode using two control points and the affine merge mode using three control points may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points may be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points may be generated.

[0803] (MV derivation > affine inter-frame mode)

[0804] For example, when the information parsed from the stream indicates that the affine inter mode is to be applied, the inter predictor 218 derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.

[0805] Figure 88 is a flowchart illustrating an example of a process of inter prediction by affine inter mode in the decoder 200 .

[0806] In the affine inter mode, first, the inter predictor 218 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_11). The control points are the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block, as shown in FIG. Figure 46A or Figure 46B shown.

[0807] The inter-frame predictor 218 obtains the MV predictor selection information included in the stream as a prediction parameter, and uses the MV identified by the MV predictor selection information to derive the MV predictor at each control point of the current block. Figure 48A and Figure 48B When the MV derivation method shown in FIG. 1 is used, the inter-frame predictor 218 selects Figure 48A or Figure 48BThe MV predictor in the coding block near the corresponding control point of the current block shown in is selected using the MV of the block identified by the information to derive the motion vector predictor (v0, v1) or (v0, v1, v2) at the control point of the current block.

[0808] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block and the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point of the current block is derived.

[0809] Next, the inter-frame predictor 218 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1 and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sj_13). The inter-frame predictor 218 then uses these affine MVs and the decoded reference picture to perform motion compensation for the sub-blocks (step Sj_14). When the processes in steps Sj_13 and Sj_14 are performed for each sub-block included in the current block, the inter-frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate a predicted image for the current block.

[0810] Note that the above MV candidate list can be generated in step Sj_11 as in step Sk_11.

[0811] (MV derivation > triangle pattern)

[0812] For example, when information parsed from the stream indicates that the triangle mode is to be applied, the inter predictor 218 derives an MV in the triangle mode and performs motion compensation (prediction) using the MV.

[0813] Figure 89 is a flowchart illustrating an example of a process of inter-frame prediction by triangular mode in the decoder 200 .

[0814] In triangular mode, first, the inter-frame predictor 218 partitions the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 may obtain partition information from the stream as a prediction parameter, which is information related to partitioning. The inter-frame predictor 218 may then partition the current block into the first partition and the second partition based on the partition information.

[0815] Next, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sx_12). In other words, the inter-frame predictor 218 generates an MV candidate list.

[0816] The inter-frame predictor 218 then selects the MV candidate for the first partition and the MV candidate for the second partition as the first MV and the second MV, respectively, from the multiple MV candidates obtained in step Sx_11 (step Sx_13). At this time, the inter-frame predictor 218 can obtain MV selection information from the stream for identifying each selected MV candidate as a prediction parameter. The inter-frame predictor 218 can then select the first MV and the second MV based on the MV selection information.

[0817] Next, the inter-frame predictor 218 performs motion compensation using the selected first MV and the decoded reference picture to generate a first predicted image (step Sx_14). Similarly, the inter-frame predictor 218 performs motion compensation using the selected second MV and the decoded reference picture to generate a second predicted image (step Sx_15).

[0818] Finally, the inter-frame predictor 218 generates a predicted image of the current block by performing weighted addition of the first predicted image and the second predicted image (step Sx_16).

[0819] (MV estimate > DMVR)

[0820] For example, the information parsed from the stream indicates that DMVR is to be applied, and the inter predictor 218 performs motion estimation using DMVR.

[0821] Figure 90 is a flowchart illustrating an example of a process of motion estimation performed by DMVR in the decoder 200 .

[0822] The inter-frame predictor 218 derives the MV of the current block according to the merge mode (step S1_11). Next, the inter-frame predictor 218 derives the final MV of the current block by searching the area around the reference picture indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined according to DMVR.

[0823] Figure 91 is a flowchart showing an example of a motion estimation process performed by DMVR in the decoder 200, and is similar to Figure 58B same.

[0824] First, in Figure 58AIn step 1 shown, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and the eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is minimum, the inter-frame predictor 218 changes the target to the search position that obtains the minimum cost and executes the process in step 2 shown in FIG. 58 . When the cost at the starting point is minimum, the inter-frame predictor 218 skips Figure 58A The process in step 2 shown in FIG and the process in step 3 are performed.

[0825] In such Figure 58A In step 2 shown, the inter-frame predictor 218 performs a search similar to the process in step 1. Based on the results of the process in step 1, the search position after the target change is regarded as a new starting point. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is minimized. Here, if the cost at one of the search positions other than the starting point is determined to be minimized, the inter-frame predictor 218 performs the process in step 4. If the cost at the starting point is minimized, the inter-frame predictor 218 performs the process in step 3.

[0826] In step 4, the inter predictor 218 regards the search position at the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a vector difference.

[0827] exist Figure 58A In step 3 shown, the inter-frame predictor 218 determines the pixel position with sub-pixel accuracy that obtains the minimum cost based on the costs at four points located at the upper, lower, left, and right positions relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position.

[0828] The pixel position with sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) for up, down, left, and right, using the cost at the corresponding search position among the four search positions as a weight. The inter-frame predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as a vector difference.

[0829] (Motion Compensation > BIO / OBMC / LIC)

[0830] For example, when the information parsed from the stream indicates that the predicted image is to be corrected, the inter-frame predictor 218 corrects the predicted image based on the correction mode when the predicted image is generated. The correction mode is, for example, one of the above-mentioned BIO, OBMC, and LIC.

[0831] Figure 92is a flowchart showing one example of a process of generating a predicted image in the decoder 200 .

[0832] The inter-frame predictor 218 generates a predicted image (step Sm_11 ), and corrects the predicted image according to any of the above-described modes (step Sm_12 ).

[0833] Figure 93 is a flowchart illustrating another example of a process of generating a predicted image in the decoder 200 .

[0834] The inter-frame predictor 218 derives the MV of the current block (step Sn_11). Next, the inter-frame predictor 218 generates a predicted image using the MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, the inter-frame predictor 218 obtains prediction parameters included in the stream and determines whether to perform a correction process based on the prediction parameters. For example, the prediction parameters are flags indicating whether one or more of the above-mentioned modes should be applied. Here, if it is determined that the correction process is to be performed (Yes in step Sn_13), the inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). Note that in LIC, both luminance and chrominance can be corrected in step Sn_14. If it is determined that the correction process is not to be performed (No in step Sn_13), the inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).

[0835] (Motion Compensation > OBMC)

[0836] For example, when the information parsed from the stream indicates that OBMC is to be performed, when a predicted image is generated, the inter predictor 218 corrects the predicted image according to OBMC.

[0837] Figure 94 : is a flowchart showing an example of a process of correction of a predicted image by OBMC in the decoder 200. Note that Figure 94 The flowchart in the figure shows the use of Figure 62 The correction process of the predicted image of the current picture and the reference picture is shown.

[0838] First, if Figure 62 As shown, a predicted image (Pred) is obtained by conventional motion compensation using the MV assigned to the current block.

[0839] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) already derived for the decoded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The inter-frame predictor 218 then performs a first correction of the predicted image by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.

[0840] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) derived for the decoded block adjacent to the current block above to the current block (reusing the motion vector for the current block). The inter-frame predictor 218 then performs a second correction on the predicted image by overlapping the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) on which the first correction has been performed. This provides an effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.

[0841] (Motion Compensation > BIO)

[0842] For example, when the information parsed from the stream indicates that BIO is to be performed, when a predicted image is generated, the inter predictor 218 corrects the predicted image according to the BIO.

[0843] Figure 95 is a flowchart illustrating an example of a process of correction of a predicted image performed by the BIO in the decoder 200 .

[0844] like Figure 63 As shown, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter-frame predictor 218 derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ), the motion vector M1 is the motion vector (MV x1 , MV y1 ).

[0845] Next, the inter-frame predictor 218 uses the motion vector M0 and the reference picture L0 to derive the interpolated image I of the current block. 0 In addition, the inter-frame predictor 218 uses the motion vector M1 and the reference picture L1 to derive the interpolated image I of the current block. 1 (Step Sy_12). Here, the interpolated image I 0 is an image contained in the reference picture Ref0 and to be derived for the current block, and the interpolated image I 1 It is an image included in the reference picture Ref1 and derived for the current block. 0 and interpolated image I 1 Each of the interpolated images I may be the same size as the current block.0 and interpolated image I 1 Each of the interpolated images I may be an image larger than the current block. 0 and interpolated image I 1 A predicted image obtained by using a motion vector (M0, M1) and a reference picture (L0, L1) and applying a motion compensation filter may be included.

[0846] In addition, the inter-frame predictor 218 obtains the interpolated image I 0 and interpolated image I 1 Derive the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1 )(Step Sy_13). Note that the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), and the vertical gradient image is (Iy 0 , Iy 1 The inter-frame predictor 218 may derive a gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may be an image each indicating a spatial variation of a pixel value in the horizontal direction or a spatial variation of a pixel value in the vertical direction.

[0847] Next, the inter-frame predictor 218 uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) derives an optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_14). As an example, the sub-block may be a 4×4 pixel sub-CU.

[0848] Next, the inter-frame predictor 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter-frame predictor 218 uses the optical flow (vx, vy) to derive correction values ​​for the pixel values ​​included in the current block (step Sy_15). The inter-frame predictor 218 then uses the correction values ​​to correct the predicted image of the current block (step Sy_16). Note that the correction values ​​can be derived in units of pixels, or can be derived in units of multiple pixels, or in units of sub-blocks, etc.

[0849] Note that the BIO process flow is not limited to Figure 95 The process disclosed in . You can only execute Figure 95 The present invention may be a part of the processes disclosed in the specification, or different processes may be added or used as a substitute, or the processes may be performed in a different processing order.

[0850] (Motion Compensation > LIC)

[0851] For example, when the information parsed from the stream indicates that LIC is to be performed, when a predicted image is generated, the inter predictor 218 corrects the predicted image according to LIC.

[0852] Figure 96 is a flowchart illustrating an example of a process of correction of a predicted image performed by the LIC in the decoder 200 .

[0853] First, the inter-frame predictor 218 obtains a reference image corresponding to the current block from the decoded reference picture using MV (step Sz_11).

[0854] Next, the inter-frame predictor 218 extracts information indicating how the luminance values ​​of the current picture and the reference picture have changed for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values ​​of the decoded left-neighboring reference region (surrounding reference region) and the decoded upper-neighboring reference region (surrounding reference region), as well as the luminance pixel values ​​at corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 218 uses this information indicating how the luminance values ​​have changed to calculate luminance correction parameters (step Sz_13).

[0855] The inter-frame predictor 218 generates a predicted image for the current block by performing a luma correction process in which a luma correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_14). In other words, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luma correction parameter. This correction can be performed for luma or chroma.

[0856] (Predictive Controller)

[0857] The prediction controller 220 selects an intra-frame predicted image or an inter-frame predicted image and outputs the selected image to the adder 208. In general, the configuration, function, and process of the prediction controller 220, the intra-frame predictor 216, and the inter-frame predictor 218 on the decoder 200 side may correspond to the configuration, function, and process of the prediction controller 128, the intra-frame predictor 124, and the inter-frame predictor 126 on the encoder 100 side.

[0858] (First aspect)

[0859] Figure 97 is a flow chart of an example of a process flow 1000 for decoding an image using a CCALF (cross-component adaptive loop filtering) process according to the first aspect. The process flow 1000 may be, for example, Figure 67 The decoder 200 and the like are executed.

[0860] In step S1001, a filtering process is applied to reconstructed image samples of a first component. For example, the first component may be a luma component. The luma component may be represented as a Y component. The reconstructed luma image samples may be the output signal of an ALF process. The output signal of the ALF process may be reconstructed luma samples generated by a SAO process. In some embodiments, the filtering process performed in step S1001 may be represented as a CCALF process. The number of reconstructed luma samples may be the same as the number of coefficients of the filter to be used in the CCALF process. In other embodiments, a clipping process may be performed on the filtered reconstructed luma samples.

[0861] In step S1002, the reconstructed image samples of the second component are modified. The second component may be a chroma component. The chroma component may be represented as a Cb and / or Cr component. The reconstructed image samples of chroma may be the output signal of the ALF process. The output signal of the ALF may be a reconstructed chroma sample generated by the SAO process. The modified reconstructed image samples may be the sum of the reconstructed samples of chroma and the filtered reconstructed samples of luminance, i.e., the output of step S1001. In other words, the modification process may be performed by adding the filtered values ​​of the reconstructed luminance samples generated by the CCALF process of step S1001 to the filtered values ​​of the reconstructed chroma samples generated by the ALF process. In some embodiments, a clipping process may be performed on the reconstructed chroma samples. The first component and the second component may belong to the same block, or may belong to different blocks.

[0862] In step S1003, the values ​​of the modified reconstructed image samples of the chrominance components are clipped. By performing the clipping process, the sample values ​​can be ensured to be within a certain range. Furthermore, clipping can promote better convergence in processes such as least squares optimization to minimize the difference between the residual (the difference between the original sample value and the reconstructed sample value) and the filtered values ​​of the chrominance samples to facilitate the determination of filter coefficients.

[0863] In step S1004, the image is decoded using the cropped reconstructed image samples of the chrominance components. In some embodiments, step S1003 need not be performed. In this case, the image is decoded using the uncropped modified reconstructed chrominance samples.

[0864] Figure 98 is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment. In this embodiment, a clipping process is applied to the modified reconstructed image samples of the chrominance components, such as Figure 97In step S1003 of FIG. 4 , for example, for a 10-bit output, the modified reconstructed image samples may be clipped to a range of [0, 1023]. In some embodiments, when clipping the filtered reconstructed image samples of the luma component generated by the CCALF process, it may not be necessary to clip the modified reconstructed image samples of the chroma components.

[0865] Figure 99 : is a block diagram showing the functional configuration of an encoder and a decoder according to an embodiment. Figure 97 In step S1003, a clipping process is applied to the modified reconstructed image samples of the chrominance components. The clipping process is not applicable to the filtered reconstructed luminance samples generated by the CCALF process. The filtered values ​​of the reconstructed chrominance samples generated by the ALF process do not need to be clipped. Figure 99 In other words, the reconstructed image samples to be modified are generated using the filtered values ​​(ALF chroma) and the difference values ​​(CCALF Cb / Cr), where no clipping is applied to the output of the generated sample values.

[0866] Figure 100 is a block diagram illustrating the functional configuration of an encoder and decoder according to an embodiment. In this embodiment, a clipping process is applied to the filtered reconstructed luma samples ("Clipped Output Samples") and the modified reconstructed image samples of the chroma components ("Summed Clipped") generated by the CCALF process. The filtered values ​​of the reconstructed chroma samples generated by the ALF process are not clipped ("No Clipping"). For example, the clipping range applied to the filtered reconstructed image samples of the luma component can be [-2^15, 2^15-1] or [-2^7, 2^7-1].

[0867] Figure 101 Another example is shown in which the clipping process is applied to: the filtered reconstructed luma samples generated by the CCALF process ("clipped output samples"), the modified reconstructed image samples of the chroma components ("clipping after summation"), and the filtered reconstructed chroma samples generated by the ALF process ("clipping"). In other words, the output values ​​of the CCALF process and the ALF chroma process are clipped individually and clipped again after they are summed. In this embodiment, the modified reconstructed image samples of the chroma components do not need to be clipped. For example, the final output of the ALF chroma process may be clipped to 10-bit values. For example, the clipping range of the filtered reconstructed image samples applied to the luma component may be [-2^15, 2^15-1] or [-2^7, 2^7-1]. The range may be fixed or may be determined adaptively. In either case, the range may be signaled in the header information, such as in the SPS (Sequence Parameter Set) or APS (Adaptation Parameter Set). In the case when a non-linear ALF is used, it may be Figure 101 Define the clipping parameters in "Clip after summing".

[0868] The reconstructed image samples of the luma component to be filtered by the CCALF process may be adjacent samples to the current reconstructed image samples of the chroma components. That is, the modified current reconstructed image samples may be generated by adding the filtered values ​​of the adjacent luma component image samples positioned adjacent to the current image sample to the filtered values ​​of the current chroma component image samples. The filtered values ​​of the luma component image samples may be represented as difference values.

[0869] The process disclosed in this aspect may reduce the hardware internal memory size required to store filtered image sample values.

[0870] (Second aspect)

[0871] Figure 102 is a flow chart of an example of a process flow 2000 for applying a CCALF process to decode an image using defined information according to the second aspect. The process flow 2000 may be, for example, Figure 67 The decoder 200 and the like are executed.

[0872] In step S2001, the cropping parameters are parsed from the bitstream. The cropping parameters can be parsed from the VPS, APS, SPS, PPS, slice header at the CTU or TU level, such as Figure 103 As described in. Figure 103 This is a conceptual diagram showing the positions of clipping parameters. Figure 103 The parameters described in may be replaced by different types of tailoring parameters, flags or indices. Two or more tailoring parameters may be parsed from two or more parameter sets in the bitstream.

[0873] In step S2002, the difference is cropped using the cropping parameters. Based on the reconstructed image samples of the first component (eg, Figures 98-101 The difference is generated by the difference value (CCALF Cb / Cr) in the first component. For example, the first component is the luma component, and the difference is the filtered reconstructed luma samples generated by the CCALF process. In this case, the clipping process is applied to the filtered reconstructed luma samples using the resolved clipping parameters.

[0874] The clipping parameter limits the value to a desired range. If the desired range is [-3, 3], for example, the value 5 is clipped to 3 using the operation clip(-3, 3, 5). In this example, the value -3 is the lower limit, and the value 3 is the upper limit.

[0875] The trim parameters may indicate the indices used to derive the lower and upper bounds, such as Figure 104(i) In this example, ccalf_luma_clip_idx[] is the index, -range_array[] is the lower bound, and range_array[] is the upper bound. In this example, range_array[] is a defined range array, which may be different from the range array used for ALF. The defined range array may be predetermined.

[0876] The trimming parameters can indicate lower and upper limits, such as Figure 104 (ii) In this example, -ccalf_luma_clip_low_range[] is the lower range, and -ccalf_luma_clip_up_range[] is the upper range.

[0877] The trim parameter may indicate a common range for both the lower and upper ranges, such as Figure 104 (iii) In this example, -ccalf_luma_clip_range is the lower limit, and ccalf_luma_clip_range is the upper limit.

[0878] The difference is generated by multiplying, dividing, adding or subtracting at least two reconstructed image samples of the first component. For example, the two reconstructed image samples may be from a current and an adjacent image sample or two adjacent image samples. The positions of the current and adjacent image samples may be predetermined.

[0879] In step S2003, reconstructed image samples of a second component different from the first component are modified using a clipping value. The clipping value may be a clipping value of the reconstructed image samples of the luma component. The second component may be a chroma component. The modification may include an operation for multiplying, dividing, adding, or subtracting the clipping value relative to the reconstructed image samples of the second component.

[0880] In step S2004, the image is decoded using the modified reconstructed image samples.

[0881] In the present disclosure, one or more pruning parameters for cross-component adaptive loop filtering are signaled in the bitstream. This signaling allows the syntax of cross-component adaptive loop filtering to be combined with the syntax of adaptive loop filters for syntactic simplification. Furthermore, this signaling allows for more flexible design of cross-component adaptive loop filtering, thereby improving coding efficiency.

[0882] The cropping parameters may be defined or predefined for both the encoder and the decoder without signaling. The cropping parameters may also be derived using luminance information without signaling. For example, if strong gradients or edges are detected in the luminance reconstructed image, cropping parameters corresponding to a large cropping range may be derived, and if weak gradients or edges are detected in the luminance reconstructed image, cropping parameters corresponding to a short cropping range may be derived.

[0883] (Third aspect)

[0884] Figure 105 is a flow chart of an example of a process flow 3000 for decoding an image using filter coefficients applying a CCALF process according to the third aspect. The process flow 3000 may be, for example, Figure 67 The filter coefficients are used in the filtering step of the CCALF process to generate filtered reconstructed image samples of the luminance component.

[0885] In step S3001, a determination is made as to whether the filter coefficients are located within a defined symmetric region of the filter. Optionally, an additional step may be performed to determine whether the shape of the filter coefficients is symmetric. Information indicating whether the samples of the filter coefficients are symmetric may be encoded into the bitstream. If the shape is symmetric, the position of the coefficients within the symmetric region may be determined or predetermined.

[0886] In step S3002 , if the filter coefficient is within the defined symmetric region (“Yes” in step S3001 ), the filter coefficient is copied to a symmetric position and a set of filter coefficients is generated.

[0887] In step S3003, the reconstructed image samples of the first component are filtered using the filter coefficients. The first component may be a luminance component.

[0888] In step S3004, the output of the filtering is used to modify the reconstructed image samples of a second component different from the first component. The second component may be a chrominance component.

[0889] In step S3005, the image is decoded using the modified reconstructed image samples.

[0890] If the filter coefficients are asymmetric (No in step S3001), all filter coefficients may be encoded from the bitstream and a set of filter coefficients may be generated without duplication.

[0891] This aspect may reduce the amount of information to be encoded into the bitstream. That is, only one of the symmetric filter coefficients may need to be encoded in the bitstream.

[0892] Figure 106 、 Figure 107 、 Figure 108 、 Figure 109 and Figure 110 is a conceptual diagram indicating examples of positions of filter coefficients to be used in the CCALF process. In these examples, some coefficients included in a set of coefficients are signaled assuming that there is symmetry.

[0893] Specifically, Figure 106 (a) Figure 106 (b) Figure 106 (c) and Figure 106 (d) shows examples where a portion of a set of CCALF coefficients (marked by diagonal lines and a grid pattern) lies within a defined symmetric region. In these examples, the symmetric region has a line-symmetric shape. Only some of the marked coefficients (marked by diagonal lines or a grid pattern) and the white coefficients can be encoded into the bitstream, while the other coefficients can be generated by using the coded coefficients. As another example, only the marked coefficients can be generated and used in the filtering process. The other white coefficients (not marked by any pattern) do not need to be used in the filtering process.

[0894] Figure 106 (e) Figur...

Claims

1. An encoder, comprising: Circuit; as well as a memory coupled to the circuit; Wherein, the circuit is in operation: generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values ​​by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the first coefficient value to the trimmed second coefficient value; clipping the third coefficient value; and The third reconstructed image samples of the chrominance component are encoded using the clipped third coefficient values.

2. A decoder comprising: Circuit; as well as a memory coupled to the circuit; Wherein, the circuit is in operation: generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values ​​by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the first coefficient value to the trimmed second coefficient value; clipping the third coefficient value; and The third reconstructed image samples of the chrominance component are decoded using the clipped third coefficient values.

3. A transmitter for a bit stream, comprising: Circuit; as well as a memory coupled to the circuit; Wherein, the circuit is in operation: sending the bit stream; The bitstream includes filter information, and the filter information causes the decoder to perform a filtering process, and the filtering process includes: generating first coefficient values ​​by applying a CCALF (cross-component adaptive loop filtering) process to first reconstructed image samples of the luminance component; If the first coefficient value is less than 64, setting the first coefficient value to zero; generating second coefficient values ​​by applying an ALF (Adaptive Loop Filter) process to the second reconstructed image samples of the chrominance component; clipping the second coefficient value; generating a third coefficient value by adding the first coefficient value to the trimmed second coefficient value; clipping the third coefficient value; and The third reconstructed image samples of the chrominance component are decoded using the clipped third coefficient values.