Systems and methods for video coding

By applying cross-component adaptive loop filtering and adaptive loop filtering processes in video encoding, the reconstruction image samples of brightness and chrominance components are filtered and cropped, and coefficient values are generated and encoded, which solves the problems of insufficient encoding efficiency and image quality in the prior art, and realizes the encoding effect of low resource utilization, small circuit scale and fast speed.

CN120302064APending Publication Date: 2025-07-11PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510714458.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-18
Filing Date
2020-09-18
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing video encoding technology processes the increasing amount of digital video data, the encoding efficiency and image quality improvement are still insufficient, and the loop filtering process has problems such as high resource utilization and large circuit scale.

Method used

The cross component adaptive loop filtering (CCALF) and adaptive loop filtering (ALF) processes are used to filter and crop the reconstruction image samples of the luminance and chrominance components, generate and encode coefficient values, and generate third coefficient values for encoding by adding them. Combining technologies such as block segmentation, intra- and inter-frame prediction, transformation and quantization, the encoding process is optimized.

Benefits of technology

Improve coding efficiency, enhance image quality, reduce the utilization rate of processing resources and circuit scale, and improve encoding/decoding speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302064A_ABST
    Figure CN120302064A_ABST
Patent Text Reader

Abstract

The encoder includes a circuit and a memory. The circuitry performs a CCALF process on a current block in operation. The circuitry sets a first flag indicating whether a CCALF process is enabled for a first block, the first block being adjacent to a left side of a current block. The circuitry sets a second flag indicating whether the CCALF process is enabled for a second block adjacent to an upper side of the current block. The circuitry sets a third flag indicating that the CCALF process is enabled for the current block. The circuitry determines a first index associated with a color component of a current block. The circuitry derives a second index indicating a context model using the first flag, the second flag, and the first index, and performs entropy encoding on a third flag indicating whether to enable a CCALF process for the current block using the context model indicated by the second index.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent with the application date of September 18, 2020, the application number of 202080057593.9, and the invention name of "Systems and Methods for Video Coding". Technical Field

[0002] The present disclosure relates to video coding, and in particular to video coding and decoding systems, components and methods in video coding and decoding, such as for performing a CCALF (Cross-Component Adaptive Loop Filtering) process. Background Art

[0003] With the advancement of video coding technologies, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec), there is still a continuous need to improve and optimize video coding technologies to handle the increasing amount of digital video data in various applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in the CCALF (Cross-Component Adaptive Loop Filtering) process. Summary of the Invention

[0004] According to one aspect, there is provided an encoder comprising circuitry and a memory coupled to the circuitry. The circuitry, in operation, generates a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component and clips the first coefficient value. The circuitry generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component and clips the second coefficient value. The circuitry generates a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value and encodes a third reconstructed image sample of the chrominance component using the third coefficient value.

[0005] According to another aspect, the first reconstructed image sample is located adjacent to the second reconstructed image sample.

[0006] According to another aspect, the circuitry, in operation, sets the first coefficient value to zero in response to the first coefficient value being less than 64.

[0007] According to another aspect, an encoder is provided, which includes: a block splitter that divides a first image into a plurality of blocks in operation; an intra predictor that predicts a block included in the first image using a reference block included in the first image in operation; an inter predictor that predicts a block included in the first image using a reference block included in a second image different from the first image in operation; a loop filter that filters a block included in the first image in operation; a transformer that transforms a prediction error between an original signal and a prediction signal generated by the intra predictor or the inter predictor to generate transform coefficients in operation; a quantizer that quantizes the transform coefficients to generate quantized coefficients in operation; and an entropy encoder that variably encodes the quantized coefficients to generate an encoded bitstream including the encoded quantized coefficients and control information. The loop filter performs the following operations:

[0008] Generate a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to a first reconstructed image sample of the luminance component;

[0009] Clip the first coefficient value;

[0010] Generate a second coefficient value by applying an ALF (Adaptive Loop Filter) process to a second reconstructed image sample of the chrominance component;

[0011] Clip the second coefficient value;

[0012] Generate a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value; and

[0013] Encode a third reconstructed image sample of the chrominance component using the third coefficient value.

[0014] According to another aspect, a decoder is provided, which includes a circuit and a memory coupled to the circuit. The circuit generates a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to a first reconstructed image sample of the luminance component and clips the first coefficient value in operation. The circuit generates a second coefficient value by applying an ALF (Adaptive Loop Filter) process to a second reconstructed image sample of the chrominance component and clips the second coefficient value. The circuit generates a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value and decodes a third reconstructed image sample of the chrominance component using the third coefficient value.

[0015] According to another aspect, there is provided a decoding apparatus including: a decoder that decodes an encoded bitstream in operation to output quantized coefficients; an inverse quantizer that inverse quantizes the quantized coefficients in operation to output transform coefficients; an inverse transformator that inverse transforms the transform coefficients in operation to output a prediction error; an intra predictor that uses a reference block included in the first image to predict a block included in the first image in operation; an inter predictor that uses a reference block included in a second image different from the first image to predict a block included in the first image in operation; a loop filter that filters a block included in the first image in operation; and an output terminal that outputs a picture including the first image in operation. The loop filter performs the following operations:

[0016] Generating a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to a first reconstructed image sample of a luminance component;

[0017] Clipping the first coefficient value;

[0018] Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to a second reconstructed image sample of a chrominance component;

[0019] Clipping the second coefficient value;

[0020] Generating a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value; and

[0021] Decoding a third reconstructed image sample of the chrominance component using the third coefficient value.

[0022] According to another aspect, there is provided an encoding method including:

[0023] Generating a first coefficient value by applying a CCALF (Cross Component Adaptive Loop Filter) process to a first reconstructed image sample of a luminance component;

[0024] Clipping the first coefficient value;

[0025] Generating a second coefficient value by applying an ALF (Adaptive Loop Filter) process to a second reconstructed image sample of a chrominance component;

[0026] Clipping the second coefficient value;

[0027] Generating a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value; and

[0028] Encoding a third reconstructed image sample of the chrominance component using the third coefficient value.

[0029] According to another aspect, a decoding method is provided, including:

[0030] generating a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component;

[0031] clipping the first coefficient value;

[0032] generating a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component;

[0033] clipping the second coefficient value;

[0034] generating a third coefficient value by adding the clipped first coefficient value and the clipped second coefficient value; and

[0035] decoding a third reconstructed image sample of the chrominance component using the third coefficient value.

[0036] According to another aspect, an encoder is provided, which includes a circuit and a memory coupled to the circuit. The circuit generates a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component during operation. The circuit generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component. The circuit generates a third coefficient value by adding the first coefficient value and the second coefficient value, and encodes a third reconstructed image sample of the chrominance component using the third coefficient value. The circuit determines a first parameter for which the Cb component and the Cr component of the chrominance component have the same value. The circuit determines an entropy coding model from multiple models using the first parameter. The circuit performs entropy coding on a second parameter of the CCALF process using the model.

[0037] According to another aspect, a decoder is provided, which includes a circuit and a memory coupled to the circuit. The circuit determines a first parameter for which the Cb component and the Cr component of the chrominance component have the same value during operation. The circuit determines an entropy coding model from multiple models using the first parameter. The circuit performs entropy coding on a second parameter of the CCALF process using the model. The circuit generates a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component. The circuit generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component. The circuit generates a third coefficient value by adding the first coefficient value and the second coefficient value, and decodes a third reconstructed image sample of the chrominance component using the third coefficient value.

[0038] According to another aspect, an encoder is provided that includes a circuit and a memory. The circuit, in operation, generates a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component. The circuit generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component. The circuit generates a third coefficient value by adding the first coefficient value and the second coefficient value, and encodes a third reconstructed image sample of the chrominance component using the third coefficient value. In the CCALF process, in response to the coordinates of the second reconstructed image sample being (x, y), the coordinates of the first reconstructed image sample are (2x, 2y - 1), (2x - 1, 2y), (2x, 2y), (2x + 1, 2y), (2x - 1, 2y + 1), (2x, 2y + 1), (2x + 1, 2y + 1), and (2x, 2y + 2).

[0039] According to another aspect, a decoder is provided that includes a circuit and a memory. The circuit, in operation, generates a first coefficient value by applying a CCALF (Cross-Component Adaptive Loop Filtering) process to a first reconstructed image sample of a luminance component. The circuit generates a second coefficient value by applying an ALF (Adaptive Loop Filtering) process to a second reconstructed image sample of a chrominance component. The circuit generates a third coefficient value by adding the first coefficient value and the second coefficient value, and encodes a third reconstructed image sample of the chrominance component using the third coefficient value. In the CCALF process, in response to the coordinates of the second reconstructed image sample being (x, y), the coordinates of the first reconstructed image sample are (2x, 2y - 1), (2x - 1, 2y), (2x, 2y), (2x + 1, 2y), (2x - 1, 2y + 1), (2x, 2y + 1), (2x + 1, 2y + 1), and (2x, 2y + 2).

[0040] According to another aspect, an encoder is provided that includes a circuit and a memory. The circuit, in operation, performs a CCALF process on a current block. The circuit sets a first flag that indicates whether the CCALF process is enabled for a first block that is adjacent to the left of the current block. The circuit sets a second flag that indicates whether the CCALF process is enabled for a second block that is adjacent to the top of the current block. The circuit sets a third flag that indicates that the CCALF process is enabled for the current block. The circuit determines a first index associated with a color component of the current block. The circuit derives a second index indicating a context model using the first flag, the second flag, and the first index, and performs entropy coding on the third flag indicating whether the CCALF process is enabled for the current block using the context model indicated by the second index.

[0041] According to another aspect, a decoder is provided that includes circuitry that, in operation, parses a first flag indicating whether the CCALF process is enabled for a first block that is adjacent to the left side of a current block. The circuitry parses a second flag indicating whether the CCALF process is enabled for a second block that is adjacent to the upper side of the current block. The circuitry determines a first index associated with a color component of the current block. The circuitry derives a second index indicating a context model using the first flag, the second flag, and the first index. The circuitry performs entropy decoding on a third flag indicating whether the CCALF process is enabled for the current block using the context model indicated by the second index, and in response to the third flag indicating that the CCALF process is enabled for the current block, performs the CCALF process on the current block.

[0042] In video coding techniques, new methods need to be proposed to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of embodiments of the present disclosure, including elements of embodiments of the present disclosure considered alone or in various combinations, may facilitate one or more of the following: improvement of coding efficiency, enhancement of image quality, reduction of utilization of processing resources associated with encoding / decoding, reduction of circuit size, improvement of encoding / decoding processing speed, and the like.

[0043] In addition, some implementations of embodiments of the present disclosure, including elements of embodiments of the present disclosure considered alone or in various combinations, may facilitate the appropriate selection of one or more elements in encoding and decoding, such as filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations. Note that the present disclosure includes the disclosure of configurations and methods that may provide advantages in addition to the above advantages. Examples of such configurations and methods include configurations or methods for increasing coding efficiency while reducing processing resource usage.

[0044] Additional benefits and advantages of the disclosed embodiments will become apparent from the specification and the drawings. The benefits and / or advantages may be obtained individually from the various embodiments and features of the specification and the drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.

[0045] It should be noted that a general or specific embodiment may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Figure 1 is a schematic diagram showing an example of a functional configuration of a transmission system according to an embodiment.

[0047] Figure 2 Figure 2 is a conceptual diagram for showing an example of a hierarchical structure of data in a stream.​​​​

[0048] Figure 3 Figure 3 is a conceptual diagram for showing an example of a slice configuration.

[0049] Figure 4 Figure 4 is a conceptual diagram for showing an example of a tile configuration.

[0050] Figure 5 Figure 5 is a conceptual diagram for showing an example of an encoding structure in scalable coding.

[0051] Figure 6 Figure 6 is a conceptual diagram for showing an example of an encoding structure in scalable coding.

[0052] Figure 7 Figure 7 is a block diagram showing the functional configuration of an encoder according to an embodiment.

[0053] Figure 8 Figure 8 is a functional block diagram showing an example of the installation of an encoder.

[0054] Fig. 9 Fig. 9 is a flowchart indicating an example of the overall encoding process performed by the encoder.

[0055] Fig.10 Fig.10 is a conceptual diagram showing an example of block splitting.

[0056] Fig.11 Fig.11 is a block diagram showing an example of the functional configuration of a splitter according to an embodiment.

[0057] Fig.12 Fig.12 is a conceptual diagram for showing an example of a splitting pattern.

[0058] Fig.13A Fig.13A is a conceptual diagram for showing an example of the syntax tree of a splitting pattern.

[0059] Fig. 13B Fig. 13B is a conceptual diagram for showing another example of the syntax tree of a splitting pattern.

[0060] Fig.14 Fig.14 ​​​​​​​​​​​​​​​​​​​​​​​​​​It is a diagram indicating example transform basis functions for various transform types.

[0061] Fig.15 Fig.15 It is a conceptual diagram for illustrating the concept of an example spatial-variation transform (SVT).

[0062] Fig.16 Fig.16 It is a flowchart showing an example of the process performed by a transducer.

[0063] Fig.17 Fig.17 It is a flowchart showing another example of the process performed by a transducer.

[0064] Fig.18 Fig.18 It is a block diagram showing an example of the functional configuration of a quantizer according to an embodiment.

[0065] Fig.19 Fig.19 It is a flowchart showing an example of the quantization process performed by a quantizer.

[0066] Fig. 20 Fig. 20 It is a block diagram showing an example of the functional configuration of an entropy encoder according to an embodiment.

[0067] Fig.21 Fig.21 It is a conceptual diagram for illustrating an example process of context-based adaptive binary arithmetic coding (CABAC) in an entropy encoder.

[0068] Fig. 22 Fig. 22 It is a block diagram showing an example of the functional configuration of a loop filter according to an embodiment.

[0069] Fig.23A Fig.23A It is a conceptual diagram for showing an example of the filter shape used in an adaptive loop filter (ALF).

[0070] Fig. 23B Fig. 23B It is a conceptual diagram for showing another example of the filter shape used in an ALF.

[0071] Fig.23C Fig.23C It is a conceptual diagram for showing another example of the filter shape used in an ALF.

[0072] Fig.23D Fig.23D ​​​​​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram for showing an example process of cross-component ALF (CC-ALF).

[0073] Fig.23E Fig.23E It is a conceptual diagram for showing an example of the filter shape used in CC-ALF.

[0074] Fig.23F Fig.23F It is a conceptual diagram for showing an example process of joint chroma CCALF (JC-CCALF).

[0075] Figure 23G Figure 23G It is a table showing example weight index candidates that can be adopted in JC-CCALF.

[0076] Fig.24 Fig.24 It is a block diagram showing an example of the specific configuration of a loop filter used as a deblocking filter (DBF).

[0077] Fig.25 Fig.25 It is a conceptual diagram for showing an example of a deblocking filter having symmetric filtering characteristics with respect to a block boundary.

[0078] Fig.26 Fig.26 It is a conceptual diagram for showing a block boundary where a deblocking filtering process is performed.

[0079] Fig. 27 Fig. 27 It is a conceptual diagram for showing an example of a boundary strength (Bs) value.

[0080] Fig.28 Fig.28 It is a flowchart showing an example of the process performed by the predictor of an encoder.

[0081] Fig.29 Fig.29 It is a flowchart showing another example of the process performed by the predictor of an encoder.

[0082] Fig.30 Fig.30 It is a flowchart showing another example of the process performed by the predictor of an encoder.

[0083] Fig.31 Fig.31 It is a conceptual diagram for showing sixty-seven intra prediction modes used in intra prediction in an embodiment.

[0084] Fig.32 Fig.32 ​​​​​​​​​​​​​​​​​​​​​​​​It is a flowchart showing an example of the process performed by the intra predictor.

[0085] Fig.33 Fig.33 It is a conceptual diagram for showing an example of a reference picture.

[0086] Fig.34 Fig.34 It is a conceptual diagram for showing an example of a reference picture list.

[0087] Fig.35 Fig.35 It is a flowchart showing an example of the basic process flow of inter prediction.

[0088] Fig.36 Fig.36 It is a flowchart showing an example of the derivation process of a motion vector.

[0089] Fig.37 Fig.37 It is a flowchart showing another example of the derivation process of a motion vector.

[0090] Fig.38A Fig.38A It is a conceptual diagram for showing an example characterization of the mode for MV derivation.

[0091] Fig.38B Fig.38B It is a conceptual diagram for showing an example characterization of the mode for MV derivation.

[0092] Fig.39 Fig.39 It is a flowchart showing an example of the inter prediction process in a conventional inter mode.

[0093] Fig.40 Fig.40 It is a flowchart showing an example of the inter prediction process in a conventional merge mode.

[0094] Fig.41 Fig.41 It is a conceptual diagram for showing an example of the motion vector derivation process in the merge mode.

[0095] Fig.42 Fig.42 It is a conceptual diagram for showing an example of the MV derivation process for the current picture through the HMVP merge mode.

[0096] Fig.43 Fig.43 It is a flowchart showing an example of the frame rate up-conversion (FRUC) process.

[0097] ​​​​​​​​​​​​​​​​​​​​​​​​​ Fig.44 Fig.44 It is a conceptual diagram showing an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.

[0098] Fig.45 Fig.45 It is a conceptual diagram showing an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture.

[0099] Fig.46A Fig.46A It is a conceptual diagram showing an example of deriving the motion vector of each sub-block based on the motion vectors of multiple adjacent blocks.

[0100] Fig.46B Fig.46B It is a conceptual diagram showing an example of deriving the motion vector of each sub-block in an affine mode using three control points.

[0101] Fig.47A Fig.47A It is a conceptual diagram showing an example of MV derivation at control points in an affine mode.

[0102] Fig.47B Fig.47B It is a conceptual diagram showing an example of MV derivation at control points in an affine mode.

[0103] Fig.47C Fig.47C It is a conceptual diagram showing an example of MV derivation at control points in an affine mode.

[0104] Fig.48A Fig.48A It is a conceptual diagram showing an affine mode using two control points.

[0105] Fig.48B Fig.48B It is a conceptual diagram showing an affine mode using three control points.

[0106] Fig.49A Fig.49A It is a conceptual diagram showing an example of a method for MV derivation at control points when the number of control points for an encoded block and the number of control points for the current block are different from each other.

[0107] Fig.49B Fig.49B It is a conceptual diagram showing another example of a method for MV derivation at control points when the number of control points for an encoded block and the number of control points for the current block are different from each other.

[0108] ​​​​​​​​​​​​​​​​​​​​​​ Fig.50 Fig.50 It is a flowchart showing an example of the process in the affine merge mode.

[0109] Fig.51 Fig.51 It is a flowchart showing an example of the process in the affine inter - frame mode.

[0110] Fig.52A Fig.52A It is a conceptual diagram for showing the generation of two triangular prediction images.

[0111] Fig.52B Fig.52B It is a conceptual diagram for showing an example of the first part of the first partition overlapping with the second partition and the first sample set and the second sample set that can be weighted as part of the correction process.

[0112] Fig.52C Fig.52C It is a conceptual diagram for showing the first part of the first partition, which is the part of the first partition overlapping with a part of the adjacent partition.

[0113] Fig.53 Fig.53 It is a flowchart showing an example of the process in the triangular mode.

[0114] Fig.54 Fig.54 It is a conceptual diagram for showing an example of the Advanced Temporal Motion Vector Prediction (ATMVP) mode in which the MV is derived in units of sub - blocks.

[0115] Fig.55 Fig.55 It is a flowchart showing the relationship between the merge mode and the Dynamic Motion Vector Refresh (DMVR).

[0116] Fig.56 Fig.56 It is a conceptual diagram for showing an example of the DMVR.

[0117] Fig.57 Fig.57 It is a conceptual diagram for showing another example of the DMVR for determining the MV.

[0118] Fig.58A Fig.58A It is a conceptual diagram for showing an example of the motion estimation in the DMVR.

[0119] Fig.58B ​​​​​​​​​​​​​​​​​​​​​​​ Fig.58B It is a flowchart showing an example of the motion estimation process in DMVR.

[0120] Fig.59 Fig.59 It is a flowchart showing an example of the generation process of a predicted image.

[0121] Fig.60 Fig.60 It is a flowchart showing another example of the generation process of a predicted image.

[0122] Fig.61 Fig.61 It is a flowchart showing an example of the correction process of a predicted image by overlapping block motion compensation (OBMC).

[0123] Fig.62 Fig.62 It is a conceptual diagram for showing an example of the predicted image correction process by OBMC.

[0124] Fig.63 Fig.63 It is a conceptual diagram for showing a model assuming uniform linear motion.

[0125] Fig.64 Fig.64 It is a flowchart showing an example of the inter-frame prediction process according to BIO.

[0126] Fig.65 Fig.65 It is a functional block diagram showing an example of the functional configuration of an inter-frame predictor that can perform inter-frame prediction according to BIO.

[0127] Fig.66A Fig.66A It is a conceptual diagram for showing an example of the process of a predicted image generation method using the luminance correction process performed by LIC.

[0128] Fig.66B Fig.66B It is a flowchart showing an example of the process of a predicted image generation method using LIC.

[0129] Fig.67 Fig.67 It is a block diagram showing the functional configuration of a decoder according to an embodiment.

[0130] Fig.68 Fig.68 It is a functional block diagram showing an example of the installation of a decoder.

[0131] Fig.69 Fig.69 ​​​​​​​​​​​​​​​​​​​​​​​​It is a flowchart showing an example of the overall decoding process performed by the decoder.

[0132] Fig.70 Fig.70 It is a conceptual diagram for showing the relationship between the segmentation determiner and other constituent elements.

[0133] Fig.71 Fig.71 It is a block diagram showing an example of the functional configuration of the entropy decoder.

[0134] Fig.72 Fig.72 It is a conceptual diagram for showing an example of the CABAC process in the entropy decoder.

[0135] Fig.73 Fig.73 It is a block diagram showing an example of the functional configuration of the inverse quantizer.

[0136] Fig.74 Fig.74 It is a flowchart showing an example of the inverse quantization process performed by the inverse quantizer.

[0137] Fig.75 Fig.75 It is a flowchart showing an example of the process performed by the inverse transformer.

[0138] Fig.76 Fig.76 It is a flowchart showing another example of the process performed by the inverse transformer.

[0139] Fig.77 Fig.77 It is a block diagram showing an example of the functional configuration of the loop filter.

[0140] Fig.78 Fig.78 It is a flowchart showing an example of the process performed by the predictor of the decoder.

[0141] Fig.79 Fig.79 It is a flowchart showing another example of the process performed by the predictor of the decoder.

[0142] Fig.80A Fig.80A It is a flowchart showing another example of the process performed by the predictor of the decoder.

[0143] Fig.80B Fig.80B It is a flowchart showing another example of the process performed by the predictor of the decoder.

[0144] ​​​​​​​​​​​​​​​​​​​​​​​​​ Fig.80C Fig.80C It is a flowchart showing another example of the process performed by the predictor of the decoder.

[0145] Fig.81 Fig.81 It is a diagram showing an example of the process performed by the intra predictor of the decoder.

[0146] Fig.82 Fig.82 It is a flowchart showing an example of the MV derivation process in the decoder.

[0147] Fig.83 Fig.83 It is a flowchart showing another example of the MV derivation process in the decoder.

[0148] Fig.84 Fig.84 It is a flowchart showing an example of the process of inter prediction through the regular inter mode in the decoder.

[0149] Fig.85 Fig.85 It is a flowchart showing an example of the process of inter prediction through the regular merge mode in the decoder.

[0150] Fig.86 Fig.86 It is a flowchart showing an example of the process of inter prediction through the FRUC mode in the decoder.

[0151] Fig.87 Fig.87 It is a flowchart showing an example of the process of inter prediction through the affine merge mode in the decoder.

[0152] Fig.88 Fig.88 It is a flowchart showing an example of the process of inter prediction through the affine inter mode in the decoder.

[0153] Fig.89 Fig.89 It is a flowchart showing an example of the process of inter prediction through the triangle mode in the decoder.

[0154] Fig.90 Fig.90 It is a flowchart showing an example of the process of motion estimation through DMVR in the decoder.

[0155] Fig.91 Fig.91 It is a flowchart showing an example process of motion estimation through DMVR in the decoder.

[0156] ​​​​​​​​​​​​​​​​​​​​​​​​ Fig.92 Fig.92 It is a flowchart showing an example of the process of generating a predicted image in a decoder.

[0157] Fig.93 Fig.93 It is a flowchart showing another example of the process of generating a predicted image in a decoder.

[0158] Fig.94 Fig.94 It is a flowchart showing an example of the process of correcting a predicted image by OBMC in a decoder.

[0159] Fig.95 Fig.95 It is a flowchart showing an example of the process of correcting a predicted image by BIO in a decoder.

[0160] Fig.96 Fig.96 It is a flowchart showing an example of the process of correcting a predicted image by LIC in a decoder.

[0161] Fig.97 Fig.97 It is a flowchart of a sample process flow for decoding an image by applying a CCALF (Cross Component Adaptive Loop Filter) process according to a first aspect.

[0162] Fig.98 Fig.98 It is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment.

[0163] Fig.99 Fig.99 It is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment.

[0164] Fig.100 Fig.100 It is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment.

[0165] Fig.101 Fig.101 It is a block diagram showing the functional configurations of an encoder and a decoder according to an embodiment.

[0166] Fig.102 Fig.102 It is a flowchart of a sample process flow for decoding an image by applying a CCALF process according to a second aspect.

[0167] Fig.103A Fig.103A ​​​​​​​​​​​​​​​​​​​​​​​Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0168] Fig.103B Fig.103B Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0169] Fig.103C Fig.103C Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0170] Fig.103D Fig.103D Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0171] Fig.103E Fig.103E Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0172] Fig.103F Fig.103F Illustrates the sample positions of the cropping parameters to be parsed from, for example, VPS, APS, SPS, PPS, slice headers, CTUs or TUs of a bitstream.

[0173] Fig.104 Fig.104 (i)-(iii) Illustrate examples of the cropping parameters.

[0174] Fig.105 Fig.105 Is a flowchart of a sample process flow for decoding an image using filter coefficients in the CCALF process according to a third aspect.

[0175] Fig.106A Fig.106A Is a conceptual diagram indicating an example of the position of the filter coefficients to be used in the CCALF process.

[0176] Fig.106B Fig.106B Is a conceptual diagram indicating an example of the position of the filter coefficients to be used in the CCALF process.

[0177] Fig.106C Fig.106C ​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0178] Fig.106D Fig.106D It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0179] Fig.106E Fig.106E It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0180] Fig.106F Fig.106F It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0181] Figure 106G Figure 106G It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0182] Fig.106H Fig.106H It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0183] Fig.107A Fig.107A It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0184] Fig.107B Fig.107B It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0185] Fig.107C Fig.107C It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0186] Fig.107D Fig.107D It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0187] Fig.107E Fig.107E It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0188] Fig.107F Fig.107F It is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process. ​​​​​​​​​​​​​​​​​​​​​​

[0189] Figure 107G Figure 107G is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0190] Figure 107H Figure 107H is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0191] Figure 108A Figure 108A is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0192] Figure 108B Figure 108B is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0193] Figure 108C Figure 108C is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0194] Figure 108D Figure 108D is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0195] Figure 108E Figure 108E is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0196] Figure 108F Figure 108F is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0197] Figure 108G Figure 108G is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0198] Figure 108H Figure 108H is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0199] Figure 109A Figure 109A is a conceptual diagram showing an example of the position of filter coefficients to be used in the CCALF process.

[0200] Figure 109B Figure 109B ​​​​​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0201] Figure 109C Figure 109C It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0202] Figure 109D Figure 109D It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0203] Figure 110A Figure 110A It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0204] Figure 110B Figure 110B It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0205] Figure 110C Figure 110C It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0206] Figure 110D Figure 110D It is a conceptual diagram showing an example of the positions of filter coefficients to be used in the CCALF process.

[0207] Figure 111 Figure 111 It is a conceptual diagram showing a further example of the positions of filter coefficients to be used in the CCALF process.

[0208] Figure 112 Figure 112 It is a conceptual diagram showing a further example of the positions of filter coefficients to be used in the CCALF process.

[0209] Figure 113 Figure 113 It is a block diagram showing the functional configuration of the CCALF process performed by an encoder and a decoder according to an embodiment.

[0210] Figure 114 Figure 114 It is a flowchart of a sample process flow for decoding an image by applying the CCALF process using a filter selected from a plurality of filters according to a fourth aspect.

[0211] Figure 115 Figure 115 It illustrates an example of a process flow for selecting a filter.​​​​​​​​​​​​​​​​​​​​​​

[0212] Figure 116-1A Figure 116-1A An example of a filter is illustrated.

[0213] Figure 116-1B Figure 116-1B An example of a filter is illustrated.

[0214] Figure 116-1C Figure 116-1C An example of a filter is illustrated.

[0215] Figure 116-1D Figure 116-1D An example of a filter is illustrated.

[0216] Figure 116-1E Figure 116-1E An example of a filter is illustrated.

[0217] Figure 116-1F Figure 116-1F An example of a filter is illustrated.

[0218] Figure 116-1G Figure 116-1G An example of a filter is illustrated.

[0219] Figure 116-1H Figure 116-1H An example of a filter is illustrated.

[0220] Figure 116-1I Figure 116-1I An example of a filter is illustrated.

[0221] Figure 117-2A Figure 117-2A An example of a filter is illustrated.

[0222] Figure 117-2B Figure 117-2B An example of a filter is illustrated.

[0223] Figure 117-2C Figure 117-2C An example of a filter is illustrated.

[0224] Figure 117-2D Figure 117-2D An example of a filter is illustrated.

[0225] Figure 117-2E Figure 117-2E An example of a filter is illustrated.

[0226] Figure 117-2F Figure 117-2F An example of a filter is illustrated.

[0227] Figure 117-2G ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​Figure 117-2G A diagram illustrates an example of a filter.

[0228] Figure 117-2H Figure 117-2H A diagram illustrates an example of a filter.

[0229] Figure 117-2I Figure 117-2I A diagram illustrates an example of a filter. Figure 118 Figure 118 is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the fifth aspect.

[0230] Figure 119 Figure 119 (i)-(iv) A diagram illustrates an example of the number of coefficients (NumCoeff) to be parsed from a bitstream.

[0231] Figure 120 Figure 120 is a flowchart of a sample process flow for decoding an image by applying the CCALF process using parameters according to the sixth aspect.

[0232] Figure 121 Figure 121 is a conceptual diagram showing an example of generating a CCALF value for the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.

[0233] Figure 122 Figure 122 is a conceptual diagram showing an example of generating a CCALF value for the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.

[0234] Figure 123 Figure 123 is a conceptual diagram showing an example of generating a CCALF value for the luminance component of a current chrominance sample by calculating the weighted average of adjacent samples.

[0235] Figure 124 Figure 124 is a conceptual diagram showing an example of generating a CCALF value for the luminance component of a current sample by calculating the weighted average of adjacent samples, where the positions of the adjacent samples are adaptively determined as the chrominance type.

[0236] Figure 125 Figure 125 is a conceptual diagram showing an example of generating a CCALF value for the luminance component of a current sample by calculating the weighted average of adjacent samples, where the positions of the adjacent samples are determined adaptively to the chrominance type.

[0237] Figure 126 Figure 126 ​​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram showing an example of generating a CCALF value for a luminance component by applying a bit shift to an output value of a weighted calculation.

[0238] Figure 127 Figure 127 It is a conceptual diagram showing an example of generating a CCALF value for a luminance component by applying a bit shift to an output value of a weighted calculation.

[0239] Figure 128 Figure 128 It is a flowchart of a sample process flow for decoding an image by applying a CCALF process using parameters according to the seventh aspect.

[0240] Figure 129A Figure 129A Illustrates the sample positions of one or more parameters to be parsed from a bitstream, where the one or more parameters may include a first parameter, a second parameter, or both.

[0241] Figure 129B Figure 129B Illustrates the sample positions of one or more parameters to be parsed from a bitstream, where the one or more parameters may include a first parameter, a second parameter, or both.

[0242] Figure 129C Figure 129C Illustrates the sample positions of one or more parameters to be parsed from a bitstream, where the one or more parameters may include a first parameter, a second parameter, or both.

[0243] Figure 129D Figure 129D Illustrates the sample positions of one or more parameters to be parsed from a bitstream, where the one or more parameters may include a first parameter, a second parameter, or both.

[0244] Figure 129E Figure 129E Illustrates the sample positions of one or more parameters to be parsed from a bitstream, where the one or more parameters may include a first parameter, a second parameter, or both.

[0245] Figure 130A Figure 130A Shows a sample process for retrieving one or more parameters, where the one or more parameters may include a first parameter, a second parameter, or both.

[0246] Figure 130B Figure 130B Shows a sample process for retrieving one or more parameters, where the one or more parameters may include a first parameter, a second parameter, or both.

[0247] Figure 130C Figure 130C ​​​​​​​​​​​​​​​​​​​​Shows a sample process for retrieving one or more parameters, which may include a first parameter, a second parameter, or both.

[0248] Figure 130D Figure 130D Shows a sample process for retrieving one or more parameters, which may include a first parameter, a second parameter, or both.

[0249] Figure 131A Figure 131A Shows a sample value of the second parameter.

[0250] Figure 131B Figure 131B Shows a sample value of the second parameter.

[0251] Figure 131C Figure 131C Shows a sample value of the second parameter.

[0252] Figure 132 Figure 132 Shows an example of parsing the second parameter using arithmetic coding.

[0253] Figure 133 Figure 133 Is a conceptual diagram of a variant of the present embodiment applied to rectangular partitions and non-rectangular partitions (such as triangular partitions).

[0254] Figure 134 Figure 134 Is a flowchart of an example process flow for decoding an image by applying the CCALF process using a parameter according to the eighth aspect.

[0255] Figure 135 Figure 135 Is a flowchart of a sample process flow for decoding an image by applying the CCALF process using a parameter according to the eighth aspect.

[0256] Figure 136 Figure 136 Shows example positions of chroma sample types 0 to 5.

[0257] Figure 137A Figure 137A Is a conceptual diagram showing sample symmetric filling.

[0258] Figure 137B Figure 137B Is a conceptual diagram showing sample symmetric filling.

[0259] Figure 137C Figure 137C Is a conceptual diagram showing sample symmetric filling.

[0260] ​​​​​​​​​​​​​​​​​​​​​​​​​Figure 137D Figure 137D It is a conceptual diagram showing symmetric filling of the sample.

[0261] Figure 138 Figure 138 It is a conceptual diagram showing symmetric filling of the sample.

[0262] Figure 139 Figure 139 It is a conceptual diagram showing symmetric filling of the sample.

[0263] Figure 140A Figure 140A It is a conceptual diagram showing asymmetric filling of the sample.

[0264] Figure 140B Figure 140B It is a conceptual diagram showing asymmetric filling of the sample.

[0265] Figure 140C Figure 140C It is a conceptual diagram showing asymmetric filling of the sample.

[0266] Figure 140D Figure 140D It is a conceptual diagram showing asymmetric filling of the sample.

[0267] Figure 141 Figure 141 It is a conceptual diagram showing asymmetric filling of the sample.

[0268] Figure 142 Figure 142 It is a conceptual diagram showing asymmetric filling of the sample.

[0269] Figure 143 Figure 143 It is a conceptual diagram showing asymmetric filling of the sample.

[0270] Figure 144A Figure 144A It is a conceptual diagram showing further asymmetric filling of the sample.

[0271] Figure 144B Figure 144B It is a conceptual diagram showing further asymmetric filling of the sample.

[0272] Figure 144C Figure 144C It is a conceptual diagram showing further asymmetric filling of the sample.

[0273] Figure 144D Figure 144D It is a conceptual diagram showing further asymmetric filling of the sample.

[0274] Figure 145 Figure 145 ​​​​​​​​​​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram showing further asymmetric filling of the sample.

[0275] Figure 146 Figure 146 It is a conceptual diagram showing further asymmetric filling of the sample.

[0276] Figure 147 Figure 147 It is a conceptual diagram showing further asymmetric filling of the sample.

[0277] Figure 148A Figure 148A It is a conceptual diagram showing further symmetric filling of the sample.

[0278] Figure 148B Figure 148B It is a conceptual diagram showing further symmetric filling of the sample.

[0279] Figure 148C Figure 148C It is a conceptual diagram showing further symmetric filling of the sample.

[0280] Figure 148D Figure 148D It is a conceptual diagram showing further symmetric filling of the sample.

[0281] Figure 149 Figure 149 It is a conceptual diagram showing further symmetric filling of the sample.

[0282] Figure 150 Figure 150 It is a conceptual diagram showing further symmetric filling of the sample.

[0283] Figure 151A Figure 151A It is a conceptual diagram showing further asymmetric filling of the sample.

[0284] Figure 151B Figure 151B It is a conceptual diagram showing further asymmetric filling of the sample.

[0285] Figure 151C Figure 151C It is a conceptual diagram showing further asymmetric filling of the sample.

[0286] Figure 152 Figure 152 It is a conceptual diagram showing further asymmetric filling of the sample.

[0287] Figure 153 Figure 153 It is a conceptual diagram showing further asymmetric filling of the sample.

[0288] Figure 154 Fig.154 ​​​​​​​​​​​​​​​​​​​​​​​​​​​​It is a conceptual diagram showing further sample asymmetric filling.

[0289] Fig.155A Fig.155A It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0290] Fig.155B Fig.155B It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0291] Fig.155C Fig.155C It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0292] Fig.155D Fig.155D It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0293] Fig.155E Fig.155E It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0294] Fig.155F Fig.155F It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0295] Figure 155G Figure 155G It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0296] Figure 155H Figure 155H It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0297] Fig.155I Fig.155I It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0298] Fig.155J Fig.155J It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0299] Figure 155K Figure 155K It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0300] Figure 155L Figure 155L It illustrates a further example of filling with horizontal and vertical virtual boundaries.

[0301] Fig.156 ​​​​​​​​​​​​​​​​​​​​​​​​​​ Fig.156 is a block diagram showing the functional configurations of an encoder and a decoder according to an example, where symmetric padding is used for virtual boundary positions of ALF and symmetric or asymmetric padding is used for virtual boundary positions of CC-ALF.

[0302] Fig.157 Fig.157 is a block diagram showing the functional configurations of an encoder and a decoder according to another example, where symmetric padding is used for virtual boundary positions of ALF and unilateral padding is used for virtual boundary positions of CC-ALF.

[0303] Fig.158A Fig.158A is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0304] Fig.158B Fig.158B is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0305] Fig.158C Fig.158C is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0306] Fig.158D Fig.158D is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0307] Fig.158E Fig.158E is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0308] Fig.158F Fig.158F is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0309] Figure 158G Figure 158G is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0310] Figure 158H Figure 158H is a conceptual diagram showing an example of unilateral padding with horizontal or vertical virtual boundaries.

[0311] Fig.159A Fig.159A is a conceptual diagram showing an example of unilateral padding with both horizontal and vertical virtual boundaries.

[0312] Fig.159B Fig.159B ​​​​​​​​​​​​​​​​​​​​​​Is a conceptual diagram showing an example of unilateral padding with horizontal and vertical virtual boundaries.

[0313] Fig.159C Fig.159C Is a conceptual diagram showing an example of unilateral padding with horizontal and vertical virtual boundaries.

[0314] Fig.160 Fig.160 Is a flowchart of a sample process flow for decoding an image applying the CCALF process using parameters according to the ninth aspect.

[0315] Fig.161 Fig.161 Is a conceptual diagram showing an example of a filter to be applied in the CCALF process.

[0316] Fig.162 Fig.162 Describes a sample equation of the filtering process.

[0317] Fig.163 Fig.163 Is a conceptual diagram showing an example of the syntax of CCALF.

[0318] Fig.164 Fig.164 Is a conceptual diagram showing an example of signaling filter coefficient values using an Exponential-Golomb code with a fixed order k (denoted as EGk).

[0319] Fig.165A Fig.165A Is a conceptual diagram showing an example of EGk applied to filter coefficients.

[0320] Fig.165B Fig.165B Is a conceptual diagram showing an example of EGk applied to filter coefficients.

[0321] Fig.165C Fig.165C Is a conceptual diagram showing an example of EGk applied to filter coefficients.

[0322] Fig.165D Fig.165D Is a conceptual diagram showing an example of EGk applied to filter coefficients.

[0323] Fig.166A Fig.166A Is a conceptual diagram showing an example of EGk applied to filter coefficients.

[0324] Fig.166B Fig.166B Is a conceptual diagram showing an example of EGk applied to filter coefficients.​​​​​​​​​​​​​​​​​​​​​​​​

[0325] Fig.167A Fig.167A is a conceptual diagram of an example of EGk applied to filter coefficients.

[0326] Fig.167B Fig.167B is a conceptual diagram of an example of EGk applied to filter coefficients.

[0327] Fig.168A Fig.168A is a conceptual diagram of an example of EGk applied to filter coefficients.

[0328] Fig.168B Fig.168B is a conceptual diagram of an example of EGk applied to filter coefficients.

[0329] Fig.169 Fig.169 is a conceptual diagram of an example of the syntax of parameters used in the ALF process.

[0330] Fig.170 Fig.170 is a conceptual diagram of an example of the syntax of parameters used in the CCALF process.

[0331] Fig.171 Fig.171 is a flowchart of an example of the process flow for decoding an image to which the CCALF process is applied using a set of coefficients.

[0332] Fig.172A Fig.172A is a conceptual diagram of an example of the shape of a filter coefficient group.

[0333] Fig.172B Fig.172B is a conceptual diagram of an example of the shape of a filter coefficient group.

[0334] Fig.172C Fig.172C is a conceptual diagram of an example of the shape of a filter coefficient group.

[0335] Fig.172D Fig.172D is a conceptual diagram of an example of the shape of a filter coefficient group.

[0336] Fig.173A Fig.173A is a conceptual diagram of an example of the positions of the reconstructed samples of the first and second components.

[0337] Fig.173B Fig.173B is a conceptual diagram of an example of the positions of the reconstructed samples of the first and second components.​​​​​​​​​​​​​​​​​​​​​​​​​​

[0338] Fig.173C Fig.173C It is a conceptual diagram showing examples of the positions of the reconstructed samples of the first and second components.

[0339] Fig.173D Fig.173D It is a conceptual diagram showing examples of the positions of the reconstructed samples of the first and second components.

[0340] Fig.174 Fig.174 It is a flowchart showing an example of the process flow for decoding an image to which the CCALF process is applied using a context model based on selection of other blocks.

[0341] Figure 175A Figure 175A It is a conceptual diagram showing examples of the positions of the first and second blocks.

[0342] Fig.175B Fig.175B It is a conceptual diagram showing examples of the positions of the first and second blocks.

[0343] Fig.175C Fig.175C It is a conceptual diagram showing examples of the positions of the first and second blocks.

[0344] Fig.175D Fig.175D It is a conceptual diagram showing examples of the positions of the first and second blocks.

[0345] Fig.176 Fig.176 It is a conceptual diagram showing another example of the positions of the first and second blocks.

[0346] Fig.177 Fig.177 It is a table showing an example of the equation for calculating ctxIdx.

[0347] Fig.178 Fig.178 It is a table showing examples of the initValue and shiftIdx of ctxIdx for the third flag.

[0348] Fig.179 Fig.179 It is a table showing examples of the calculated ctxIdx.

[0349] Fig.180 Fig.180 It is a table showing another example of the initValue and shiftIdx of ctxIdx for the third flag.

[0350] Fig.181 ​​​​​​​​​​​​​​​​​​​​​​​​​​ Fig.181 A table of another example for calculating ctxIdx.

[0351] Fig.182 Fig.182 A diagram showing an example of the overall configuration of a content providing system for implementing a content distribution service.

[0352] Fig.183 Fig.183 A conceptual diagram showing an example of the display screen of a web page.

[0353] Fig.184 Fig.184 A conceptual diagram showing an example of the display screen of a web page.

[0354] Fig.185 Fig.185 A block diagram showing an example of a smartphone.

[0355] Fig.186 Fig.186 A block diagram showing an example of the functional configuration of a smartphone. Detailed Description of the Invention

[0356] In the drawings, unless otherwise indicated by the context, the same reference numerals denote similar elements. The sizes and relative positions of the elements in the drawings are not necessarily drawn to scale.

[0357] Hereinafter, embodiments will be described with reference to the drawings. Note that each of the embodiments described below shows a general or specific example. The numerical values, shapes, materials, components, arrangements and connections of components, steps, relationships and orders of steps, etc. indicated in the following embodiments are only examples and are not intended to limit the scope of the claims.

[0358] Hereinafter, embodiments of an encoder and a decoder will be described. The embodiments are examples of the encoder and the decoder, and the processes and / or configurations presented in the description of the aspects of the present disclosure can be applied to the encoder and the decoder. The processes and / or configurations can also be implemented in encoders and decoders different from those according to the embodiments. For example, regarding the processes and / or configurations applied to the embodiments, any of the following can be implemented:

[0359] (1) Any component of the encoder or decoder according to the embodiments presented in the description of the aspects of the present disclosure can be replaced with or combined with another component presented anywhere in the description of the aspects of the present disclosure.

[0360] ​​​​​​​​​​(2) In an encoder or decoder according to an embodiment, any change can be made to the functions or processes performed by one or more components of the encoder or decoder, such as addition, replacement, removal, etc. of the functions or processes. For example, any function or process can be replaced or combined with another function or process presented anywhere in the description of the aspects of the present disclosure.

[0361] (3) In a method implemented by an encoder or decoder according to an embodiment, any change can be made, such as addition, replacement, and removal of one or more processes included in the method. For example, any process in the method can be replaced or combined with another process presented anywhere in the description of the aspects of the present disclosure.

[0362] (4) One or more components included in an encoder or decoder according to an embodiment can be combined with components presented anywhere in the description of the aspects of the present disclosure content, can be combined with components including one or more functions presented anywhere in the description of the aspects of the present disclosure, and can be combined with components implementing one or more processes implemented by the components presented in the description of the aspects of the present disclosure.

[0363] (5) A component including one or more functions of an encoder or decoder according to an embodiment, or a component implementing one or more processes of an encoder or decoder according to an embodiment, can be combined with or replaced by components presented anywhere in the description of the aspects of the present disclosure, can be combined with or replaced by components including one or more functions presented anywhere in the description of the aspects of the present disclosure, or can be combined with or replaced by components implementing one or more processes presented anywhere in the description of the aspects of the present disclosure.

[0364] (6) In a method implemented by an encoder or decoder according to an embodiment, any process included in the method can be replaced or combined with a process presented anywhere in the description of the aspects of the present disclosure or with any corresponding or equivalent process.

[0365] (7) One or more processes included in a method implemented by an encoder or decoder according to an embodiment can be combined with processes presented anywhere in the description of the aspects of the present disclosure content.

[0366] (8) The implementation manners of the processes and / or configurations presented in the description of the aspects of the present disclosure are not limited to an encoder or decoder according to an embodiment. For example, the processes and / or configurations can be implemented in a device for purposes different from those of the motion image encoder or motion image decoder disclosed in the embodiment.

[0367] (Term Definition)

[0368] The corresponding terms can be defined as indicated by the following examples.

[0369] An image is a data unit configured with a set of pixels, is a picture, or includes blocks smaller than pixels. In addition to video, an image also includes still images.

[0370] A picture is an image processing unit configured with a set of pixels and can also be referred to as a frame or a field. For example, a picture can take the form of an array of luminance samples in a monochrome format or an array of luminance samples and two corresponding arrays of chrominance samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0371] A block is a processing unit that is a set of a determined number of pixels. A block can have any number of different shapes. For example, a block can have a rectangle of M×N (M columns × N rows) pixels, a square of M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, shards, bricks, CTUs, superblocks, basic segmentation units, VPDUs, processing segmentation units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and sub-blocks. A block can take the form of an M×N sample array or an M×N transform coefficient array. For example, a block can be a square or rectangular pixel region including a luminance matrix and two chrominance matrices.

[0372] A pixel or a sample is the smallest point of an image. A pixel or a sample includes pixels at integer positions and pixels at sub-pixel positions, such as those generated based on pixels at integer positions.

[0373] A pixel value or a sample value is a characteristic value of a pixel. A pixel value or a sample value can include one or more of a luminance value, a chrominance value, an RGB gray level, a depth value, a binary value of zero or 1, etc.

[0374] Chroma or chrominance is the intensity of a color, usually represented by the symbols Cb and Cr, which specify the value of an array of samples or the value of a single sample representing one of two color difference signals related to the primary colors.

[0375] Luma or luminance is the brightness of an image, usually represented by the symbol or subscript Y or L, which specify the value of an array of samples or the value of a single sample representing a monochrome signal related to the primary colors.

[0376] A flag includes one or more bits indicating a value of, for example, a parameter or an index. A flag can be a binary flag that indicates the binary value of the flag, and it can also indicate a non-binary value of a parameter.

[0377] A signal conveys information that is symbolized or encoded into the signal. Signals include discrete digital signals and continuous analog signals.

[0378] A stream or bitstream is a digital data string of a digital data stream. A stream or bitstream can be a single stream or can be configured with multiple streams having multiple hierarchical layers. A stream or bitstream can be transmitted in serial communication using a single transmission path or can be transmitted in packet communication using multiple transmission paths.

[0379] Difference refers to various mathematical differences, such as simple difference (x - y), absolute value of difference (|x - y|), difference of squares (x^2 - y^2), square root of difference (√(x – y)), weighted difference (ax - by: a and b are constants), offset difference (x - y + a: a is an offset), etc. In the case of scalars, a simple difference is sufficient and difference calculation is included.

[0380] Sum refers to various mathematical sums, such as simple sum (x + y), absolute value of sum (|x + y|), sum of squares (x^2 + y^2), square root of sum (√(x + y)), weighted sum (ax + by: a and b are constants), offset sum (x + y + a: a is an offset), etc. In the case of scalars, a simple sum is sufficient and sum calculation is included.

[0381] A frame is a combination of a top field and a bottom field, where sampling lines 0, 2, 4,... are from the top field and sampling lines 1, 3, 5,... are from the bottom field.

[0382] A slice is an integer number of coding tree units in all subsequent dependent slices (if any) that are included between one independent slice segment and the next independent slice segment (if any) within the same access unit.

[0383] A tile is a rectangular region of coding tree blocks within a specific tile column and a specific tile row in a picture. A tile can be a rectangular region of a frame that is intended to be capable of independent decoding and encoding, although loop filtering across tile edges can still be applied.

[0384] A coding tree unit (CTU) can be a coding tree block of the luma samples of a picture having three sample arrays, or two corresponding coding tree blocks of the chroma samples. Alternatively, a CTU can be a coding tree block of the samples of a monochrome picture and a picture encoded using three separate color planes and a syntax structure for encoding samples. A superblock can be a 64×64 pixel square block consisting of 1 or 2 mode information blocks, or recursively divided into four 32×32 blocks, which can themselves be further divided.

[0385] (System configuration)

[0386] First, a transmission system according to an embodiment will be described. Figure 1 FIG. 1 is a schematic diagram showing an example of the configuration of a transmission system 400 according to an embodiment.

[0387] The transmission system 400 is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. As shown, the transmission system 400 includes an Figure 1 encoder 100, a network 300, and a decoder 200 as shown.

[0388] An image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image and outputs the stream to the network 300. The stream includes, for example, an encoded image and control information for decoding the encoded image. The image is compressed by encoding.

[0389] It should be noted that the image before being encoded by the encoder 100 is also referred to as a raw image, a raw signal, or a raw sample. The image can be a video or a still image. The image is a general concept of a sequence, a picture, and a block, and thus is not limited to a spatial region having a specific size and a temporal region having a specific size unless otherwise specified. The image is an array of pixels or pixel values, and a signal representing the image or the pixel value is also referred to as a sample. The stream can be referred to as a bitstream, an encoded bitstream, a compressed bitstream, or an encoded signal. In addition, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.

[0390] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 200 can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a two-way communication network and can be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasting and satellite broadcasting. Alternatively, the network 300 can be replaced by a recording medium such as a digital versatile disc (DVD) and a Blu-ray disc (BD) on which the stream is recorded.

[0391] The decoder 200 generates a decoded image as an uncompressed image by decoding, for example, the stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method adopted by the encoder 100.

[0392] It should be noted that the decoder 200 can also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 can also be referred to as a decoding method, an image decoding method, or a video decoding method.

[0393] (Data Structure)

[0394] Figure 2It is a conceptual diagram showing an example of the hierarchical structure of data in a stream. For convenience, reference will be made to Figure 1 transmission system 400 to describe Figure 2 . The stream includes, for example, a video sequence. As Figure 2 shown in (a) of

[0395] , the video sequence includes one or more video parameter sets (VPSs), one or more sequence parameter sets (SPSs), one or more picture parameter sets (PPSs), supplementary enhancement information (SEI), and a plurality of pictures.

[0396] In a video with multiple layers, the VPS may include encoding parameters shared between some of the multiple layers, as well as encoding parameters related to some of the multiple layers included in the video or related to a single layer.

[0397] The SPS includes parameters for the sequence, that is, the encoding parameters that the decoder 200 refers to for decoding the sequence. For example, the encoding parameters may indicate the width or height of a picture. It should be noted that there may be multiple SPSs.

[0397] The PPS includes parameters for the picture, that is, the encoding parameters that the decoder 200 refers to for decoding each picture in the sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. It should be noted that there may be multiple PPSs. Each of the SPS and PPS may be abbreviated as a parameter set.

[0398] As Figure 2 shown in (b) of

[0399] , the picture may include a picture header and one or more slices. The picture header includes the encoding parameters that the decoder 200 refers to for decoding one or more slices.

[0399] As Figure 2 shown in (c) of

[0400] , the slice includes a slice header and one or more tiles. The slice header includes the encoding parameters that the decoder 200 refers to for decoding one or more tiles.

[0400] As Figure 2 shown in (d) of

[0401] Note that a picture may not include any slices and may include slice groups instead of slices. In this case, the slice group includes at least one slice. In addition, a tile may include slices.

[0402] The CTU is also referred to as a superblock or a basis splitting unit. As Figure 2As shown in (e), the CTU includes a CTU header and at least one coding unit (CU). As shown in the figure, the CTU includes four coding units CU(10), CU(11), CU(12), and CU(13). The CTU header includes coding parameters that the decoder 200 references to decode at least one CU.

[0403] The CU can be divided into multiple smaller CUs. As shown in the figure, CU(10) is not divided into smaller coding units; CU(11) is divided into four smaller coding units CU(110), CU(111), CU(112), and CU(113); CU(12) is not divided into smaller coding units; and CU(13) is divided into seven smaller coding units CU(1310), CU(1311), CU(1312), CU(1313), CU(132), CU(133), and CU(134). As Figure 2 As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information used to predict the CU, and the residual coefficient information is information representing the prediction residual to be described later. Although the CU is basically the same as the prediction unit (PU) and the transform unit (TU), it should be noted that, for example, the sub-block transform (SBT) to be described later may include multiple TUs smaller than the CU. In addition, the CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. The VPDU is, for example, a fixed unit that can be processed in one stage when performing pipeline processing in hardware.

[0404] It should be noted that the stream may not include Figure 2 all the hierarchical layers shown. The order of the hierarchical layers can be exchanged, or any hierarchical layer can be replaced by another hierarchical layer. Here, the picture that is the target of the process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as the current picture. When the process is an encoding process, the current picture represents the current picture to be encoded, and when the process is a decoding process, the current picture represents the current picture to be decoded. Similarly, for example, the CU or CU block that is the target of the process to be performed by a device such as the encoder 100 or the decoder 200 is referred to as the current block. When the process is an encoding process, the current block represents the current block to be encoded, and when the process is a decoding process, the current block represents the current block to be decoded.

[0405] (Picture Structure: Slice / Tile)

[0406] The picture can be configured with one or more slice units or one or more tile units to facilitate parallel encoding / decoding of the picture.

[0407] A slice is a basic coding unit included in a picture. A picture may include, for example, one or more slices. In addition, a slice includes one or more coding tree units (CTUs).

[0408] Figure 3 is a conceptual diagram for showing an example of slice configuration. For example, in Figure 3 the picture includes 11×8 CTUs and is divided into four slices (slice 1 to 4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is the shape obtained by horizontally dividing the picture. The boundary of each slice does not need to coincide with the image end and may coincide with any boundary between CTUs in the image. The processing order (encoding order or decoding order) of CTUs in a slice is, for example, the raster scan order. A slice includes a slice header and encoded data. The characteristics of a slice may be written in the slice header. The characteristics may include the CTU address of the top CTU in the slice, slice type, etc.

[0409] A tile is a unit of a rectangular region included in a picture. Tiles of a picture may be assigned numbers called TileId in the raster scan order.

[0410] Figure 4 is a conceptual diagram for showing an example of tile configuration. For example, in Figure 4 the picture includes 11×8 CTUs and is divided into four tiles (tile 1 to 4) of rectangular regions. When using tiles, the processing order of CTUs may be different from the processing order in the case of not using tiles. When not using tiles, multiple CTUs in the picture are usually processed in the raster scan order. When using multiple tiles, at least one CTU in each of the multiple tiles is processed in the raster scan order. For example, as Figure 4 shown, the processing order of CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then continues from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0411] It should be noted that one tile may include one or more slices, and one slice may include one or more tiles. It should be noted that a picture may be configured with one or more tile sets. A tile set may include one or more tile groups, or one or more tiles. A picture may be configured with one of a tile set, a tile group, and a tile. For example, assume that the order of scanning multiple tiles for each tile set in the raster scan order is the basic encoding order of the tiles. Assume that a set of one or more tiles consecutive in the basic encoding order in each tile set is a tile group. Such a picture may be processed by a splitter 102 described later (see Figure 7 ) to configure.

[0412] (Scalable coding)

[0413] Figure 5 and Figure 6 is a conceptual diagram showing an example of a scalable stream structure, and for convenience will be described with reference to Figure 1 to describe.

[0414] As Figure 5 shown, the encoder 100 can generate a temporally / spatially scalable stream by dividing each of a plurality of pictures into any of a plurality of layers and encoding the pictures in the layers. For example, the encoder 100 encodes the pictures for each layer, thereby achieving scalability in the case where the enhancement layer exists above the base layer. This encoding of each picture is also referred to as scalable coding. In this way, the decoder 200 can switch the image quality of the image displayed by decoding the stream. In other words, the decoder 200 can determine which layer to decode based on internal factors such as the processing power of the decoder 200 and external factors such as the communication bandwidth state. As a result, the decoder 200 can decode the content while freely switching between low resolution and high resolution. For example, a user of the stream watches a video of the stream halfway through on a smartphone on the way home and continues to watch the video on a device (e.g., a TV connected to the Internet) at home. It should be noted that each of the above smartphone and device includes a decoder 200 with the same or different performance. In this case, when the device decodes the layer to a higher layer in the stream, the user can watch a high-quality video at home. In this way, the encoder 100 does not need to generate multiple streams with different image qualities of the same content, and thus can reduce the processing load.

[0415] In addition, the enhancement layer may include meta information based on statistical information about the image. The decoder 200 can generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the metadata. Super-resolution imaging may include, for example, an increase in the signal-to-noise ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information for identifying linear or non-linear filter coefficients used in the super-resolution process, or information for identifying parameter values in a filtering process, machine learning, or least squares method (used in the super-resolution process), etc.

[0416] In an embodiment, a configuration may be provided in which a picture is divided into slices, for example, according to the meaning of objects in the picture. In this case, the decoder 200 may decode only a partial area in the picture by selecting the slice to be decoded. In addition, the attributes of the objects (such as people, cars, balls, etc.) and the positions of the objects in the picture (coordinates in the same image) may be stored as metadata. In this case, the decoder 200 is able to identify the position of the desired object based on the metadata and determine the slice including the object. For example, as Figure 6 shown, a data storage structure different from the image data may be used to store the metadata, such as the SEI (Supplemental Enhancement Information) message in HEVC. The metadata indicates, for example, the position, size, or color of the main object.

[0417] The metadata may be stored in units of multiple pictures (such as a stream, a sequence, or a random access unit). In this way, the decoder 200 is able to obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the picture unit information, is able to identify the picture in which the object (person) appears and determine the position of the object in the picture.

[0418] (Encoder)

[0419] An encoder according to an embodiment will be described. Figure 7 FIG. is a block diagram showing a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes video in units of blocks.

[0420] As Figure 7 shown, the encoder 100 is a device that encodes an image in units of blocks, and includes a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra predictor 124, an inter predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra predictor 124 and the inter predictor 126 are part of the prediction controller.

[0421] The encoder 100 is implemented as, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor acts as a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, an intra predictor 124, an inter predictor 126, and a prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra predictor 124, the inter predictor 126, and the prediction controller 128.

[0422] (Installation example of the encoder)

[0423] Figure 8 is a functional block diagram showing an installation example of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 a plurality of components of the encoder 100 shown are installed on Figure 8 the processor a1 and the memory a2 shown.

[0424] The processor a1 is a circuit that performs information processing and is coupled to the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit for encoding images. The processor a1 may be a processor such as a CPU. In addition, the processor a1 may be an aggregate of multiple electronic circuits. Additionally, for example, the processor a1 may assume Figure 7 the roles of two or more components among the multiple components of the encoder 100 shown, etc.

[0425] The memory a2 is a dedicated or general-purpose memory for storing information used by the processor a1 to encode images. The memory a2 may be an electronic circuit and may be connected to the processor a1. In addition, the memory a2 may be included in the processor a1. In addition, the memory a2 may be an aggregate of multiple electronic circuits. Additionally, the memory a2 may be a magnetic disk, an optical disk, etc., or may be represented as a storage device, a recording medium, etc. In addition, the memory a2 may be a non-volatile memory or a volatile memory.

[0426] For example, the memory a2 may store an image to be encoded or a bitstream corresponding to the encoded image. In addition, the memory a2 may store a program for causing the processor a1 to encode images.

[0427] In addition, for example, the memory a2 may act as Figure 7The roles of two or more elements for storing information among the multiple elements such as the encoder 100 shown. For example, the memory a2 can act as Figure 7 the roles of the block memory 118 and the frame memory 122 shown. More specifically, the memory a2 can store reconstructed blocks, reconstructed pictures, etc.

[0428] It should be noted that in the encoder 100, all of the multiple elements etc. shown may not be implemented, and all of the processes described here may not be executed. Figure 7 A part of the elements etc. shown may be included in another device, or a part of the processes described here may be executed by another device. Figure 7 A part of the elements etc. shown may be included in another device, or a part of the processes described here may be executed by another device.

[0429] Hereinafter, the overall flow of the process executed by the encoder 100 will be described, and then each element included in the encoder 100 will be described.

[0430] (Overall flow of the encoding process)

[0431] Fig. 9 is a flowchart showing an example of the overall encoding process executed by the encoder 100, and will be described for convenience with reference to Figure 7 for description.

[0432] First, the splitter 102 of the encoder 100 splits each picture included in the input image into a plurality of blocks having a fixed size (e.g., 128×128 pixels) (step Sa_1). The splitter 102 then selects a splitting mode for the fixed-size blocks (also referred to as the block shape) (step Sa_2). In other words, the splitter 102 further splits the fixed-size blocks into a plurality of blocks forming the selected splitting mode. The encoder 100 executes steps Sa_3 to Sa_9 for each of the plurality of blocks, for that block (i.e., the current block to be encoded).

[0433] The prediction controller 128 and the prediction executor (which includes the intra-frame predictor 124 and the inter-frame predictor 126) generate a predicted image of the current block (step Sa-3). The predicted image may also be referred to as a prediction signal, a prediction block, or a prediction sample.

[0434] Next, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Sa_4). The prediction residual may also be referred to as a prediction error.

[0435] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients may sometimes be referred to as a coefficient block.

[0436] Next, the entropy encoder 110 encodes (specifically, entropy-codes) the plurality of quantized coefficients and prediction parameters related to generation of a predicted image to generate a stream (step Sa_6). This stream may sometimes be referred to as an encoded bitstream or a compressed bitstream.

[0437] Next, the inverse quantizer 112 performs inverse quantization of the plurality of quantized coefficients, and the inverse transformer 114 performs inverse transformation of the result to recover the prediction residual (step Sa_7).

[0438] Next, the adder 116 adds the predicted image and the recovered prediction residual to reconstruct the current block (step Sa_8). Thus, a reconstructed image is generated. The reconstructed image may also be referred to as a reconstructed block or a decoded image block.

[0439] When the reconstructed image is generated, the loop filter 120 performs filtering of the reconstructed image as needed (step Sa_9).

[0440] The encoder 100 then determines whether the encoding of the entire picture has been completed (step Sa_10). When it is determined that the encoding has not been completed (No in step Sa_10), the processing starting from step Sa_2 is repeated for the next block of the image.

[0441] Although in the above example the encoder 100 selects a splitting mode for a fixed-size block and encodes each block according to the splitting mode, it should be noted that each block may be encoded according to a corresponding splitting mode among a plurality of splitting modes. In this case, the encoder 100 may evaluate the cost of each of the plurality of splitting modes, and may, for example, select the stream that can be obtained by encoding according to the splitting mode that yields the minimum cost as the stream to be output.

[0442] As shown, the processes in steps Sa_1 to Sa_10 are sequentially executed by the encoder 100. Alternatively, two or more processes may be executed in parallel, the processes may be reordered, and so on.

[0443] The encoding process employed by the encoder 100 is a hybrid encoding using predictive encoding and transform encoding. Further, the predictive encoding is executed by an encoding loop configured with a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, a block memory 118, a frame memory 122, an intra predictor 124, an inter predictor 126, and a prediction controller 128. In other words, the prediction executor configured with the intra predictor 124 and the inter predictor 126 is part of the encoding loop.

[0444] (Splitter)

[0445] The splitter 102 splits each picture included in the original image into a plurality of blocks and outputs each block to the subtractor 104. For example, the splitter 102 first splits the picture into blocks of a fixed size (e.g., 128×128 pixels). Other fixed block sizes may be employed. The blocks of the fixed size are also referred to as coding tree units (CTUs). The splitter 102 then splits each fixed-size block into variable-size (e.g., 64×64 pixels or smaller) blocks based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a splitting mode. The variable-size blocks may also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in the picture may be processed in units of CUs, PUs, or TUs.

[0446] Fig.10 is a conceptual diagram for showing an example of block splitting according to an embodiment. In Fig.10 it, the solid lines represent the block boundaries of the blocks split by quadtree block splitting, and the dashed lines represent the block boundaries of the blocks split by binary tree block splitting.

[0447] Here, the block 10 is a square block (128×128 block) having 128×128 pixels. This 128×128 block 10 is first split into four square 64×64 pixel blocks (quadtree block splitting).

[0448] The upper-left 64×64 pixel block is further vertically split into two rectangular 32×64 pixel blocks, and the left 32×64 pixel block is further vertically split into two rectangular 16×64 pixel blocks (binary tree block splitting). As a result, the upper-left 64×64 pixel block is split into two 16×64 pixel blocks 11 and 12 and a 32×64 pixel block 13.

[0449] The upper-right 64×64 pixel block is horizontally split into two rectangular 64×32 pixel blocks 14 and 15 (binary tree block splitting).

[0450] The 64×64 pixel block in the lower left corner is first divided into four square 32×32 pixel blocks (quad tree block division). The upper left and lower right blocks among the four square 32×32 pixel blocks are further divided. The square 32×32 pixel block in the upper left corner is vertically divided into two rectangular 16×32 pixel blocks, and the right 16×32 pixel block is further horizontally divided into two 16×16 pixel blocks (binary tree block division). The 32×32 pixel block in the lower right corner is horizontally divided into two 32×16 pixel blocks (binary tree block division). The square 32×32 pixel block in the upper right corner is horizontally divided into two rectangular 32×16 pixel blocks (binary tree block division). As a result, the square 64×64 pixel block in the lower left corner is divided into rectangular 16×32 pixel blocks 16, two square 16×16 pixel blocks 17 and 18, two square 32×32 pixel blocks 19 and 20, and two rectangular 32×16 pixel blocks 21 and 22.

[0451] The 64×64 pixel block 23 in the lower right corner is not divided.

[0452] As described above, in Fig.10 , based on recursive quad tree and binary tree block division, block 10 is divided into 13 variable-sized blocks 11 to 23. This type of division is also called quad tree plus binary tree (QTBT) division.

[0453] It should be noted that in Fig.10 , a block is divided into four or two blocks (quad tree or binary tree block division), but the division is not limited to these examples. For example, a block can be divided into three blocks (ternary block division). The division including such ternary block division is also called multi-type tree (MBT) division.

[0454] Fig.11 is a block diagram showing an example of the functional configuration of the divider 102 according to an embodiment. As Fig.11 shown, the divider 102 may include a block division determiner 102a. As an example, the block division determiner 102a may perform the following process.

[0455] For example, the block division determiner 102a may obtain or retrieve block information from the block memory 118 and / or the frame memory 122, and determine a division pattern (e.g., the division pattern described above) based on the block information. The divider 102 divides the original image according to the division pattern and outputs at least one block obtained by the division to the subtractor 104.

[0456] In addition, for example, the block segmentation determiner 102a outputs one or more parameters indicating the determined segmentation pattern (e.g., the above-mentioned segmentation pattern) to the transformer 106, the inverse transformer 114, the intra predictor 124, the inter predictor 126, and the entropy encoder 110. The transformer 106 can transform the prediction residual based on one or more parameters. The intra predictor 124 and the inter predictor 126 can generate a prediction image based on one or more parameters. In addition, the entropy encoder 110 can perform entropy encoding on one or more parameters.

[0457] As an example, the parameters related to the segmentation pattern can be written in the stream as follows.

[0458] Fig.12 FIG. is a conceptual diagram for showing an example of the segmentation pattern. Examples of the segmentation pattern include: segmentation into four regions (QT), where one block is divided into two regions both horizontally and vertically; segmentation into three regions (HT or VT), where one block is divided in the same direction at a ratio of 1:2:1; segmentation into two regions (HB or VB), where one block is divided in the same direction at a ratio of 1:1; and no segmentation (NS).

[0459] It should be noted that the segmentation pattern does not have a block segmentation direction in the case of segmentation into four regions and no segmentation, and the segmentation pattern has segmentation direction information in the case of segmentation into two regions or three regions.

[0460] Fig.13A FIG. is a conceptual diagram for showing an example of the syntax tree of the segmentation pattern.

[0461] Fig. 13B FIG. is a conceptual diagram for showing another example of the syntax tree of the segmentation pattern.

[0462] Fig.13A and Fig. 13B FIG. are conceptual diagrams for showing examples of the syntax tree of the segmentation pattern. In Fig.13AIn the example, first, there is information indicating whether to perform segmentation (S: segmentation flag), and next, there is information indicating whether to perform segmentation into four regions (QT: QT flag). Next, there is information indicating which of the three-region and two-region segmentation to perform (TT: TT flag or BT: BT flag), and then there is information indicating the division direction (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of at least one block obtained by segmenting according to such a segmentation pattern can be further repeatedly segmented in a similar process. In other words, as an example, whether to perform segmentation, whether to perform segmentation into four regions, which of the horizontal and vertical directions is the direction of the segmentation method to be performed, which of the three-region and two-region segmentation to perform can be determined recursively, and can be encoded in the stream according to the encoding order disclosed by the syntax tree as shown in Fig.13A shown.

[0463] In addition, although the information items indicating S, QT, TT, and Ver are arranged in the listed order in the syntax tree as shown in Fig.13A shown, the information items indicating S, QT, Ver, and BT can also be arranged in the listed order. In other words, in the example of Fig. 13B , first, there is information indicating whether to perform segmentation (S: segmentation flag), and next, there is information indicating whether to perform segmentation into four regions (QT: QT flag). Next, there is information indicating the segmentation direction (Ver: vertical flag, or Hor: horizontal flag), and next, there is information indicating which of the two-region and three-region segmentation to perform (BT: BT flag or TT: TT flag).

[0464] It should be noted that the above segmentation pattern is an example, and a segmentation pattern other than the described segmentation pattern can be used, or a part of the described segmentation pattern can be used.

[0465] (Subtractor)

[0466] The subtractor 104 subtracts the predicted image (predicted samples input from the prediction controller 128 indicated below) from the original image in units of blocks, and the original image is input from the segmenter 102 and segmented by the segmenter 102. In other words, the subtractor 104 calculates the prediction residual of the current block (also referred to as the error). The subtractor 104 then outputs the calculated prediction residual to the transformer 106.

[0467] The original image can be an image whose signal representing each picture included in the video (for example, a luminance signal and two chrominance signals) has been input to the encoder 100. The signal representing the image can also be referred to as a sample.

[0468] (Transformer)

[0469] Transformer 106 transforms the prediction residual in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to quantizer 108. More specifically, Transformer 106 applies, for example, a defined discrete cosine transform (DCT) or discrete sine transform (DST) to the prediction residual in the spatial domain. The defined DCT or DST can be predefined.

[0470] It should be noted that Transformer 106 can adaptively select a transform type from multiple transform types and transform the prediction residual into transform coefficients by using a transform basis function corresponding to the selected transform type. This transform is also referred to as explicit multi-core transform (EMT) or adaptive multi-core transform (AMT). The transform basis function can also be referred to as a basis.

[0471] The transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transform types can be represented as DCT2, DCT5, DCT8, DST1, and DST7. Fig.14 is a chart of example transform basis functions indicating example transform types. In Fig.14 where N represents the number of input pixels. For example, selecting a transform type from multiple transform types can depend on the prediction type (one of intra prediction and inter prediction) and can depend on the intra prediction mode.

[0472] Information indicating whether to apply such EMT or AMT (e.g., referred to as an EMT flag or AMT flag) and information indicating the selected transform type are typically signaled at the CU level. It should be noted that signaling of such information does not necessarily need to be performed at the CU level and can also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0473] In addition, Transformer 106 can perform a re - transform on the transform coefficients (which are the transform results). This re - transform is also referred to as adaptive secondary transform (AST) or non - separable secondary transform (NSST). For example, Transformer 106 performs the re - transform in units of sub - blocks (e.g., 4×4 pixel sub - blocks) included in a transform coefficient block corresponding to the intra prediction residual. Information indicating whether to apply NSST and information related to the transform matrix used in NSST are typically signaled at the CU level. It should be noted that signaling of such information does not necessarily need to be performed at the CU level and can also be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0474] The transformer 106 can employ separable transforms and non-separable transforms. A separable transform is a method in which the transform is performed multiple times by separately performing the transform for each of a plurality of directions according to the dimension of the input. A non-separable transform is a method of performing a collective transform, in which two or more dimensions in a multi-dimensional input are collectively regarded as a single dimension.

[0475] In one example of a non-separable transform, when the input is a 4×4 pixel block, the 4×4 pixel block is considered as a single array containing 16 elements, and the transform applies a 16×16 transform matrix to this array.

[0476] In another example of a non-separable transform, an input block of 4×4 pixels is regarded as a single array containing 16 elements, and then a transform that performs a given rotation on this array multiple times (a given transform of a hypercube) can be performed.

[0477] In the transform in the transformer 106, the type of transform of the transform basis function to be transformed into the frequency domain can be switched according to the region in the CU. Examples include a spatially varying transform (SVT).

[0478] Fig.15 is a conceptual diagram for showing an example of SVT.

[0479] In SVT, as Fig.15 shown, the CU is horizontally or vertically divided into two equal regions, and only one of the regions is transformed into the frequency domain. The transform basis type can be set for each region. For example, DST7 and DST8 are used. For example, in the two regions obtained by vertically dividing the CU into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, in the two regions, DST7 can be used for the region at position 1. Similarly, in the two regions obtained by horizontally dividing the CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Although in Fig.15 the example shown, one of the two regions in the CU is transformed and the other region is not transformed, each of the two regions can be transformed. In addition, the division method can include not only division into two regions, but also division into four regions. Furthermore, the division method can be more flexible. For example, the information indicating the division method can be encoded and can be signaled in the same way as the CU division. It should be noted that SVT can also be referred to as a sub-block transform (SBT).

[0480] The AMT and EMT described above can be referred to as MTS (Multiple Transform Selection). When applying MTS, transform types such as DST7, DCT8, etc. can be selected, and the information indicating the selected transform type can be encoded as index information for each CU. There is another process called IMTS (Implicit MTS) as a process for selecting the transform type to be used for the orthogonal transform to be performed without encoded index information. When applying IMTS, for example, when a CU has a rectangular shape, the orthogonal transform of the rectangular shape can be performed using DST7 (for the short side) and DST2 (for the long side). Additionally, for example, when a CU has a square shape, the orthogonal transform of the rectangular shape can be performed by using DCT2 when MTS is valid in the sequence and using DST7 when MTS is invalid in the sequence. DCT2 and DST7 are only examples. Other transform types can be used, and the combination of the used transform types can also be changed to a different combination of transform types. IMTS can be used only for intra-prediction blocks, or can be used for both intra-prediction blocks and inter-prediction blocks.

[0481] The three processes of MTS, SBT, and IMTS have been described above as selection processes for selectively switching the transform type used for the orthogonal transform. However, all three selection processes can be adopted, or only some of the selection processes can be selectively adopted. For example, it can be identified whether to adopt one or more selection processes based on flag information in headers such as SPS, etc. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching the transform type can be a selection process different from the above three selection processes, or each of the three selection processes can be replaced by another process. Generally, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing the orthogonal transform of the entire CU and determining the transform type based on a determined rule without encoding the information indicating the transform type. Function [3] is a function for performing the orthogonal transform of a partial region of the CU and encoding the information indicating the transform type used in the transform. Function [4] is a function for performing the orthogonal transform of a partial region of the CU and determining the transform type based on a determined rule without encoding the information indicating the transform type used in the transform. The determined rule can be predetermined.

[0482] It should be noted that it can be determined for each processing unit whether to apply MTS, IMTS, and / or SBT. For example, it can be determined for each sequence, picture, tile, slice, CTU, or CU whether to apply MTS, IMTS, and / or SBT.

[0483] It should be noted that the tool of the selective switching transformation type in the present invention can be described as a method, a selection process, or a process for selecting a basis used in the transformation process for selectively selecting a basis. Additionally, the tool for selectively switching the transformation type can be described as a mode for adaptively selecting the transformation type.

[0484] Fig.16 is a flowchart showing an example of the process performed by the transformer 106, and for convenience, reference will be made to Figure 7 for description.

[0485] For example, the transformer 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined to perform an orthogonal transformation (yes in step St_1), the transformer 106 selects a transformation type for the orthogonal transformation from among a plurality of transformation types (step St_2). Next, the transformer 106 performs an orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transformation type to the entropy encoder 110 so as to allow the entropy encoder 110 to encode this information (step St_4). On the other hand, when it is determined not to perform an orthogonal transformation (no in step St_1), the transformer 106 outputs information indicating that no orthogonal transformation is performed so as to allow the entropy encoder 110 to encode this information (step St_5). It should be noted that whether to perform an orthogonal transformation in step St_1 can be determined based on, for example, the size of the transformation block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transformation can also be performed using a defined transformation type without encoding information indicating the transformation type used in the orthogonal transformation. The defined transformation type can be predefined.

[0486] Fig.17 is a flowchart showing an example of the process performed by the transformer 106, and for convenience, reference will be made to Figure 7 for description. It should be noted that Fig.17 the example shown in Fig.16 is an example of an orthogonal transformation in the case where the transformation type used in the orthogonal transformation is selectively switched (as in the case of the example shown in

[0487] As an example, the first transformation type group may include DCT2, DST7, and DCT8. As another example, the second transformation type group may include DCT2. The transformation types included in the first transformation type group and the transformation types included in the second transformation type group may partially overlap with each other, or may be completely different from each other.

[0488] The transformer 106 determines whether the transform size is less than or equal to a determined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the determined value (yes in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform type included in the first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one transform type included in the first transform type group to the entropy encoder 110 so as to allow the entropy encoder 110 to encode this information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to the predetermined value (no in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The determined value can be a threshold value and can be a predetermined value.

[0489] In step Su_3, the information indicating the transform type used in the orthogonal transform can be information indicating a combination of the transform type to be vertically applied to the current block and the transform type to be horizontally applied to the current block. The first type group can include only one transform type, and the information indicating the transform type used for the orthogonal transform may not be encoded. The second transform type group can include multiple transform types, and the information indicating the transform type used for the orthogonal transform among one or more transform types included in the second transform type group may be encoded.

[0490] Alternatively, the transform type can be indicated based on the transform size without encoding the information indicating the transform type. It should be noted that such determination is not limited to the determination of whether the transform size is less than or equal to the determined value, and other processes for determining the transform type used in the orthogonal transform based on the transform size are also possible.

[0491] (Quantizer)

[0492] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in a determined scan order and quantizes the scanned transform coefficients based on the quantization parameter (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients (hereinafter also referred to as quantized coefficients) of the current block to the entropy encoder 110 and the inverse quantizer 112. The determined scan order can be predetermined.

[0493] The determined scan order is the order for quantizing / inverse quantizing the transform coefficients. For example, the determined scan order can be defined as an ascending order of frequencies (from low frequency to high frequency) or a descending order of frequencies (from high frequency to low frequency).

[0494] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, when the value of the quantization parameter increases, the quantization step also increases. In other words, when the value of the quantization parameter increases, the error of the quantized coefficient (quantization error) increases.

[0495] In addition, quantization matrices can be used for quantization. For example, multiple quantization matrices can be used corresponding to the frequency transform size (e.g., 4×4, 8×8), prediction mode (e.g., intra prediction, inter prediction), and pixel component (e.g., luminance, chrominance pixel components). It should be noted that quantization means digitizing the sampled values at determined intervals corresponding to determined levels. In the present technical field, quantization can be referred to using other expressions, such as rounding and scaling, and rounding and scaling can be employed. The determined intervals and determined levels can be pre-determined.

[0496] The method of using a quantization matrix can include: a method of using a quantization matrix directly set on the encoder 100 side, and a method of using a quantization matrix set as the default (default matrix). On the encoder 100 side, a quantization matrix suitable for the image characteristics can be set by directly setting the quantization matrix. However, this situation may have the disadvantage of increasing the coding amount for encoding the quantization matrix. It should be noted that a quantization matrix for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.

[0497] There is a method of quantizing high-frequency coefficients and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be regarded as equivalent to a method of using a quantization matrix (flat matrix) whose coefficients have the same value.

[0498] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level. The quantization matrix can be specified using, for example, the sequence parameter set (SPS) or the picture parameter set (PPS). The SPS includes parameters for the sequence, and the PPS includes parameters for the picture. Each of the SPS and PPS can be abbreviated as a parameter set.

[0499] When using a quantization matrix, the quantizer 108 scales the quantization width, which can be calculated based on, for example, the quantization parameter, for each transform coefficient using the value of the quantization matrix. The quantization process performed without using a quantization matrix can be a process of quantizing the transform coefficients according to the quantization width calculated based on, for example, the quantization parameter. It should be noted that in the quantization process performed without using any quantization matrix, the quantization width can be multiplied by a determined value common to all transform coefficients in the block. The determined value can be pre-determined.

[0500] Fig.18It is a block diagram showing an example of the functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a differential quantization parameter generator 108a, a predictive quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization executor 108e.

[0501] Fig.19 It is a flowchart showing an example of the quantization process executed by the quantizer 108, and for convenience, reference will be made to Figure 7 and 18 for description.

[0502] As an example, the quantizer 108 can perform quantization on each CU based on the Fig.19 flowchart shown. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (Yes in step Sv_1), the quantization parameter generator 108c generates quantization parameters for the current block (step Sv_2), and stores the quantization parameters in the quantization parameter storage device 108d (step Sv_3).

[0503] Next, the quantization executor 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). The predictive quantization parameter generator 108b then obtains the quantization parameters of a processing unit different from the current block from the quantization parameter storage device 108d (step Sv_5). The predictive quantization parameter generator 108b generates predictive quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generator 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generator 108c and the predictive quantization parameters of the current block generated by the predictive quantization parameter generator 108b (step Sv_7). The differential quantization parameters can be generated by calculating the difference. The differential quantization parameter generator 108a outputs the differential quantization parameters to the entropy encoder 110 so that the entropy encoder 110 can encode the differential quantization parameters (step Sv_8).

[0504] It should be noted that the differential quantization parameters can be encoded at, for example, the sequence level, picture level, slice level, tile level, or CTU level. In addition, the initial values of the quantization parameters can be encoded at the sequence level, picture level, slice level, tile level, or CTU level. At initialization, the quantization parameters can be generated using the initial values of the quantization parameters and the differential quantization parameters.

[0505] It should be noted that the quantizer 108 can include multiple quantizers, and dependent quantization can be applied, where the transform coefficients are quantized using a quantization method selected from multiple quantization methods.

[0506] (Entropy Encoder)

[0507] Fig. 20 is a block diagram showing an example of the functional configuration of the entropy encoder 110 according to an embodiment, and will be described for convenience with reference to Figure 7 The entropy encoder 110 generates a bitstream by entropy encoding the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy coding. More specifically, the entropy encoder 110 shown in the figure includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, in which a multi-level signal such as a quantized coefficient and a prediction parameter is transformed into a binary signal. Examples of binarization methods include truncated Rice binarization, exponential Golomb code, and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntax element or the surrounding state (i.e., the occurrence probability of the binary signal). Examples of methods for deriving the context value include bypassing, referring to syntax elements, referring to the upper and left adjacent blocks, referring to hierarchical information, etc. The binary arithmetic encoder 110c performs arithmetic coding on the binary signal using the derived context.

[0508] Fig.21 is a conceptual diagram for explaining an example process of CABAC in the entropy encoder 110. First, initialization is performed in the entropy encoder 110 with CABAC. In the initialization, initialization in the binary arithmetic encoder 110c and setting of the initial context value are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic coding of multiple quantized coefficients in the CTU. Each time arithmetic coding is performed, the context controller 110b can update the context value. The context controller 110b can then save the context value as post-processing. For example, the saved context value can be used to initialize the context value of the next CTU.

[0509] (Inverse Quantizer)

[0510] The inverse quantizer 112 inverse-quantizes the quantized coefficients input from the quantizer 108. More specifically, the inverse quantizer 112 inverse-quantizes the quantized coefficients of the current block in a determined scan order. The inverse quantizer 112 then outputs the inverse-quantized transform coefficients of the current block to the inverse transform unit 114. The determined scan order can be predetermined.

[0511] (Inverse Transform Unit)

[0512] The inverse transformer 114 restores the prediction residual by performing an inverse transform on the transform coefficients input from the inverse quantizer 112. More specifically, the inverse transformer 114 restores the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the restored prediction residual to the adder 116.

[0513] It should be noted that since information is usually lost in quantization, the restored prediction residual does not match the prediction residual calculated by the subtractor 104. In other words, the restored prediction residual usually includes quantization errors.

[0514] (Adder)

[0515] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transformer 114 and the predicted image input from the prediction controller 128. Subsequently, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block may also be referred to as a local decoded block.

[0516] (Block Memory)

[0517] The block memory 118 is a storage device for storing, for example, blocks in the current picture for intra prediction. More specifically, the block memory 118 stores the reconstructed image output from the adder 116.

[0518] (Frame Memory)

[0519] The frame memory 122 is a storage device for storing, for example, reference pictures used in inter prediction, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.

[0520] (Loop Filter)

[0521] The loop filter 120 applies a loop filter to the reconstructed image output from the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter used in the coding loop (intra-loop filter). Examples of the loop filter include, for example, an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, etc.

[0522] Fig. 22 is a block diagram showing an example of the functional configuration of the loop filter 120 according to an embodiment. For example, as Fig. 22As shown, the loop filter 120 includes a deblocking filter executor 120a, an SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs a deblocking filter process on the reconstructed image. The SAO executor 120b performs an SAO process on the reconstructed image after the deblocking filter process. The ALF executor 120c performs an ALF process on the reconstructed image after the SAO process. The ALF and the deblocking filter will be described in detail later. The SAO process is a process for improving image quality by reducing ringing (a phenomenon in which pixel values are distorted like waves around an edge) and correcting pixel value deviations. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that in some embodiments, the loop filter 120 may not include Fig. 22 all the constituent elements disclosed in Fig. 22 , and may include some constituent elements and may include additional elements. In addition, the loop filter 120 may be configured to perform the above processes in a processing order different from that disclosed in

[0523] (Loop filter > Adaptive loop filter)

[0524] In the ALF, a least squares error filter for removing compression artifacts is applied. For example, for each 2×2 pixel sub-block in the current block, one filter selected from a plurality of filters is applied based on the local gradient direction and activity.

[0525] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of a plurality of classes (e.g., fifteen or twenty-five classes). The classification of the sub-block can be based on, for example, gradient directionality and activity. In an example, a class index C (e.g., C = 5D + A) is calculated or determined based on the gradient directionality D (e.g., 0 to 2 or 0 to 4) and the gradient activity A (e.g., 0 to 4). Then, based on the classification index C, each sub-block is classified into one of the plurality of classes.

[0526] For example, the gradient directionality D is calculated by comparing the gradients in a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). In addition, for example, the gradient activity A is calculated by adding the gradients in a plurality of directions and quantifying the added result.

[0527] Based on such classification results, the filter to be used for each sub-block can be determined from a plurality of filters.

[0528] The filter shape to be used in the ALF is, for example, a circularly symmetric filter shape. FIG. 23A to FIG. 23C is a conceptual diagram for showing an example of the filter shape used in the ALF. Fig.23A A 5×5 diamond filter is illustrated, Fig. 23B a 7×7 diamond filter is illustrated, and Fig.23C a 9×9 diamond filter is illustrated. Information indicating the filter shape is typically signaled at the picture level. It should be noted that the signaling of such information indicating the filter shape does not necessarily need to be performed at the picture level and can be performed at another level (e.g., at the sequence level, slice level, tile level, CTU level, or CU level).

[0529] For example, the turning on or off of the ALF can be determined at the picture level or CU level. For example, a decision on whether to apply the ALF to the luminance can be made at the CU level, and a decision on whether to apply the ALF to the chrominance can be made at the picture level. Information indicating the turning on or off of the ALF is typically signaled at the picture level or CU level. It should be noted that the signaling of the information indicating the turning on or off of the ALF does not necessarily need to be performed at the picture level or CU level and can be performed at another level (e.g., at the sequence level, slice level, tile level, or CTU level).

[0530] In addition, as described above, one filter is selected from multiple filters, and the ALF process for the sub-block is performed. The set of coefficients for each of the multiple filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that the signaling of the set of coefficients does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0531] (Loop Filter > Cross-component Adaptive Loop Filter)

[0532] Fig.23D is a conceptual diagram for showing an example process of the cross-component ALF (CC-ALF). Fig.23E is a conceptual diagram for showing an example of the filter shape used in the CC-ALF, such as Fig.23D the CC-ALF of Fig.23D and Fig.23E the example CC-ALF of operates by applying a linear diamond filter to the luminance channel of each chrominance component. For example, the filter coefficients can be transmitted in the APS, scaled by a factor of 2^10, and rounded for fixed-point representation. For example, in Fig.23D the Y samples (the first component) are used for the CCALF of Cb and the CCALF of Cr (components different from the first component).

[0533] The application of the filter can be controlled on variable block sizes and signaled by flags encoded for each sample block's context. The block size together with the CC-ALF enable flag can be received at the slice level for each chrominance component. CC-ALF can support various block sizes, e.g., 16×16 pixels, 32×32 pixels, 64×64 pixels, 128×128 pixels (in chrominance samples).

[0534] (Loop Filter > Joint Chrominance Cross-Component Adaptive Loop Filter)

[0535] An example of joint chrominance - CCALF is in Fig.23F and Figure 23G is shown. Fig.23F is a conceptual diagram for showing an example process of joint chrominance CCALF. Figure 23G is a table showing example weight index candidates. As shown, one CCALF filter is used to generate a CCALF filtered output as a chrominance refinement signal for one color component while applying a weighted version of the same chrominance refinement signal to another color component. In this way, the complexity of the existing CCALF is reduced by about half. The weight value can be encoded as a sign flag and a weight index. The weight index (denoted as weight_index) can be encoded into 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight, which is non-zero. For example, the magnitude of JcCcWeight can be determined as follows:

[0536] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index >> 2;

[0537] Otherwise, JcCcWeight is equal to 4 / (weight_index – 4).

[0538] The block-level on / off control for ALF filtering of Cb and Cr can be separate. This is the same as in CCALF, and two separate groups of block-level on / off control flags can be encoded. Different from CCALF, the Cb, Cr on / off control block sizes here are the same, so only one block size variable can be encoded.

[0539] (Loop Filter > Deblocking Filter)

[0540] During the deblocking filtering process, the loop filter 120 performs a filtering process on the block boundaries in the reconstructed image to reduce the distortion occurring at the block boundaries.

[0541] Fig.24 is showing the loop filter 120 acting as a deblocking filter (see Figure 7 and Fig. 22Block diagram of an example of the specific configuration of the deblocking filter actuator 120a.

[0542] The deblocking filter actuator 120a includes: a boundary determiner 1201; a filter determiner 1203; a filtering actuator 1205; a process determiner 1208; a filter characteristic determiner 1207; and switches 1202, 1204, and 1206.

[0543] The boundary determiner 1201 determines whether the pixel to be deblocked filtered (i.e., the current pixel) exists around the block boundary. The boundary determiner 1201 then outputs the determination result to the switch 1202 and the process determiner 1208.

[0544] In the case where the boundary determiner 1201 determines that the current pixel exists around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1204. In the opposite case (where the boundary determiner 1201 determines that the current pixel does not exist around the block boundary), the switch 1202 outputs the unfiltered image to the switch 1206. Note that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.

[0545] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.

[0546] In the case where the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the filtering actuator 1205. In the opposite case (where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel), the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the switch 1206.

[0547] When the unfiltered image is obtained through the switches 1202 and 1204, the filtering actuator 1205 performs deblocking filtering on the current pixel with the filtering characteristics determined by the filter characteristic determiner 1207. The filtering actuator 1205 then outputs the filtered pixel to the switch 1206.

[0548] Under the control of the process determiner 1208, the switch 1206 selectively outputs one of the pixels that have not been deblocked filtered and the pixels that have been deblocked filtered by the filtering actuator 1205.

[0549] The processing determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary and when the filter determiner 1203 has determined that deblocking filtering of the current pixel is to be performed, the processing determiner 1208 causes the switch 1207 to output the pixel for which deblocking filtering has been performed. In addition, except for the above cases, the processing determiner 1208 causes the switch 1206 to output the pixel for which deblocking filtering has not been performed. By repeating the output of the pixels in this way, the filtered image is output from the switch 1206. It should be noted that Fig.24 The configuration shown in

[0550] Fig.25 is a conceptual diagram for showing an example of a deblocking filter having symmetric filtering characteristics with respect to the block boundary.

[0551] During the deblocking filtering process, a pixel value and a quantization parameter can be used to select one of two deblocking filters (i.e., a strong filter and a weak filter) having different characteristics. In the case of the strong filter, when pixels p0 to p2 and pixels q0 to q2 exist across the block boundary, as Fig.25 shown, by performing calculations according to the following expressions, for example, the pixel values of the corresponding pixels q0 to q2 are changed to pixel values q'0 to q'2.

[0552] q'0 = (p1 + 2×p0 + 2×q0 + 2×q1 + q2 + 4) / 8

[0553] q'1 = (p0 + q0 + q1 + q2 + 2) / 4

[0554] q'2 = (p0 + q0 + q1 + 3×q2 + 2×q3 + 4) / 8

[0555] It should be noted that in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the corresponding pixels p0 to p2 and pixels q0 to q2. In addition, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of the pixel q2 with respect to the block boundary. In addition, on the right side of each expression, the coefficient multiplied by the corresponding pixel value of the pixel to be used for deblocking filtering is the filter coefficient.

[0556] In addition, in deblocking filtering, clipping can be performed so that the change in the calculated pixel value does not exceed a threshold. For example, during the clipping process, the pixel value calculated according to the above expressions can be clipped to a value obtained according to "calculated pixel value ± 2×threshold" (using a threshold determined based on the quantization parameter). In this way, over-smoothing can be prevented.

[0557] Fig.26 It is a conceptual diagram for showing a block boundary on which a deblocking filtering process is performed. Fig. 27 It is a conceptual diagram for showing an example of a boundary strength (Bs) value.

[0558] The block boundary on which the deblocking filtering process is performed is, for example, a boundary between CUs, Pus, or TUs having 8×8 pixel blocks, as Fig.26 shown. The deblocking filtering process can be performed, for example, in units of four rows or four columns. First, as Fig. 27 shown for block P and block Q ( Fig.26 shown), a boundary strength (Bs) value is determined.

[0559] According to the Fig. 27 Bs value in, it can be determined whether to perform a deblocking filtering process on a block boundary belonging to the same image with different strengths. When the Bs value is 2, a deblocking filtering process for a chrominance signal is performed. When the Bs value is 1 or greater and a determined condition is satisfied, a deblocking filtering process for a luminance signal is performed. The determined condition can be predetermined. Note that the conditions for determining the Bs value are not limited to those Fig. 27 shown in, and the Bs value can be determined based on another parameter.

[0560] (Predictor (intra predictor, inter predictor, prediction controller))

[0561] Fig.28 It is a flowchart showing an example of a process performed by the predictor of the encoder 100. It should be noted that the predictor includes all or part of the following constituent elements: an intra predictor 124; an inter predictor 126; and a prediction controller 128. The prediction executor includes, for example, the intra predictor 124 and the inter predictor 126.

[0562] The predictor generates a predicted image of the current block (step Sb_1). This predicted image can also be referred to as a prediction signal or a predicted block. It should be noted that the prediction signal is, for example, an intra predicted image (image prediction signal) or an inter predicted image (inter prediction signal). The predictor uses a reconstructed image that has been obtained through the generation of a prediction image, the generation of a prediction residual, the generation of quantized coefficients, the recovery of the prediction residual, and the addition to the prediction image by another block, to generate a predicted image of the current block.

[0563] The reconstructed image can be, for example, an image in a reference picture, or an image of an encoded block (i.e., the above-mentioned other block) in the current picture, and the current picture is a picture including the current block. The encoded block in the current picture is, for example, an adjacent block of the current block.

[0564] Fig.29 It is a flowchart showing another example of a process performed by the predictor of the encoder 100.

[0565] The predictor generates a prediction image using a first method (step Sc_1a), generates a prediction image using a second method (step Sc_1b), and generates a prediction image using a third method (step Sc_1c). The first method, the second method, and the third method may be mutually different methods for generating a prediction image. Each of the first to third methods may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.

[0566] Next, the prediction processor evaluates the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates a cost C for the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the prediction images by comparing the costs C of the prediction images. It should be noted that the cost C can be calculated according to the expression of the R-D optimization model, for example, C = D + λ × R. In this expression, D represents the compression artifacts of the prediction image and is expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the prediction image. In addition, R represents the bit rate of the stream. In addition, λ represents, for example, a multiplier according to the Lagrange method multiplier.

[0567] Then, the predictor selects one of the prediction images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects the method or mode for obtaining the final prediction image. For example, the predictor selects the prediction image with the minimum cost C based on the cost C calculated for the prediction image. Alternatively, the evaluation in step Sc_2 and the selection of the prediction image in step Sc_3 may be performed based on the parameters used in the encoding process. The encoder 100 may transform the information for identifying the selected prediction image, method, or mode into a stream. This information may be, for example, a flag or the like. In this way, the decoder 200 can generate a prediction image based on this information according to the method or mode selected by the encoder 100. It should be noted that in Fig.29 the example shown, after generating the prediction image using the corresponding method, the predictor selects any prediction image. However, the predictor may select the method or mode based on the parameters used in the above encoding process before generating the prediction image, and may generate the prediction image according to the selected method or mode.

[0568] For example, the first method and the second method may be intra-frame prediction and inter-frame prediction, respectively, and the predictor may select the final prediction image of the current block from the prediction images generated according to the prediction method.

[0569] Fig.30 is a flowchart showing another example of the process executed by the predictor of the encoder 100.

[0570] First, the predictor generates a prediction image using intra prediction (step Sd_1a) and generates a prediction image using inter prediction (step Sd_1b). It should be noted that the prediction image generated by intra prediction is also referred to as an intra prediction image, and the prediction image generated by inter prediction is also referred to as an inter prediction image.

[0571] Next, the predictor evaluates each of the intra prediction image and the inter prediction image (step Sd_2). The above cost C can be used in the evaluation. The predictor can then select, from the intra prediction image and the inter prediction image, the prediction image for which the minimum cost C has been calculated as the final prediction image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the prediction image for the current block is selected.

[0572] The prediction processor then selects, from the intra prediction image and the inter prediction image, the prediction image for which the minimum cost C has been calculated as the final prediction image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the prediction image for the current block is selected.

[0573] (Intra Predictor)

[0574] The intra predictor 124 generates a prediction signal (i.e., an intra prediction image) by performing intra prediction (also referred to as prediction within a frame) of the current block by referring to one or more blocks in the current picture and stored in the block memory 118. More specifically, by referring to the pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block i, the intra predictor 124 generates an intra prediction image and then outputs the intra prediction image to the prediction controller 128.

[0575] For example, the intra predictor 124 performs intra prediction by using one of a plurality of defined intra prediction modes. Intra prediction modes generally include one or more non - directional prediction modes and a plurality of directional prediction modes. The defined modes can be predefined.

[0576] One or more non - directional prediction modes include, for example, the planar prediction mode and the DC prediction mode defined in the H.265 / High Efficiency Video Coding (HEVC) standard.

[0577] The plurality of directional prediction modes include, for example, thirty - three directional prediction modes defined in the H.265 / HEVC standard. It should be noted that, in addition to the thirty - three directional prediction modes, the plurality of directional prediction modes can also include thirty - two directional prediction modes (a total of sixty - five directional prediction modes). Fig.31It is a conceptual diagram for showing a total of sixty-seven intra prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra prediction. Solid arrows represent thirty-three directions defined in the H.265 / HEVC standard, and dashed arrows represent an additional thirty-two directions ( Fig.31 The two non-directional prediction modes are not shown).

[0578] In various processing examples, the luminance block can be referred to in the intra prediction of the chrominance block. In other words, the chrominance component of the current block can be predicted based on the luminance component of the current block. This intra prediction is also referred to as cross-component linear model (CCLM) prediction. An intra prediction mode of the chrominance block that references such a luminance block (also referred to as, for example, the CCLM mode) can be added as one of the intra prediction modes of the chrominance block.

[0579] The intra predictor 124 can correct the pixel values of the intra prediction based on the horizontal / vertical reference pixel gradients. The intra prediction accompanied by such correction is also referred to as position-dependent intra prediction combination (PDPC). Information indicating whether PDPC is applied (e.g., referred to as the PDPC flag) is typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).

[0580] Fig.32 It is a flowchart showing an example of the process performed by the intra predictor 124.

[0581] The intra predictor 124 selects one intra prediction mode from a plurality of intra prediction modes (step Sw_1). The intra predictor 124 then generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra prediction modes. For example, two of the six intra prediction modes can be the planar mode and the DC prediction mode, and the other four modes can be directional prediction modes. The intra predictor 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0582] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the intra prediction mode selected from these MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as prediction parameters by the entropy encoder 110.

[0583] When it is determined that the selected intra prediction mode is not included in the MPM (No in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flags. The intra predictor 124 then generates information indicating the intra prediction mode selected from among at least one intra prediction mode not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded as prediction parameters by the entropy encoder 110. The information indicating the intra prediction mode indicates any one of, for example, 0 to 60.

[0584] (Inter-frame predictor)

[0585] The inter-frame predictor 126 performs inter-frame prediction of the current block by referring to one or more blocks in a reference picture (also referred to as prediction between frames), and generates a predicted picture (inter-frame predicted picture), where the reference picture is different from the current picture and is stored in the frame memory 122. Inter-frame prediction is performed in units of the current block or a current sub-block in the current block (e.g., a 4×4 block). The sub-block is included in the block and is a unit smaller than the block. The size of the sub-block can be in the form of a slice, a tile, a picture, etc.

[0586] For example, the inter-frame predictor 126 performs motion estimation in the reference picture of the current block or current sub-block, and finds a reference block or reference sub-block that best matches the current block or current sub-block. The inter-frame predictor 126 then obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or reference sub-block to the current block or sub-block. The inter-frame predictor 126 generates an inter-frame predicted picture of the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter-frame predictor 126 outputs the generated inter-frame predicted picture to the prediction controller 128.

[0587] The motion information used in motion compensation can be signaled as an inter-frame prediction signal in various forms. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a motion vector predictor can be signaled.

[0588] (Reference picture list)

[0589] Fig.33 is a conceptual diagram for showing an example of a reference picture. Fig.34 is a conceptual diagram for showing an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. It should be noted that in Fig.33In it, each rectangle represents a picture, each arrow represents a picture reference relationship, the horizontal axis represents time, I, P, and B in the rectangle respectively represent an intra-predicted picture, a uni-predicted picture, and a bi-predicted picture, and the numbers in the rectangle represent the decoding order. As Fig.33 shown, the decoding order of the pictures is in the order of I0, P1, B2, B3, B4, and the display order of the pictures is in the order of I0, B3, B2, B4, P1. As Fig.34 shown, the reference picture list is a list representing reference picture candidates. For example, a picture (or slice) may include at least one reference picture list. For example, one reference picture list is used when the current picture is a uni-predicted picture, and two reference picture lists are used when the current picture is a bi-predicted picture. In Fig.33 and Fig.34 the example of, the picture B3 as the current picture currPic has two reference picture lists, namely the L0 list and the L1 list. When the current picture currPic is the picture B3, the reference picture candidates of the current picture currPic are I0, P1, B2, and the reference picture lists (i.e., the L0 list and the L1 list) indicate these pictures. The inter-frame predictor 126 or the prediction controller 128 specifies which picture in each reference picture list is actually to be referenced in the form of a reference picture index refidxLx. In Fig.34 it, the reference pictures P1 and B2 are specified by the reference picture indexes refIdxL0 and refIdxL1.

[0590] Such reference picture lists can be generated for each unit such as a sequence, a picture, a slice, a block, a CTU, or a CU. Additionally, among the reference pictures indicated in the reference picture list, the reference picture indexes indicating the reference pictures to be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, block level, CTU level, or CU level. Furthermore, a common reference picture list can be used in multiple inter-frame prediction modes.

[0591] (Basic Process of Inter-Frame Prediction)

[0592] Fig.35 is a flowchart showing an example basic processing flow of inter-frame prediction processing.

[0593] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Then, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Se_4).

[0594] Here, in the generation of the predicted image, the inter-frame predictor 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). In addition, in the determination of the MV, the inter-frame predictor 126 determines the MV by selecting a motion vector candidate (MV candidate) (step Se_1) and deriving the MV (step Se_2). The selection of the MV candidate is performed, for example, by the inter-frame predictor 126 generating an MV candidate list and selecting at least one MV candidate from the MV candidate list. It should be noted that the previously derived MV can be added to the MV candidate list. Alternatively, in the derivation of the MV, the inter-frame predictor 126 can also select at least one MV candidate from at least one MV candidate and determine the selected at least one MV candidate as the MV of the current block. Alternatively, the inter-frame predictor 126 can determine the MV of the current block by performing an estimation in the reference picture region specified by each of the selected at least one MV candidates. It should be noted that the estimation in the reference picture region can be referred to as motion estimation.

[0595] In addition, although steps Se_1 to Se_3 are performed by the inter-frame predictor 126 in the above example, the processes such as step Se_1 and step Se_2 can be performed by another component included in the encoder 100, for example.

[0596] It should be noted that an MV candidate list can be generated for each process in the inter-frame prediction mode, or a common MV candidate list can be used in multiple inter-frame prediction modes. The processes in steps Se_3 and Se_4 respectively correspond to Fig. 9 steps Sa_3 and Sa_4 shown in Fig.30 The process in step Se_3 corresponds to the process in step Sd_1b in

[0597] (Motion Vector Derivation Process)

[0598] Fig.36 is a flowchart showing an example of the derivation process of the motion vector.

[0599] The inter-frame predictor 126 can derive the MV of the current block in a mode for encoding motion information (e.g., MV). In this case, for example, the motion information can be encoded as a prediction parameter and signaled. In other words, the encoded motion information is included in the stream.

[0600] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where the motion information is not encoded. In this case, the motion information is not included in the stream.

[0601] Here, the MV derivation mode may include a conventional inter-frame mode, a conventional merge mode, a FRUC mode, an affine mode, etc., which will be described later. The modes for encoding motion information in the modes include a conventional inter-frame mode, a conventional merge mode, an affine mode (specifically, an affine inter-frame mode and an affine merge mode), etc. It should be noted that the motion information may include not only the MV but also the motion vector predictor selection information which will be described later. The modes that do not encode motion information include the FRUC mode, etc. The inter-frame predictor 126 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.

[0602] Fig.37 is a flowchart showing another example of the derivation of the motion vector.

[0603] The inter-frame predictor 126 may derive the MV of the current block in a mode that encodes the MV difference. In this case, for example, the MV difference may be encoded as a prediction parameter and signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. It should be noted that the MV predictor is a motion vector predictor.

[0604] Alternatively, the inter-frame predictor 126 may derive the MV in a mode that does not encode the MV difference. In this case, the encoded MV difference is not included in the stream.

[0605] Here, as described above, the MV derivation mode includes a conventional inter-frame mode, a conventional merge mode, a FRUC mode, an affine mode, etc., which will be described later. The modes that encode the MV difference in the modes include a conventional inter-frame mode, an affine mode (specifically, an affine inter-frame mode), etc. The modes that do not encode the MV difference include the FRUC mode, a conventional merge mode, an affine mode (specifically, an affine merge mode), etc. The inter-frame predictor 126 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.

[0606] (Motion Vector Derivation Mode)

[0607] Fig.38A and Fig.38B is a conceptual diagram for showing an example classification of the modes for MV derivation. For example, as Fig.38A shown, according to whether the motion information is encoded and whether the MV difference is encoded, the MV derivation mode is roughly classified into three modes. The three modes are an inter-frame mode, a merge mode, and a frame rate up-conversion (FRUC) mode. The inter-frame mode is a mode that performs motion estimation and encodes the motion information and the MV difference. For example, as Fig.38BAs shown, the inter - frame modes include the affine inter - frame mode and the regular inter - frame mode. The merge mode is a mode that does not perform motion estimation, selects the MV from the encoded surrounding blocks, and derives the MV of the current block using this MV. The merge mode is a mode that basically encodes the motion information without encoding the MV difference. For example, as Fig.38B shown, the merge modes include the regular merge mode (also known as the normal merge mode or the regular merge mode), the merge with motion vector difference (MMVD) mode, the combined inter - frame merge / intra - prediction (CIIP) mode, the triangular mode, the ATMVP mode, and the affine merge mode. Here, in the MMVD mode among the modes included in the merge mode, the MV difference is encoded exceptionally. It should be noted that the affine merge mode and the affine inter - frame mode are the modes included in the affine mode. The affine mode is a mode used to derive the MV of each of the multiple sub - blocks included in the current block as the MV of the current block under the assumption of an affine transformation. The FRUC mode is a mode that is used to derive the MV of the current block by performing an estimation between the encoded regions, and neither encodes the motion information nor encodes any MV difference. It should be noted that the corresponding modes will be described in more detail later.

[0608] It should be noted that Fig.38A and Fig.38B the classification of the modes shown in

[0609] (MV derivation > regular inter - frame mode)

[0610] The regular inter - frame mode is an inter - frame prediction mode that is used to derive the MV of the current block from the reference picture region specified by the MV candidates based on a block similar to the image of the current block. In this regular inter - frame mode, the MV difference is encoded.

[0611] Fig.39 is a flowchart showing an example of the inter - frame prediction process in the regular inter - frame mode.

[0612] First, the inter - frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple encoded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter - frame predictor 126 generates an MV candidate list.

[0613] Next, the inter-frame predictor 126 extracts N (an integer greater than or equal to 2) MV candidates from the plurality of MV candidates obtained in step Sg_1 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the determined priority order (step Sg_2). It should be noted that the priority order can be predetermined for each of the N MV candidates.

[0614] Next, the inter-frame predictor 126 selects one motion vector predictor candidate from the N motion vector predictor candidates as the motion vector predictor for the current block (also referred to as the MV predictor) (step Sg_3). At this time, the inter-frame predictor 126 encodes in the stream the motion vector predictor selection information for identifying the selected motion vector predictor. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.

[0615] Next, the inter-frame predictor 126 derives the MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter-frame predictor 126 also encodes in the stream the difference between the derived MV and the motion vector predictor as the MV difference. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130. It should be noted that the encoded reference picture is a picture including a plurality of blocks that have been reconstructed after being encoded.

[0616] Finally, by performing motion compensation on the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image of the current block (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a slice, the inter-frame prediction of the slice using the normal inter-frame mode ends. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode ends. It should be noted that in steps Sg_1 to Sg_5, not all blocks included in the slice can go through these processes, and when some blocks go through the processes, the inter-frame prediction of the slice using the normal inter-frame mode can end. This also applies to the processes in steps Sg_1 to Sg_5. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode can end.

[0617] It should be noted that the predicted image is an inter-frame prediction signal as described above. In addition, information indicating the inter-frame prediction mode (the normal inter-frame mode in the above example) for generating the predicted image is encoded as a prediction parameter in the encoded signal, for example.

[0618] Note that the MV candidate list can also be used as a list for use in another mode. Further, the processes related to the MV candidate list can be applied to the processes related to the list for use in another mode. The processes related to the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering of MV candidates, or deletion of MV candidates.

[0619] (MV Derivation > Regular Merge Mode)

[0620] The regular merge mode is an inter prediction mode for deriving an MV by selecting an MV candidate from the MV candidate list as the MV of the current block. Note that the regular merge mode is a type of merge mode and can be abbreviated as the merge mode. In the present embodiment, the regular merge mode and the merge mode are distinguished, and the merge mode is used in a broader sense.

[0621] Fig.40 is a flowchart showing an example of inter prediction in the regular merge mode.

[0622] First, the inter predictor 126 obtains a plurality of MV candidates for the current block based on information such as MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sh_1). In other words, the inter predictor 126 generates an MV candidate list.

[0623] Next, the inter predictor 126 selects one MV candidate from the plurality of MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter predictor 126 encodes in the stream MV selection information for identifying the selected MV candidate. In other words, the inter predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.

[0624] Finally, by performing motion compensation of the current block using the derived MV and the coded reference picture, the inter predictor 126 generates a predicted picture of the current block (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a slice, the inter prediction of the slice using the regular merge mode ends. Further, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a picture, the inter prediction of the picture using the regular merge mode ends. Note that not all blocks included in a slice can undergo the processes in steps Sh_1 to Sh_3, and when some blocks undergo the processes, the inter prediction of the slice using the regular merge mode can end. This also applies to the processes in steps Sh_1 to Sh_3. When the processes are performed for some blocks in a picture, the inter prediction of the picture using the regular merge mode can be completed.

[0625] In addition, information included in the encoded signal and representing an inter-frame prediction mode (in the above example, the regular merge mode) used to generate a predicted image is encoded as a prediction parameter in the stream, for example.

[0626] Fig.41 FIG. 5 is a conceptual diagram for showing an example of a motion vector derivation process of a current picture by the regular merge mode. First, the inter-frame predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of a plurality of encoded blocks located spatially around the current block; temporally adjacent MV candidates, which are MVs of surrounding blocks onto which the position of the current block in the encoded reference picture is projected; combined MV candidates, which are MVs generated by combining MV values of a spatially adjacent MV predictor and MV values of a temporally adjacent MV predictor; and zero MV candidates, which are MVs having a zero value.

[0627] Next, the inter-frame predictor 126 selects one MV candidate from among the plurality of MV candidates registered in the MV candidate list and determines the selected MV candidate as the MV of the current block.

[0628] In addition, the entropy encoder 110 writes and encodes in the stream a merge_idx, which is a signal indicating which MV candidate has been selected.

[0629] It should be noted that the MV candidates registered in the Fig.41 MV candidate list described in FIG. 5 are examples. The number of MV candidates may be different from the number of MV candidates in the figure, and the MV candidate list may be configured in such a way that some types of MV candidates in the figure may not be included, or one or more types of MV candidates other than the types of MV candidates in the figure may be included.

[0630] By using the MV of the current block derived by the regular merge mode to perform dynamic motion vector refresh (DMVR) to be described later, the final MV can be determined. It should be noted that in the regular merge mode, motion information is encoded and no MV difference is encoded. In the MMVD mode, one MV candidate is selected from the MV candidate list, just as in the case of the regular merge mode, and the MV difference is encoded. As shown in Fig.38B FIG. 6, MMVD can be classified as a merge mode together with the regular merge mode. It should be noted that the MV difference in the MMVD mode does not always need to be the same as the MV difference for the inter-frame mode. For example, the MV difference derivation in the MMVD mode may be a process that requires a smaller amount of processing than the MV difference derivation in the inter-frame mode.

[0631] In addition, a combined inter-frame merge / intra-frame prediction (CIIP) mode can be performed. This mode is used to overlap the prediction image generated in inter-frame prediction and the prediction image generated in intra-frame prediction to generate the prediction image of the current block.

[0632] It should be noted that the MV candidate list can be referred to as the candidate list. Additionally, merge_idx is MV selection information.

[0633] (MV derivation > HMVP mode)

[0634] Fig.42 is a conceptual diagram showing an example of the MV derivation process for the current picture using the HMVP merge mode. In the regular merge mode, the MV of a CU, for example, which is the current block, is determined by selecting one MV candidate from the MV list generated from the reference coded block (e.g., CU). Here, another MV candidate can be registered in the MV candidate list. The mode of registering such another MV candidate is called the HMVP mode.

[0635] In the HMVP mode, a first-in-first-out (FIFO) server for HMVP is used to manage MV candidates, separately from the MV candidate list of the regular merge mode.

[0636] In the FIFO buffer, motion information such as the MV of a block processed in the past is stored first and foremost. In managing the FIFO buffer, each time a block is processed, the MV of the latest block (i.e., the CU processed immediately before) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the CU processed earliest) is deleted from the FIFO buffer. In Fig.42 the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the oldest MV.

[0637] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is an MV different from all the MV candidates already registered in the MV candidate list of the regular merge mode starting from HMVP1. When it is determined that the MV is different from all the MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for the regular merge mode as an MV candidate. At this time, one or more MV candidates in the FIFO buffer can be registered (added to the MV candidate list).

[0638] By using the HMVP mode in this way, not only can the MVs of blocks adjacent to the current block in space or time be added, but also the MVs of blocks processed in the past can be added. As a result, the variation of the MV candidates in the regular merge mode is expanded, which increases the possibility of improving the coding efficiency.

[0639] Note that the MV can be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer can include not only the MV value, but also reference picture information, reference direction, number of pictures, etc. Additionally, the block can be, for example, a CU.

[0640] Note that Fig.42 the MV candidate list and the FIFO buffer shown in Fig.42 are examples. The sizes of the MV candidate list and the FIFO buffer can be different from Fig.42 those in

[0641] Note that the HMVP mode can be applied to modes other than the regular merge mode. For example, motion information such as the MV of a block that was previously processed in the affine mode can also be stored most recently and used as an MV candidate, which can promote better efficiency. The mode obtained by applying the HMVP mode to the affine mode can be referred to as the historical affine mode.

[0642] (MV derivation > FRUC mode)

[0643] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, on the decoder side, motion estimation is performed without using any pixel values in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values in the current block include frame rate up-conversion (FRUC) mode, pattern matching motion vector derivation (PMMVD) mode, etc.

[0644] Fig.43 An example of the FRUC process in flowchart form is shown in

[0645] Next, the best MV candidate is selected from among the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation value of the corresponding MV candidate included in the MV candidate list is calculated, and one MV candidate is selected based on the evaluation value. Based on the selected motion vector candidate, the motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (the best MV candidate) is directly derived as the motion vector of the current block. Additionally, for example, pattern matching in the surrounding area of the position in the reference picture corresponding to the selected motion vector candidate can be used to derive the motion vector of the current block. In other words, the estimation using pattern matching and the evaluation value can be performed in the surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces a better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, the update of the motion vector that produces a better evaluation value may not be performed.

[0646] Finally, by performing motion compensation on the current block using the derived MV and the encoded reference picture, the inter-frame predictor 126 generates a predicted image of the current block (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a picture, the inter-frame prediction of the picture using the FRUC mode ends. Note that not all blocks included in a slice go through the processes in steps Si_1 to Si_5, and when some blocks go through the processes, the inter-frame prediction of the slice using the FRUC mode can end. When the processes in steps Si_1 to Si_5 are performed for some blocks included in a picture in a similar manner, the inter-frame prediction of the picture using the FRUC mode can end.

[0647] A similar process can be performed on a sub-block basis.

[0648] The evaluation value can be calculated according to various methods. For example, a comparison is made between the reconstructed image in the region in the reference picture corresponding to the motion vector and the reconstructed image in a determined region (this region can be, for example, a region in another reference picture or a region in an adjacent block of the current picture, as described below). The determined region can be predetermined.

[0649] The difference between the pixel values of the two reconstructed images can be used for the evaluation value of the motion vector. Note that information other than the difference can be used to calculate the evaluation value.

[0650] Next, an example of pattern matching will be described in detail. First, one MV candidate included in the MV candidate list (e.g., the merge list) is selected as the estimated starting point through pattern matching. For example, as the pattern matching, the first pattern matching or the second pattern matching can be used. The first pattern matching and the second pattern matching can be respectively referred to as bilateral matching and template matching.

[0651] (MV Derivation>FRUC>Bilateral Matching)

[0652] In the first pattern matching, pattern matching is performed between two blocks that are located along the motion trajectory of the current block and are included in two different reference pictures. Therefore, in the first pattern matching, the region in the other reference picture along the motion trajectory of the current block is used as the determined region for calculating the evaluation value of the above candidate. The determined region can be predetermined.

[0653] Fig.44 is a conceptual diagram showing an example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along the motion trajectory. As Fig.44 shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the pair of the best matches among the pairs in two blocks that are included in two different reference pictures (Ref0, Ref1) and are located along the motion trajectory of the current block (Cur block). More specifically, for the current block, the difference between the reconstructed image at the specified position in the first coded reference picture (Ref0) specified by the MV candidate and the reconstructed image at the specified position in the second coded reference picture (Ref1) specified by the symmetric MV obtained by scaling the MV candidate by the display time interval is derived, and the obtained difference value is used to calculate the evaluation value. The MV candidate that produces the best evaluation value and may produce good results can be selected from multiple MV candidates as the final MV.

[0654] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) specifying the two reference blocks are proportional to the time distances (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the time distances from the current picture to the corresponding two reference pictures are equal to each other, mirror-symmetric bidirectional motion vectors are derived in the first pattern matching.

[0655] (MV Derivation>FRUC>Template Matching)

[0656] In the second mode matching (template matching), mode matching is performed between a block in a reference picture and a template in a current picture (the template is a block adjacent to the current block in the current picture (the adjacent block is, for example, an upper and / or left adjacent block)). Therefore, in the second mode matching, the adjacent blocks of the current block in the current picture are used as a determined region for calculating the evaluation value of the above-mentioned MV candidate.

[0657] Fig.45 is a conceptual diagram for showing an example of mode matching (template matching) between a template in a current picture and a block in a reference picture. As Fig.45 shown, in the second mode matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference picture (Ref0) that best matches the adjacent blocks of the current block in the current picture (Cur Pic). More specifically, the difference between the reconstructed image in the coding region adjacent to the left and above or left or above and the reconstructed image in the corresponding region in the coded reference picture (Ref0) specified by the MV candidate is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.

[0658] This information indicating whether to apply the FRUC mode (e.g., referred to as the FRUC flag) can be signaled at the CU level. In addition, when the FRUC mode is applied (e.g., when the FRUC flag is true), the information indicating the applicable mode matching method (e.g., the first mode matching or the second mode matching) can be signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0659] (MV derivation > Affine mode)

[0660] The affine mode is a mode that generates an MV using an affine transformation. For example, an MV can be derived for each sub-block based on the motion vectors of multiple adjacent blocks. This mode is also referred to as the affine motion compensation prediction mode.

[0661] Fig.46A is a conceptual diagram for showing an example of MV derivation for each sub-block based on the motion vectors of multiple adjacent blocks. In Fig.46A it, the current block includes, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the upper left control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (vx , v y )。

[0662] [Mathematical expression 1]

[0663]

[0664] Here, x and y represent the horizontal and vertical positions of the sub-block respectively, and w represents a determined weighting coefficient. The determined weighting coefficient can be pre-determined.

[0665] This information indicating the affine mode (e.g., called an affine flag) can be signaled at the CU level. Note that the signaling of the information indicating the affine mode does not necessarily need to be performed at the CU level and can be performed at another level (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0666] In addition, the affine mode can include several modes for different methods of deriving the motion vectors at the upper left and upper right control points. For example, the affine mode includes two modes: an affine inter mode (also called an affine normal inter mode) and an affine merge mode.

[0667] (MV derivation > affine mode)

[0668] Fig.46B is a conceptual diagram showing an example of MV derivation in units of sub-blocks in an affine mode using three control points. In Fig.46B , the current block includes, for example, sixteen 4×4 blocks. Here, the motion vector V0 at the upper left control point in the current block is derived based on the motion vectors of adjacent blocks. Here, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V2 at the lower left control point of the current block is derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v x , v y ) of the corresponding sub-blocks in the current block can be derived.

[0669] [Mathematical expression 2]

[0670] Here, x and y represent the horizontal and vertical positions of the sub-block respectively, and w and h can be weighting coefficients, which can be pre-determined

[0671]

[0672] weighting coefficients. In an embodiment, w can represent the width of the current block, and h can represent the height of the current block.

[0673] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0674] In addition, this affine mode using three control points may include different methods for deriving motion vectors at the upper left, upper right and lower left control points. For example, as in the case of the affine mode using two control points, the affine mode using three control points may include two modes, the affine inter-frame mode and the affine merge mode.

[0675] Note that in the affine mode, the size of each sub-block included in the current block may not be limited to 4×4 pixels, and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.

[0676] (MV Derivation > Affine Mode > Control Points)

[0677] Fig.47A , Fig.47B and Fig.47C is a conceptual diagram for illustrating an example of MV derivation at a control point in an affine mode.

[0678] like Fig.47A As shown, in the affine mode, for example, based on a plurality of motion vectors corresponding to blocks encoded according to the affine mode among the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) adjacent to the current block, a motion vector predictor at a corresponding control point of the current block is calculated. More specifically, the encoded blocks A (left), B (upper), C (upper right), D (lower left), and E (upper left) are checked in the order listed, and the first valid block encoded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on a plurality of motion vectors corresponding to the identified blocks.

[0679] For example, Fig.47B As shown, when block A adjacent to the left side of the current block has been encoded according to the affine mode using two control points, motion vectors v3 and v4 projected at the upper left corner position and the upper right corner position of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated.

[0680] For example, Fig.47CAs shown, when the block A adjacent to the left side of the current block has been encoded according to the affine mode using three control points, the motion vectors v3, v4, and v5 projected at the upper left corner position, upper right corner position, and lower left corner position of the encoded block including the block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block are calculated.

[0681] Figures 47A to 47C The MV derivation method shown can be used in the MV derivation at each control point of the current block in Fig.50 the step Sk_1 shown, or can be used for the MV predictor derivation at each control point of the current block in Fig.51 the step Sj_1 shown described later.

[0682] Fig.48A and Fig.48B are conceptual diagrams for showing examples of MV derivation at control points in the affine mode.

[0683] Fig.48A is a conceptual diagram for showing an exemplary affine mode using two control points.

[0684] In the affine mode, as Fig.48A shown, the MV selected from the MVs at the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block.

[0685] Fig.48B is a conceptual diagram for showing an exemplary affine mode using three control points.

[0686] In the affine mode, as Fig.48B shown, the MV selected from the MVs at the encoded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block. In addition, the MV selected from the MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 at the lower left corner control point of the current block.

[0687] Note that Fig.48A and Fig.48B the MV derivation method shown can be used in the MV derivation at each control point of the current block in Fig.50 the step Sk_1 shown described later, or can be used for Fig.51 MV predictor derivation at each control point of the current block in step Sj_1 shown in

[0688] Here, when affine modes with different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level, the number of control points of the coded block and the number of control points of the current block can be different from each other.

[0689] Fig.49A and Fig.49B are conceptual diagrams showing examples of methods for MV derivation at control points when the number of control points of the coded block and the number of control points of the current block are different from each other.

[0690] For example, as Fig.49A shown, the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the left side of the current block has been coded according to the affine mode using two control points. In this case, the motion vectors v3 and v4 projected at the upper left position and upper right position in the coded block including block A are derived. Then, the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3 and v4. In addition, the motion vector v2 at the lower left control point is calculated based on the derived motion vectors v0 and v1.

[0691] For example, as Fig.49B shown, the current block has two control points at the upper left corner and upper right corner, and the block A adjacent to the left side of the current block has been coded according to the affine mode using three control points. In this case, the motion vectors v3, v4, and v5 projected at the upper left position, upper right position, and lower left position in the coded block including block A are derived. Then, the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.

[0692] Note that Fig.49A and Fig.49B the MV derivation methods shown can be used for MV derivation at each control point of the current block in step Sk_1 shown in Fig.50 or can be used for MV predictor derivation at each control point of the current block in step Sj_1 shown in Fig.51 shown later.

[0693] (MV Derivation > Affine Mode > Affine Merge Mode)

[0694] Fig.50 is a flowchart showing an example of the process in the affine merge mode.

[0695] In the affine merge mode as shown in the figure, first, the inter-frame predictor 126 derives the MVs at the corresponding control points of the current block (step Sk_1). As Fig.46A shown, the control points are the upper left corner point and the upper right corner point of the current block, or as Fig.46B shown, the upper left corner point, the upper right corner point, and the lower left corner point of the current block. The inter-frame predictor 126 can encode MV selection information for identifying two or three derived MVs in the stream.

[0696] For example, when using the FIG. 47A to FIG. 47C shown MV derivation method, as Fig.47A shown, the inter-frame predictor 126 checks the encoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) in the listed order and identifies the first valid block encoded according to the affine mode.

[0697] The inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine mode to derive the MVs at the control points. For example, when block A is identified and block A has two control points, as Fig.47B shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left corner control point of the current block and the motion vector v1 at the upper right corner control point of the current block according to the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block including block A. For example, the inter-frame predictor 126 projects the motion vectors v3 and v4 at the upper left corner and the upper right corner of the encoded block onto the current block to calculate the motion vector v0 at the upper left corner control point of the current block and the motion vector v1 at the upper right corner control point of the current block.

[0698] Alternatively, when block A is identified and block A has three control points, as Fig.47C shown, the inter-frame predictor 126 calculates the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block according to the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the encoded block including block A. For example, the inter-frame predictor 126 projects the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the encoded block onto the current block to calculate the motion vector v0 at the upper left corner control point of the current block, the motion vector v1 at the upper right corner control point of the current block, and the motion vector v2 at the lower left corner control point of the current block.

[0699] Note that, as described above in Fig.49A shown, when block A is identified and block A has two control points, the MVs at three control points can be calculated, and as described above in Fig.49B As shown, when block A is recognized and block A has three control points, the MVs at two control points can be calculated.

[0700] Next, the inter-frame predictor 126 performs motion compensation on each of a plurality of sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the MV of each of the plurality of sub-blocks as an affine MV (step Sk_2) using, for example, two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-blocks (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all the sub-blocks included in the current block, the process of generating the prediction image using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate the prediction image of the current block.

[0701] Note that the above MV candidate list can be generated in step Sk_1. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example Figures 47A to 47C the MV derivation method shown in Fig.48A and Fig.48B the MV derivation method shown in Fig.49A and Fig.49B the MV derivation method shown in and any combination of other MV derivation methods.

[0702] Note that in addition to the affine mode, the MV candidate list can include MV candidates in a mode that performs prediction on a sub-block basis.

[0703] Note that, for example, an MV candidate list (which includes MV candidates in the affine merge mode using two control points and the affine merge mode using three control points) can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points can be generated. The MV candidate can be, for example, the MV for the encoded block A (left), block B (upper), block C (upper right), block D (lower left), and block E (upper left), or the MV of the valid block in the block.

[0704] Note that the index indicating one of the MVs in the MV candidate list can be transmitted as the MV selection information.

[0705] (MV Derivation > Affine Mode > Affine Inter - Frame Mode)

[0706] Fig.51 It is a flowchart showing an example of the process in the affine inter - frame mode.

[0707] In the affine inter - frame mode, first, the inter - frame predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the upper - left corner point, the upper - right corner point, and the lower - right corner point of the current block, as Fig.46A or Fig.46B shown.

[0708] For example, when using the MV derivation method shown in Fig.48A and Fig.48B , the inter - frame predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block by selecting the MV of any block in the coded blocks near the corresponding control points of the current block shown in Fig.48A or Fig.48B . At this time, the inter - frame predictor 126 encodes in the stream the MV predictor selection information for identifying the selected two or three MV predictors.

[0709] For example, the inter - frame predictor 126 can determine, using cost evaluation, etc., the block from which the MV is selected as the MV predictor at the control point from among the coded blocks adjacent to the current block, and can write a flag in the bitstream indicating which MV predictor has been selected. In other words, the inter - frame predictor 126 outputs the MV predictor selection information such as a flag as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.

[0710] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4), while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 calculates the MV of each sub-block corresponding to the updated MV predictor as an affine MV using the above expression (1A) or expression (1B) (step Sj_3). The inter-frame predictor 126 then performs motion compensation for the sub-blocks using these affine MVs and the encoded reference pictures (step Sj_4). When the MV predictor is updated in step Sj_2, the processes in steps Sj_3 and Sj_4 are performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor in the stream as an MV difference. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.

[0711] Finally, the inter-frame predictor 126 generates a predicted image of the current block by performing motion compensation for the current block using the determined MV and the encoded reference pictures (step Sj_6).

[0712] Note that the above MV candidate list can be generated in step Sj_1. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example, Figures 47A to 47C the MV derivation methods shown in Fig.48A and Fig.48B the MV derivation methods shown in Fig.49A and Fig.49B the MV derivation methods shown in

[0713] Note that in addition to the affine mode, the MV candidate list can include MV candidates in a mode that performs prediction in units of sub-blocks.

[0714] Note that, for example, an MV candidate list including MV candidates in an affine inter - frame mode using two control points and an affine inter - frame mode using three control points can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine inter - frame mode using two control points and an MV candidate list including MV candidates in the affine inter - frame mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine inter - frame mode using two control points and the affine inter - frame mode using three control points can be generated. The MV candidate can be, for example, the MV for block A (left), block B (above), block C (upper - right), block D (lower - left), and block E (upper - left) for encoding, or the MV of the valid block in the block.

[0715] Note that an index indicating one of the MV candidates in the MV candidate list can be transmitted as MV predictor selection information.

[0716] (MV derivation > triangle mode)

[0717] In the above example, the inter - frame predictor 126 generates a rectangular prediction image for the current rectangular block. However, the inter - frame predictor 126 can generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and can combine multiple prediction images to generate a final rectangular prediction image. The shape different from the rectangle can be, for example, a triangle.

[0718] Fig.52A is a conceptual diagram for showing the generation of two triangle prediction images.

[0719] The inter - frame predictor 126 generates a triangle prediction image by performing motion compensation on a first partition with a triangle shape in the current block using the first MV of the first partition to generate a triangle prediction image. Similarly, the inter - frame predictor 126 generates a triangle prediction image by performing motion compensation on a second partition with a triangle shape in the current block using the second MV of the second partition to generate a triangle prediction image. Then, the inter - frame predictor 126 generates a prediction image with a rectangular shape identical to the rectangular shape of the current block by combining these prediction images.

[0720] Note that a first prediction image with a rectangular shape corresponding to the current block can be generated using the first MV as the prediction image for the first partition. In addition, a second prediction image with a rectangular shape corresponding to the current block can be generated using the second MV as the prediction image for the second partition. The prediction image of the current block can be generated by performing weighted addition of the first prediction image and the second prediction image. Note that the part where the weighted addition is performed can be a partial region across the boundary between the first partition and the second partition.

[0721] Fig.52BIt is a conceptual diagram for showing the first part of the first partition that overlaps with the second partition and examples of the first and second sample sets that can be weighted as part of the correction process. The first part can be, for example, one-fourth of the width or height of the first partition. In another example, the first part can have a width corresponding to N samples adjacent to the edge of the first partition, where N is an integer greater than zero. For example, N can be the integer 2. As shown in the figure, Fig.52B The left example in shows a rectangular partition with a rectangular part, whose width is one-fourth of the width of the first partition, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Fig.52B The central example in shows a rectangular partition with a rectangular part, whose height is one-fourth of the height of the first partition, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part. Fig.52B The right example in shows a triangular partition with a polygonal part, whose height corresponds to two samples, where the first sample set includes samples outside the first part and samples inside the first part, and the second sample set includes samples within the first part.

[0722] The first part can be the part of the first partition that overlaps with the adjacent partition. Fig.52C It is a conceptual diagram for showing the first part of the first partition, which is the part of the first partition that overlaps with a part of the adjacent partition. For ease of illustration, a rectangular partition with an overlapping part with a spatially adjacent rectangular partition is shown. Partitions with other shapes, such as triangular partitions, can be used, and the overlapping part can overlap with a partition adjacent in space or time.

[0723] In addition, although examples of generating a predicted image for each of the two partitions using inter-frame prediction are given, intra-frame prediction can be used to generate a predicted image for at least one partition.

[0724] Fig.53 It is a flowchart showing an example of the process in the triangular mode.

[0725] In the triangular mode, first, the inter-frame predictor 126 divides the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode the partition information (which is information related to the divided partitions) as prediction parameters in the stream. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.

[0726] First, the inter-frame predictor 126 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sx_2). In other words, the inter-frame predictor 126 generates a list of MV candidates.

[0727] The inter-frame predictor 126 then respectively selects an MV candidate for the first partition and an MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_1 as the first MV and the second MV (step Sx_3). At this time, the inter-frame predictor 126 encodes, in the stream, MV selection information for identifying the selected MV candidates as prediction parameters. In other words, the inter-frame predictor 126 outputs the MV selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.

[0728] Next, the inter-frame predictor 126 generates a first prediction image by performing motion compensation using the selected first MV and the coded reference picture (step Sx_4). Similarly, the inter-frame predictor 126 generates a second prediction image by performing motion compensation using the selected second MV and the coded reference picture (step Sx_5).

[0729] Finally, the inter-frame predictor 126 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_6).

[0730] Note that although the first partition and the second partition are triangles in the Fig.52A illustrated example, the first partition and the second partition may be trapezoids, or other shapes different from each other. Further, although the current block includes two partitions in the Fig.52A and Fig.52C illustrated examples, the current block may include three or more partitions.

[0731] In addition, the first partition and the second partition may overlap each other. In other words, the first partition and the second partition may include the same pixel region. In this case, the prediction image in the first partition and the prediction image in the second partition may be used to generate the prediction image for the current block.

[0732] In addition, although an example of generating a prediction image for each of the two partitions using inter-frame prediction has been shown, an intra-frame prediction may be used to generate a prediction image for at least one partition.

[0733] Note that the list of MV candidates for selecting the first MV and the list of MV candidates for selecting the second MV may be different from each other, or the list of MV candidates for selecting the first MV may also be used as the list of MV candidates for selecting the second MV.

[0734] Note that the partition information may include an index indicating a partitioning direction in which at least the current block is divided into a plurality of partitions. The MV selection information may include an index indicating the selected first MV and an index indicating the selected second MV. One index may indicate multiple pieces of information. For example, one index that commonly indicates a part or all of the partition information and a part or all of the MV selection information may be encoded.

[0735] (MV Derivation > ATMVP Mode)

[0736] Fig.54 is a conceptual diagram showing an example of an Advanced Temporal Motion Vector Prediction (ATMVP) mode for deriving an MV in units of sub - blocks.

[0737] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, the MV candidates for each sub - block are registered in the MV candidate list for the regular merge mode.

[0738] More specifically, in the ATMVP mode, first, as Fig.54 shown, a temporal MV reference block associated with the current block is identified in the coded reference picture specified by the MV (MV0) of an adjacent block located at the lower - left position relative to the current block. Next, in each sub - block of the current block, an MV used to encode the region corresponding to the sub - block in the temporal MV reference block is identified. The MV identified in this way is included in the MV candidate list as an MV candidate for the sub - block in the current block. When an MV candidate for each sub - block is selected from the MV candidate list, the sub - block undergoes motion compensation, where the MV candidate is used as the MV of the sub - block. In this way, a predicted image for each sub - block is generated.

[0739] Although in the Fig.54 example shown, the block located at the lower - left position relative to the current block is used as the surrounding MV reference block, it should be noted that another block may be used. Additionally, the size of the sub - block may be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub - block may be switched for units such as slices, tiles, pictures, etc.

[0740] (MV Derivation > DMVR)

[0741] Fig.55 is a flowchart showing the relationship between the merge mode and Decoder Motion Vector Refinement (DMVR).

[0742] The inter-frame predictor 126 derives the motion vector of the current block according to the merge mode (step S1_1). Next, the inter-frame predictor 126 determines whether to perform the estimation of the motion vector, i.e., motion estimation (step S1_2). Here, when it is determined not to perform motion estimation (No in step S1_2), the inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merge mode.

[0743] When it is determined to perform motion estimation in step S1_1 (Yes in step S1_2), the inter-frame predictor 126 derives the final motion vector of the current block by estimating the surrounding area of the reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to DMVR.

[0744] Fig.56 is a conceptual diagram showing an example of the DMVR process for determining the MV.

[0745] First, for example, in the merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference picture (L0) which is an encoded picture in the L0 list according to the MV candidate (L0). Similarly, reference pixels are identified from the second reference picture (L1) which is an encoded picture in the L1 list according to the MV candidate (L1). A template is generated by calculating the average value of these reference pixels.

[0746] Next, each of the surrounding areas of the MV candidates in the first reference picture (L0) and the second reference picture (L1) is estimated using the template, and the MV that generates the minimum cost is determined as the final MV. Note that the cost can be calculated, for example, using the difference between each pixel value in the template and the corresponding pixel value in the estimated area, the value of the MV candidate, etc.

[0747] It is not always necessary to perform exactly the same process described here. Other processes can be used for deriving the final MV through the estimation in the surrounding area of the MV candidate.

[0748] Fig.57 is a conceptual diagram showing another example of DMVR for determining the MV. Different from Fig.56 the example of DMVR shown in Fig.57 in the example shown in

[0749] First, the inter-frame predictor 126 estimates the surrounding area of the reference block in each reference picture included in the L0 list and the L1 list based on the initial MVs that are MV candidates obtained from each MV candidate list. For example, as Fig.57 shown, the initial MV corresponding to the reference block in the L0 table is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets the search position for the reference picture in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set (specifically, the initial MV (i.e., InitMV_L0)), the vector difference from the search position is MVd_L0. The inter-frame predictor 126 then determines the estimated position in the reference picture in the L1 list. This search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference as MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines the position symmetric to the position indicated by the initial MV as the search position in each reference picture in the L0 list and the L1 list. The inter-frame predictor 126 calculates the sum of the absolute differences (SAD) of the pixel values at the search positions in the block as the cost for each search position and finds the search position that produces the minimum cost.

[0750] Fig.58A is a conceptual diagram for illustrating an example of motion estimation in DMVR, and Fig.58B is a flowchart showing an example of the process of motion estimation.

[0751] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position (also referred to as the starting point) indicated by the initial MV and the eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at the search position other than the starting point is the minimum, the inter-frame predictor 126 changes the target to the search position that obtains the minimum cost and executes the process in step 2. When the cost at the starting point is the minimum, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.

[0752] In step 2, the inter-frame predictor 126 performs a search similar to the process in step 1, and regards the search position after the target change as the new starting point according to the result of the process in step 1. Then the inter-frame predictor 126 determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at the search position other than the starting point is the minimum, the inter-frame predictor 126 executes the process in step 4. When the cost at the starting point is the minimum, the inter-frame predictor 126 executes the process in step 3.

[0753] In step 4, the inter-frame predictor 126 regards the search position at the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0754] In step 3, the inter-frame predictor 126 determines the pixel position with sub-pixel accuracy that obtains the minimum cost based on the costs at four points located above, below, left, and right with respect to the starting point in step 1 or step 2, and regards this pixel position as the final search position. The pixel position at sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) above, below, left, and right, using the cost at the corresponding search position among the four search positions as a weight. The inter-frame predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0755] (Motion Compensation > BIO / OBMC / LIC)

[0756] Motion compensation involves modes for generating a predicted image and correcting the predicted image. Examples of the modes are bidirectional optical flow (BIO), overlapping block motion compensation (OBMC), local illumination compensation (LIC), etc., which will be described later.

[0757] Fig.59 is a flowchart showing an example of the generation process of the predicted image.

[0758] The inter-frame predictor 126 generates a predicted image (step Sm_1), and corrects the predicted image according to any of the above-mentioned modes, for example (step Sm_2).

[0759] Fig.60 is a flowchart showing another example of the generation process of the predicted image.

[0760] The inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame predictor 126 generates a predicted image using the motion vector (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter-frame predictor 126 generates a final predicted image by correcting the predicted image (step Sn_4). Note that in the LIC described later, the luminance and chrominance can be corrected in step Sn_4. When it is determined not to perform the correction process (No in step Sn_3), the inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).

[0761] (Motion Compensation > OBMC)

[0762] Note that, in addition to the motion information of the current block obtained by motion estimation, the motion information of adjacent blocks can also be used to generate an inter-predicted image. More specifically, for each sub-block in the current block, an inter-predicted image can be generated by performing a weighted addition of a predicted image (in a reference picture) based on the motion information obtained by motion estimation and a predicted image (in the current picture) based on the motion information of adjacent blocks. This inter-prediction (motion compensation) is also referred to as overlapping block motion compensation (OBMC) or OBMC mode.

[0763] In the OBMC mode, information indicating the sub-block size of OBMC (e.g., referred to as the OBMC block size) can be signaled at the sequence level. In addition, information indicating whether the OBMC mode is applied (e.g., referred to as the OBMC flag) can be signaled at the CU level. Note that the signaling of such information does not necessarily need to be performed at the sequence level and the CU level, and can be performed at another level (e.g., picture level, slice level, tile level, CTU level, or sub-block level).

[0764] The OBMC mode will be described in more detail. Fig.61 and Fig.62 are a flowchart and a conceptual diagram for showing an outline of a predicted image correction process performed by OBMC.

[0765] First, as Fig.62 shown, a predicted image (Pred) obtained by conventional motion compensation is obtained using the MV assigned to the current block. In Fig.62 , the arrow "MV" points to the reference picture and indicates what the current block of the current picture refers to in order to obtain the predicted image.

[0766] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) that has been derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which indicates the reference picture from the current block. A first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.

[0767] Similarly, a predicted image (Pred_U) is obtained by applying an MV (MV_U) that has been derived for an encoded block adjacent to the current block above (reusing the MV of the current block) to the current block. The MV (MV_U) is indicated by an arrow "MV_U" that indicates the reference picture from the current block. A second correction of the predicted image is performed by overlapping the predicted image Pred_U with predicted images (e.g., Pred and Pred_L) for which a first correction has already been performed. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is thus the final predicted image of the current block.

[0768] Although the above example is a two-way correction method using left and upper adjacent blocks, note that the correction method can also be a three-way or more-way correction method that also uses right and / or lower adjacent blocks.

[0769] Note that the region where this overlapping is performed can be only a part of the region near the block boundary, rather than the entire pixel region of the block.

[0770] Note that the predicted image correction process for obtaining one predicted image Pred from one reference picture by overlapping additional predicted images Pred_L and Pred_U according to OBMC has been described above. However, when correcting a predicted image based on multiple reference images, a similar process can be applied to each of the multiple reference pictures. In this case, after obtaining corrected predicted images from the respective reference pictures by performing OBMC image correction based on the multiple reference pictures, the obtained corrected predicted images are further overlapped to obtain a final predicted image.

[0771] Note that in OBMC, the current block unit can be a PU, or a sub-block unit obtained by further dividing a PU.

[0772] An example of a method for determining whether to apply OBMC is a method for using a signal obmc_flag that indicates whether to apply OBMC. As a specific example, the encoder 100 can determine whether the current block belongs to a region with complex motion. The encoder 100 sets the obmc_flag to the value "1" when the block belongs to a region with complex motion and applies OBMC during encoding, and sets the obmc_flag to the value "0" when the block does not belong to a region with complex motion and encodes the block without applying OBMC. The decoder 200 switches between the application and non-application of OBMC by decoding the obmc_flag written in the stream.

[0773] (Motion Compensation > BIO)

[0774] Next, the MV derivation method is described. First, the mode of deriving MV based on the model assuming uniform linear motion is described. This mode is also called the bidirectional optical flow (BIO) mode. Additionally, this bidirectional optical flow can be written as BDOF instead of BIO.

[0775] Fig.63 is a conceptual diagram for showing the model assuming uniform linear motion. In Fig.63 , (v x , v y ) represents the velocity vector, and τ0 and τ1 represent the time distances between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) represents the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) represents the MV corresponding to the reference picture Ref1.

[0776] Here, under the assumption of uniform linear motion exhibited by the velocity vector (v x , v y ), (MV x0 , MV y0 ) and (MV x1 , MV y1 ) are respectively expressed as (v xτ0 , v yτ0 ) and (-v xτ1 , -v yτ1 ), and the following optical flow equation (2) is given.

[0777] [Mathematical expression 3]

[0778] Here, I(k) represents the motion-compensated luminance value of the reference picture k (k = 0, 1) after motion compensation. This optical flow equation

[0779]

[0780] shows that the sum of the following items is zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected on a pixel-by-pixel basis.

[0781] Note that a method different from the model based on the assumption of uniform linear motion can be used on the decoder side 200 to derive the motion vector. For example, the motion vector can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0782] Fig.64 is a flowchart showing an example of an inter-frame prediction process according to BIO. Fig.65 is a functional block diagram showing an example of the functional configuration of an inter-frame predictor 126 that can perform inter-frame prediction according to BIO.

[0783] As Fig.65 shown, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolated image derivator 126b, a gradient image derivator 126c, an optical flow derivator 126d, a correction value derivator 126e, and a predicted image corrector 126f. Note that the memory 126a can be a frame memory 122.

[0784] The inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) including a current block. Then, the inter-frame predictor 126 derives a predicted image of the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.

[0785] Next, the interpolated image derivator 126b derives an interpolated image I 0 of the current block by referring to the memory 126a using the motion vector M0 and the reference picture L0. Next, the interpolated image derivator 126b derives an interpolated image I 1 of the current block by referring to the memory 126a using the motion vector M1 and the reference picture L1 (step Sy_2). Here, the interpolated image I 0 is an image included in the reference picture Ref0 and to be derived for the current block, and the interpolated image I 1 is an image included in the reference picture Ref1 and to be derived for the current block. Each of the interpolated image I 0 and the interpolated image I 1 can be the same size as the current block. Alternatively, each of the interpolated image I 0 and the interpolated image I 1 can be an image larger than the current block. In addition, the interpolated image I 0 and the interpolated image I 1 can include a predicted image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.

[0786] In addition, the gradient image derivator 126c derives the gradient image of the current block from the interpolated image I 0 and the interpolated image I 1 (Ix 0 , Ix 1 , Iy 0 , Iy 1 )(step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image derivator 126c can derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image can indicate the amount of spatial change in the pixel values in the horizontal direction, in the vertical direction, or both.

[0787] Next, the optical flow derivator 126d uses the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) to derive the optical flow (vx, vy) as the velocity vector for each sub-block of the current block (step Sy_4). The optical flow indicates the coefficient for correcting the amount of spatial pixel movement and can be referred to as a local motion estimate, a corrected motion vector, or a corrected weighted vector. As an example, the sub-block can be a 4×4 pixel sub-CU. Note that the optical flow derivation can be performed for each pixel unit or the like, rather than for each sub-block.

[0788] Next, the inter-frame predictor 126 uses the optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value derivator 126e uses the optical flow (vx, vy) to derive the correction value of the pixel values included in the current block (step Sy_5). The predicted image corrector 126f then uses the correction value to correct the predicted image of the current block (step Sy_6). Note that the correction value can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.

[0789] Note that the BIO process flow is not limited to Fig.64 the process disclosed in Fig.64 . For example, only a part of the process disclosed in

[0790] (Motion Compensation > LIC)

[0791] Next, an example of a mode for generating a predicted image (prediction) using a local illumination compensation (LIC) process is described.

[0792] Fig.66A is a conceptual diagram of an example of a process for a predicted image generation method that shows a luminance correction process performed by LIC. Fig.66B is a flowchart of an example of a process for a predicted image generation method using LIC.

[0793] First, the inter-frame predictor 126 derives an MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1).

[0794] Next, the inter-frame predictor 126 extracts information indicating how the luminance value changes between the current block and the reference picture for the current block (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region) in the current picture, and the luminance pixel values at the corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 126 uses the information indicating how the luminance value changes to calculate a luminance correction parameter (step Sz_3).

[0795] The inter-frame predictor 126 generates a predicted image of the current block by performing a luminance correction process in which the luminance correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_4). In other words, the predicted image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameter. In this correction, the luminance can be corrected, or the chrominance can be corrected, or both. In other words, a chrominance correction parameter can be calculated using the information indicating how the chrominance changes, and a chrominance correction process can be performed.

[0796] Note that Fig.66A the shape of the surrounding reference region shown in

[0797] is an example; another shape can be used.

[0798] An example of a method for determining whether to apply the LIC is a method using a lic_flag as a signal indicating whether to apply the LIC. As a specific example, the encoder 100 determines whether the current block belongs to a region with a luminance change. When the block belongs to a region with a luminance change, the encoder 100 sets the lic_flag to the value "1" and applies the LIC during encoding, and when the block does not belong to a region with a luminance change, the lic_flag is set to the value "0" and encoding is performed without applying the LIC. The decoder 200 can decode the lic_flag in the written stream and decode the current block by switching between the application and non-application of the LIC according to the flag value.

[0799] An example of a different method for determining whether to apply the LIC process is a determination method based on whether the LIC process has been applied to surrounding blocks. As a specific example, when the current block has been processed in the merge mode, the inter-frame predictor 126 determines whether the encoded surrounding blocks selected in the MV derivation in the merge mode have been encoded using the LIC. The inter-frame predictor 126 performs encoding by switching between the application and non-application of the LIC according to the result. Note that also in this example, the same process is applied in the process on the decoder 200 side.

[0800] The luminance correction (LIC) process has been described with reference to Fig.66A and Fig.66B and is further described below.

[0801] First, the inter-frame predictor 126 derives an MV for obtaining a reference image corresponding to the current block to be encoded from a reference picture that is an encoded picture.

[0802] Next, the inter-frame predictor 126 uses the luminance pixel values of the encoded surrounding reference regions adjacent to the left and above of the current block and the luminance values at the corresponding positions in the reference picture specified by the MV to extract information indicating how the luminance value of the reference picture changes to the luminance value of the current picture, and calculates a luminance correction parameter. For example, assume that the luminance pixel value of a given pixel in the surrounding reference region in the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region in the reference picture is p1. The inter-frame predictor 126 calculates the coefficients A and B for optimizing A×p1 + B = p0 as the luminance correction parameters for multiple pixels in the surrounding reference region.

[0803] Next, the inter-frame predictor 126 performs a luminance correction process using the luminance correction parameter of the reference image in the reference picture specified by the MV to generate a predicted image of the current block. For example, assume that the luminance pixel value in the reference image is p2, and the luminance pixel value after luminance correction of the predicted image is p3. The inter-frame predictor 126 generates a predicted image after undergoing the luminance correction process by calculating A×p2 + B = p3 for each pixel in the reference image.

[0804] For example, a region having a determined number of pixels extracted from each of the upper adjacent pixel and the left adjacent pixel can be used as a surrounding reference region. Additionally, the surrounding reference region is not limited to a region adjacent to the current block, and can be a region not adjacent to the current block. In Fig.66A the example shown, the surrounding reference region in the reference picture can be a region specified by another MV in the surrounding reference region in the current picture. For example, the another MV can be an MV in the surrounding reference region in the current picture.

[0805] Although the operations performed by the encoder 100 have been described herein, it should be noted that the decoder 200 performs similar operations.

[0806] Note that the LIC can be applied not only to luminance but also to chrominance. At this time, the correction parameters can be separately derived for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of Y, Cb, and Cr.

[0807] Furthermore, the LIC process can be applied in units of sub-blocks. For example, the correction parameters can be derived using the surrounding reference region in the current sub-block and the surrounding reference region in the reference sub-block in the reference picture specified by the MV of the current sub-block.

[0808] (Prediction Controller)

[0809] The prediction controller 128 selects one of an intra-frame prediction signal (an image or signal output from the intra-frame predictor 124) and an inter-frame prediction signal (an image or signal output from the inter-frame predictor 126), and outputs the selected predicted image to the subtractor 104 and the adder 116 as a prediction signal.

[0810] (Prediction Parameter Generator)

[0811] The prediction parameter generator 130 may output information related to intra prediction, inter prediction, selection of a predicted image in the prediction controller 128, etc. as prediction parameters to the entropy encoder 110. The entropy encoder 110 may generate a bitstream based on the prediction parameters input from the prediction parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameters may be used in the decoder 200. The decoder 200 may receive and decode the bitstream and perform the same processes as the prediction processes performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameters may include, for example, (i) selection of a prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) an optional index, flag, or value based on or indicating the prediction processes performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128.

[0812] (Decoder)

[0813] Next, the decoder 200 capable of decoding the bitstream output from the above-described encoder 100 will be described. Fig.67 FIG. is a block diagram showing a functional configuration of the decoder 200 according to the present embodiment. The decoder 200 is a device that decodes a bitstream as an encoded image in units of blocks.

[0814] As Fig.67 shown, the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformator 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a prediction parameter generator 222, and a segmentation determiner 224. Note that the intra predictor 216 and the inter predictor 218 are configured as part of prediction executors.

[0815] (Example of Decoder Installation)

[0816] Fig.68 FIG. is a functional block diagram showing an example of the installation of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, Fig.67 as shown, a plurality of components of the decoder 200 are installed on Fig.68 the processor b1 and the memory b2 shown.

[0817] The processor b1 is a circuit that performs information processing and is coupled to the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a bitstream. The processor b1 may be a processor such as a CPU. In addition, the processor b1 may be an aggregate of a plurality of electronic circuits. Further, for example, the processor b1 may undertake Fig.67The roles of two or more of the various elements of the decoder 200 shown, other than the elements for storing information.

[0818] Memory b2 is a dedicated or general-purpose memory for storing information used by processor b1 to decode the stream. Memory b2 can be an electronic circuit and can be connected to processor b1. In addition, memory b2 can be included in processor b1. In addition, memory b2 can be an aggregate of multiple electronic circuits. Additionally, memory b2 can be a magnetic disk, an optical disk, etc., or can be represented as a storage device, a recording medium, etc. In addition, memory b2 can be a non-volatile memory or a volatile memory.

[0819] For example, memory b2 can store an image or a stream. Additionally, memory b2 can store a program for causing processor b1 to decode the stream.

[0820] In addition, for example, memory b2 can act as Fig.67 the role of two or more of the elements for storing information among the various elements of the decoder 200 shown in etc. More specifically, memory b2 can act as Fig.67 the role of block memory 210 and frame memory 214 shown. More specifically, memory b2 can store a reconstructed image (specifically, a reconstructed block, a reconstructed picture, etc.).

[0821] Note that in decoder 200, not all of the elements of the various elements shown in Fig.67 can be implemented, and not all of the processes described herein can be executed. Fig.67 A part of the elements shown, etc., can be included in another device, or a part of the processes described herein can be executed by another device.

[0822] Hereinafter, the overall flow of the process executed by decoder 200 will be described, and then each element included in decoder 200 will be described. Note that some of the elements included in decoder 200 execute the same processes as some of those in encoder 100, so the same processes will not be described in detail again. For example, the inverse quantizer 204, inverse transformator 206, adder 208, block memory 210, frame memory 214, intra predictor 216, inter predictor 218, prediction controller 220, and loop filter 212 included in decoder 200 respectively execute processes similar to those executed by the inverse quantizer 112, inverse transformator 114, adder 116, block memory 118, frame memory 122, intra predictor 124, inter predictor 126, prediction controller 128, and loop filter 120 included in decoder 200.

[0823] (Overall flow of the decoding process)

[0824] Fig.69 is a flowchart showing an example of the overall decoding process performed by the decoder 200.

[0825] First, the segmentation determiner 224 in the decoder 200 determines the segmentation pattern of each of a plurality of fixed-size blocks (e.g., 128×128 pixels) included in the picture based on the parameters input from the entropy decoder 202 (step Sp_1). The segmentation pattern is the segmentation pattern selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks of the segmentation pattern.

[0826] The entropy decoder 202 decodes (specifically, entropy-decodes) the encoded and quantized coefficients and the prediction parameters of the current block (step Sp_2).

[0827] Next, the inverse quantizer 204 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 206 performs an inverse transform on the result to recover the prediction residual (i.e., the difference block) (step Sp_3).

[0828] Next, the prediction executor including all or part of the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220 generates a prediction signal for the current block (step Sp_4).

[0829] Next, the adder 208 adds the predicted image and the prediction residual to generate a reconstructed image of the current block (also referred to as a decoded image block) (step Sp_5).

[0830] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).

[0831] The decoder 200 then determines whether the decoding of the entire picture has ended (step Sp_7). When it is determined that the decoding has not ended (No in step Sp_7), the decoder 200 repeats the process starting from step Sp_1.

[0832] Note that the processes of these steps Sp_1 to Sp_7 can be sequentially performed by the decoder 200, or two or more processes can be performed in parallel. The processing order of two or more processes can be modified.

[0833] (Segmentation determiner)

[0834] Fig.70 is a conceptual diagram for showing the relationship between the segmentation determiner 224 and other components in the embodiment. As an example, the segmentation determiner 224 can perform the following processes.

[0835] For example, the segmentation determiner 224 collects block information from the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoder 202. Then, the segmentation determiner 224 can determine the segmentation pattern of the fixed-size block based on the block information and the parameters. The segmentation determiner 224 then outputs information indicating the determined segmentation pattern to the inverse transformer 206, the intra predictor 216, and the inter predictor 218. The inverse transformer 206 can perform inverse transformation of the transform coefficients based on the segmentation pattern indicated by the information from the segmentation determiner 224. The intra predictor 216 and the inter predictor 218 can generate a prediction image based on the segmentation pattern indicated by the information from the segmentation determiner 224.

[0836] (Entropy decoder)

[0837] Fig.71 is a block diagram showing an example of the functional configuration of the entropy decoder 202.

[0838] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the segmentation pattern by performing entropy decoding on the stream. For example, CABAC is used for entropy decoding. More specifically, the entropy decoding 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a de-binarizer 202c. The binary arithmetic decoder 202a arithmetically decodes the stream into a binary signal using the context value derived by the context controller 202b. The context controller 202b derives the context value according to the characteristics of the syntax element or the surrounding state (i.e., the occurrence probability of the binary signal) in the same manner as the context controller 110b on the encoder 100 side. The de-binarizer 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. This binarization can be performed according to the above-described binarization method.

[0839] In this way, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can output the prediction parameters included in the stream (see Figure 1 ) to the intra predictor 216, the inter predictor 218, and the prediction controller 220. The intra predictor 216, the inter predictor 218, and the prediction controller 220 can perform the same prediction process as the prediction process performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.

[0840] Fig.72 is a conceptual diagram of the flow showing an exemplary CABAC process in the entropy decoder 202.

[0841] First, initialization is performed in the CABAC in the entropy decoder 202. In the initialization, initialization in the binary arithmetic decoder 202a and setting of initial context values are performed. The binary arithmetic decoder 202a and the de-binarizer 202c then perform arithmetic decoding and de-binarization of the encoded data of, for example, a CTU. At this time, the context controller 202b updates the context value each time arithmetic decoding is performed. The context controller 202b then saves the context value as post-processing. For example, the saved context value is used to initialize the context value of the next CTU.

[0842] (Inverse Quantizer)

[0843] The inverse quantizer 204 inverse quantizes the quantized coefficients of the current block input from the entropy decoder 202. More specifically, the inverse quantizer 204 inverse quantizes the quantized coefficients of the current block based on the quantization parameter corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inverse quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.

[0844] Fig.73 is a block diagram showing an example of the functional configuration of the inverse quantizer 204.

[0845] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage device 204d, and an inverse quantization executor 204e.

[0846] Fig.74 is a flowchart showing an example of the inverse quantization process performed by the inverse quantizer 204.

[0847] The inverse quantizer 204 can perform an inverse quantization process for each CU based on Fig.74 the process shown as an example. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameter of the current block from the entropy decoder 202 (step Sv_12).

[0848] Next, the predicted quantization parameter generator 204b obtains the quantization parameter of a processing unit different from the current block from the quantization parameter storage device 204d (step Sv_13). The predicted quantization parameter generator 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).

[0849] The quantization parameter generator 204a then generates the quantization parameter for the current block based on the differential quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the differential quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b may be added together to generate the quantization parameter for the current block. In addition, the quantization parameter generator 204a stores the quantization parameter for the current block in the quantization parameter storage device 204d (step Sv_16).

[0850] Next, the inverse quantization executor 204e inverse quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).

[0851] Note that the differential quantization parameter can be decoded at the bit sequence level, picture level, slice level, tile level, or CTU level. Additionally, the initial value of the quantization parameter can be decoded at the sequence level, picture level, slice level, tile level, or CTU level. At this time, the initial value of the quantization parameter and the differential quantization parameter can be used to generate the quantization parameter.

[0852] Note that the inverse quantizer 204 may include multiple inverse quantizers, and the inverse quantization method selected from multiple inverse quantization methods may be used to inverse quantize the quantized coefficients.

[0853] (Inverse transformer)

[0854] The inverse transformer 206 restores the prediction residual by inverse-transforming the transform coefficients that are the input to the inverse quantizer 204.

[0855] For example, when the information parsed from the stream indicates that EMT or AMT is to be applied (e.g., when the AMT flag is true), the inverse transformer 206 inverse-transforms the transform coefficients of the current block based on the information indicating the parsed transform type.

[0856] In addition, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.

[0857] Fig.75 is a flowchart showing an example of the process performed by the inverse transformer 206.

[0858] For example, the inverse transform unit 206 determines whether information indicating that an orthogonal transform has not been performed exists in the stream (step St_11). Here, when it is determined that such information does not exist (No in step St_11) (for example: there is no indication regarding whether an orthogonal transform is performed; there is an indication that an orthogonal transform will be performed), the inverse transform unit 206 obtains information indicating the transform type decoded by the entropy decoder 202 (step St_12). Next, based on this information, the inverse transform unit 206 determines the transform type for the orthogonal transform in the encoder 100 (step St_13). The inverse transform unit 206 then performs an inverse orthogonal transform using the determined transform type (step St_14). As Fig.75 shown, when it is determined that information indicating that an orthogonal transform has not been performed exists (Yes in step St_11) (for example, an explicit indication that an orthogonal transform has not been performed; there is no indication to perform an orthogonal transform), the orthogonal transform is not performed.

[0859] Fig.76 is a flowchart showing an example of the process performed by the inverse transform unit 206.

[0860] For example, the inverse transform unit 206 determines whether the transform size is less than or equal to a determined value (step Su_11). The determined value can be predetermined. Here, when it is determined that the transform size is less than or equal to the determined value (Yes in step Su_11), the inverse transform unit 206 obtains from the entropy decoder 202 information indicating which transform type among at least one transform type included in the first transform type group the encoder 100 has used (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transform unit 206.

[0861] Based on this information, the inverse transform unit 206 determines the transform type for the orthogonal transform in the encoder 100 (step Su_13). The inverse transform unit 206 then performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). When it is determined that the transform size is not less than or equal to the determined value (No in step Su_11), the inverse transform unit 206 performs an inverse transform on the transform coefficients of the current block using the second transform type group (step Su_15).

[0862] Note that, as an example, the inverse orthogonal transform of the inverse transform unit 206 can be performed for each TU according to Fig.75 or Fig.76 shown in the flow. In addition, the inverse orthogonal transform can be performed by using a defined transform type without decoding the information indicating the transform type for the orthogonal transform. The defined transform type can be a predefined transform type or a default transform type. Additionally, the transform type can specifically be DST7, DCT8, etc. In the inverse orthogonal transform, the inverse transform basis function corresponding to the transform type is used.

[0863] (Adder)

[0864] The adder 208 reconstructs the current block by adding the prediction residual that is the input from the inverse transformer 206 and the prediction residual that is the input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.

[0865] (Block Memory)

[0866] The block memory 210 is a storage device for storing blocks included in the current picture and that can be referenced in intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.

[0867] (Loop Filter)

[0868] The loop filter 212 applies a loop filter to the reconstructed image generated by the adder 208, outputs the filtered reconstructed image to the frame memory 214, and provides the output of the decoder 200, for example, and outputs it to a display device or the like.

[0869] When information indicating the turning on or off of the ALF parsed from the stream indicates that the ALF is on, one filter can be selected from a plurality of filters, for example, based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.

[0870] Fig.77 is a block diagram showing an example of the functional configuration of the loop filter 212. Note that the configuration of the loop filter 212 is similar to the configuration of the loop filter 120 of the encoder 100.

[0871] For example, as Fig.77 shown, the loop filter 212 includes a deblocking filter executor 212a, an SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filter process on the reconstructed image. The SAO executor 212b performs an SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include Fig.77 all the constituent elements disclosed in Fig.77 and can include only a part of the constituent elements. In addition, the loop filter 212 can be configured to perform the above processes in a processing order different from the processing order disclosed in Fig.77 and can not perform all the processes shown in

[0872] (Frame Memory)

[0873] The frame memory 214 is, for example, a storage device for storing reference pictures used in inter-frame prediction, and may also be referred to as a frame buffer. More specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter 212.

[0874] (Predictor (intra predictor, inter predictor, prediction controller))

[0875] Fig.78 is a flowchart showing an example of the process executed by the predictor of the decoder 200. Note that the prediction executor may include all or part of the following constituent elements: the intra predictor 216; the inter predictor 218; and the prediction controller 220. The prediction executor includes, for example, the intra predictor 216 and the inter predictor 218.

[0876] The predictor generates a prediction image of the current block (step Sq_1). This prediction image is also referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra prediction signal or an inter prediction signal. More specifically, the predictor uses the reconstructed image that has been obtained for another block through the generation of the prediction image, the recovery of the prediction residual, and the addition of the prediction images to generate the prediction image of the current block. The predictor of the decoder 200 generates the same prediction image as the prediction image generated by the predictor of the encoder 100. In other words, the prediction image is generated according to a method common to or corresponding to each other between the predictors.

[0877] The reconstructed image may be, for example, an image in a reference picture, or an image of a decoded block (i.e., the above other block) in the current picture including the current block. The decoded block in the current picture is, for example, an adjacent block of the current block.

[0878] Fig.79 is a flowchart showing another example of the process executed by the predictor of the decoder 200.

[0879] The predictor determines a method or mode for generating the prediction image (step Sr_1). For example, the method or mode may be determined based on, for example, prediction parameters, etc.

[0880] When the first method is determined as the mode for generating the prediction image, the predictor generates the prediction image according to the first method (step Sr_2a). When the second method is determined as the mode for generating the prediction image, the predictor generates the prediction image according to the second method (step Sr_2b). When the third method is determined as the mode for generating the prediction image, the predictor generates the prediction image according to the third method (step Sr_2c).

[0881] The first method, the second method, and the third method may be mutually different methods for generating a predicted image. Each of the first to third methods may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.

[0882] FIG. 80A to FIG. 80C (Collectively referred to as FIG. 80) is a flowchart showing another example of the process executed by the predictor of the decoder 200.

[0883] As an example, the predictor may execute a prediction process according to the flow shown in FIG. 80. Note that the intra-block copy shown in FIG. 80 is a mode belonging to inter-frame prediction, and the block included in the current picture is referred to as a reference image or a reference block. In other words, in intra-block copy, a picture different from the current picture is not referred to. Additionally, the PCM mode shown in FIG. 80 is a mode belonging to intra-frame prediction, and transformation and quantization are not performed therein.

[0884] (Intra-frame predictor)

[0885] The intra-frame predictor 216 performs intra-frame prediction by referring to the block in the current picture stored in the block memory 210, based on the intra-frame prediction mode parsed from the stream, to generate a predicted image of the current block (i.e., the intra-frame prediction block). More specifically, the intra-frame predictor 216 performs intra-frame prediction by referring to the pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-frame prediction image, and then outputs the intra-frame prediction image to the prediction controller 220.

[0886] Note that when an intra-frame prediction mode in which the luminance block is referred to in the intra-frame prediction of the chrominance block is selected, the intra-frame predictor 216 may predict the chrominance component of the current block based on the luminance component of the current block.

[0887] In addition, when the information parsed from the stream indicates that PDPC is to be applied, the intra-frame predictor 216 corrects the pixel values of the intra-frame prediction based on the horizontal / vertical reference pixel gradients.

[0888] Fig.81 is a diagram showing an example of the process executed by the intra-frame predictor 216 of the decoder 200.

[0889] The intra-frame predictor 216 first determines whether to adopt MPM. As Fig.81As shown, the intra predictor 216 determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that an MPM flag indicating 1 exists (yes in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 among the MPMs from the entropy decoder 202. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode (step Sw_14), which is included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_12.

[0890] When it is determined that an MPM flag indicating 1 does not exist (no in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains from the entropy decoder 202 information indicating the intra prediction mode selected from at least one intra prediction mode that has never been included in the MPM in the encoder 100. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Then, the intra predictor 216 determines the intra prediction mode (step Sw_17), which is not included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15.

[0891] The intra predictor 216 generates a prediction image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).

[0892] (Inter - frame predictor)

[0893] The inter - frame predictor 218 predicts the current block by referring to the reference pictures stored in the frame memory 214. Prediction is performed in units of the current block or the current sub - block in the current block. Note that the sub - block is included in the block and is a unit smaller than the block. The size of the sub - block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub - block can be switched in units such as slices, bricks, pictures, etc.

[0894] For example, the inter - frame predictor 218 generates an inter - frame prediction image of the current block or the current sub - block by performing motion compensation using motion information (e.g., MV) parsed from the stream (e.g., prediction parameters output from the entropy decoder 202), and outputs the inter - frame prediction image to the prediction controller 220.

[0895] When the information parsed from the stream indicates that the OBMC mode is to be applied, in addition to the motion information of the current block obtained through motion estimation, the inter-frame predictor 218 also uses the motion information of adjacent blocks to generate an inter-frame predicted image.

[0896] In addition, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives motion information by performing motion estimation according to the mode matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter-frame predictor 218 then performs motion compensation (prediction) using the derived motion information.

[0897] In addition, when the BIO mode is to be applied, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. In addition, when the information parsed from the stream indicates that the affine mode is to be applied, the inter-frame predictor 218 derives the MV of each sub-block based on the MVs of multiple adjacent blocks.

[0898] (MV Derivation Process)

[0899] Fig.82 is a flowchart showing an example of the MV derivation process in the decoder 200.

[0900] For example, the inter-frame predictor 218 determines whether to decode motion information (e.g., MV). For example, the inter-frame predictor 218 can make the determination according to the prediction mode included in the stream, or can make the determination based on other information included in the stream. Here, when it is determined to decode motion information, the inter-frame predictor 218 derives the MV of the current block in the mode of decoding the motion information. When it is determined not to decode motion information, the inter-frame predictor 218 derives the MV in the mode of not decoding the motion information.

[0901] Here, the MV derivation modes include the conventional inter-frame mode, the conventional merge mode, the FRUC mode, the affine mode, etc., which will be described later. The modes of decoding motion information in the modes include the conventional inter-frame mode, the conventional merge mode, the affine mode (specifically, the affine inter-frame mode and the affine merge mode), etc. Note that the motion information can include not only the MV, but also the MV predictor selection information described later. The modes of not decoding motion information include the FRUC mode, etc. The inter-frame predictor 218 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.

[0902] Fig.83 is a flowchart showing an example of the process of MV derivation in the decoder 200.

[0903] For example, the inter - frame predictor 218 can determine whether to decode the MV difference, i.e., it can make a determination based on, for example, the prediction mode included in the stream, or it can make a determination based on other information included in the stream. Here, when it is determined to decode the MV difference, the inter - frame predictor 218 can derive the MV of the current block in the mode of decoding the MV difference. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.

[0904] When it is determined not to decode any MV difference, the inter - frame predictor 218 derives the MV in the mode of not decoding the MV difference. In this case, the encoded MV difference is not included in the stream.

[0905] Here, as described above, the MV derivation modes include the conventional inter - frame mode, the conventional merge mode, the FRUC mode, the affine mode, etc., which will be described later. The modes in which the MV difference is encoded include the conventional inter - frame mode and the affine mode (specifically, the affine inter - frame mode), etc. The modes in which the MV difference is not encoded include the FRUC mode, the conventional merge mode, the affine mode (specifically, the affine merge mode), etc. The inter - frame predictor 218 selects a mode for deriving the MV of the current block from multiple modes and uses the selected mode to derive the MV of the current block.

[0906] (MV Derivation > Conventional Inter - Frame Mode)

[0907] For example, when the information parsed from the stream indicates that the conventional inter - frame mode is to be applied, the inter - frame predictor 218 derives the MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.

[0908] Fig.84 is a flowchart showing an example of the process of inter - frame prediction by the conventional inter - frame mode in the decoder 200.

[0909] The inter - frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter - frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially around the current block (step Sg_11). In other words, the inter - frame predictor 218 generates an MV candidate list.

[0910] Next, the inter - frame predictor 218 extracts N (an integer of 2 or greater) MV candidates as motion vector predictor candidates (also referred to as MV predictor candidates) from the multiple MV candidates obtained in step Sg_11 according to the determined ranking in the priority order (step Sg_12). Note that the ranking in the priority order can be determined in advance for the corresponding N MV predictor candidates, and the ranking can be pre - determined.

[0911] Next, the inter-frame predictor 218 decodes the MV predictor selection information from the input stream, and uses the decoded MV predictor selection information to select one MV predictor candidate from among N MV predictor candidates as the MV predictor for the current block (step Sg_13).

[0912] Next, the inter-frame predictor 218 decodes the MV difference from the input stream, and derives the MV for the current block by adding the difference, which is the decoded MV difference, to the selected MV predictor (step Sg_14).

[0913] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference picture (step Sg_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed for each block among all the blocks in a slice, the inter-frame prediction of the slice using the normal inter-frame mode ends. For example, when the processes in steps Sg_11 to Sg_15 are performed for each block among all the blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode ends. Note that not all the blocks included in a slice undergo the processes in steps Sg_11 to Sg_15, and when some blocks undergo the processes, the inter-frame prediction of the slice using the normal inter-frame mode can end. This also applies to the picture in steps Sg_11 to Sg_15. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the normal inter-frame mode can end.

[0914] (MV Derivation > Normal Merge Mode)

[0915] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter-frame predictor 218 derives the MV and performs motion compensation (prediction) using the MV.

[0916] Fig.85 is a flowchart showing an example of the process of inter-frame prediction by the normal merge mode in the decoder 200.

[0917] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks temporally or spatially surrounding the current block (step Sh_11). In other words, the inter-frame predictor 218 generates an MV candidate list.

[0918] Next, the inter-frame predictor 218 selects one MV candidate from among the multiple MV candidates obtained in step Sh_11 and derives the MV for the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains the MV selection information included in the stream as a prediction parameter, and selects the MV candidate identified by the MV selection information as the MV for the current block.

[0919] Finally, the inter-frame predictor 218 generates a predicted image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed for each block among all the blocks in a slice, the inter-frame prediction of the slice using the regular merge mode ends. In addition, when the processes in steps Sh_11 to Sh_13 are performed for each block among all the blocks in a picture, the inter-frame prediction of the picture using the regular merge mode ends. Note that not all the blocks included in a slice go through the processes of steps Sh_11 to Sh_13, and when some blocks go through the processes, the inter-frame prediction of the slice using the regular merge mode can end. This also applies to the picture in steps Sh_11 to Sh_13. When the processes are performed for some blocks in a picture, the inter-frame prediction of the picture using the regular merge mode can end.

[0920] (MV derivation > FRUC mode)

[0921] For example, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives the MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is derived on the decoder 200 side without being signaled from the encoder 100 side. For example, the decoder 200 can derive the motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values in the current block.

[0922] Fig.86 is a flowchart showing an example of the process of inter-frame prediction by the FRUC mode in the decoder 200.

[0923] First, the inter-frame predictor 218 generates a list of MVs indicating decoded blocks adjacent to the current block spatially or temporally as MV candidates by referring to the MVs (this list is an MV candidate list and can also be used as an MV candidate list for a normal merge mode, for example (step Si_11)). Next, the best MV candidate is selected from among the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates the evaluation value of each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV of the current block (step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV of the current block. Additionally, for example, the MV of the current block is derived using pattern matching in the surrounding area of the position corresponding to the selected best MV candidate included in the reference picture. In other words, estimation using pattern matching and evaluation values in the reference picture can be performed in the surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces a better evaluation value, and the updated MV can be determined as the final MV of the current block. In an embodiment, the update to the MV that produces a better evaluation value may not be performed.

[0924] Finally, the inter-frame predictor 218 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all the blocks in a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all the blocks in a picture, the inter-frame prediction of the picture using the FRUC mode ends. Each sub-block can be processed similarly to the case of each block.

[0925] (MV Derivation > FRUC Mode)

[0926] For example, when the information parsed from the stream indicates that the affine merge mode is to be applied, the inter-frame predictor 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using the MV.

[0927] Fig.87 is a flowchart showing an example of the process of inter-frame prediction in the decoder 200 by the affine merge mode.

[0928] In the affine merge mode, first, the inter-frame predictor 218 derives the MV at the corresponding control points of the current block (step Sk_11). As Fig.46AAs shown, the control points are the upper left point and the upper right point of the current block, or as Fig.46B shown, are the upper left point, the upper right point, and the lower left point of the current block.

[0929] For example, when using the Figures 47A to 47C shown MV derivation method, as Fig.47A shown, the inter-frame predictor 218 checks the decoded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left) in this order and identifies the first valid block decoded according to the affine mode. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine mode to derive the MV at the control points. For example, when block A is identified and block A has two control points, as Fig.47B shown, the inter-frame predictor 218 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left and upper right corners of the decoded block including block A. In this way, the MV at each control point is derived.

[0930] Note that, as Fig.49A shown, when block A is identified and block A has two control points, the MV at three control points can be calculated, and as Fig.49B shown, when block A is identified and when block A has three control points, the MV at two control points can be calculated.

[0931] In addition, when the MV selection information is included in the stream as a prediction parameter, the inter-frame predictor 218 can use the MV selection information to derive the MV at each control point of the current block.

[0932] Next, the inter-frame predictor 218 performs motion compensation for each block included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV for each of the multiple sub-blocks as the affine MV (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and the decoded reference picture to perform motion compensation for the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all the sub-blocks included in the current block, the inter-frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate the predicted image of the current block.

[0933] Note that the above MV candidate list can be generated in step Sk_11. The MV candidate list can be, for example, a list including MV candidates derived using multiple MV derivation methods for each control point. The multiple MV derivation methods can be, for example FIG. 47A to FIG. 47C the MV derivation method shown in Fig.48A and Fig.48B the MV derivation method shown in Fig.49A and Fig.49B the MV derivation method shown in and any combination of other MV derivation methods.

[0934] Note that, except for the affine mode, the MV candidate list may include MV candidates in the mode that performs prediction in units of sub-blocks.

[0935] Note that, for example, an MV candidate list including MV candidates in the affine merge mode using two control points and the affine merge mode using three control points may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points may be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points may be generated.

[0936] (MV Derivation > Affine Inter Frame Mode)

[0937] For example, when the information parsed from the stream indicates that the affine inter frame mode will be applied, the inter frame predictor 218 derives the MV in the affine inter frame mode and performs motion compensation (prediction) using the MV.

[0938] Fig.88 is a flowchart showing an example of the process of inter frame prediction by the affine inter frame mode in the decoder 200.

[0939] In the affine inter frame mode, first, the inter frame predictor 218 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_11). The control points are the upper left point, the upper right point, and the lower left point of the current block, as Fig.46A or Fig.46B shown.

[0940] The inter frame predictor 218 obtains the MV predictor selection information included in the stream as a prediction parameter and uses the MV identified by the MV predictor selection information to derive the MV predictors at each control point of the current block. For example, when using Fig.48A and Fig.48B the MV derivation methods shown, the inter frame predictor 218 selects Fig.48A or Fig.48BThe motion vectors of the blocks identified by the MV predictor selection information in the coded blocks near the corresponding control points of the current block shown in are used to derive the motion vector predictors (v0, v1) or (v0, v1, v2) at the control points of the current block.

[0941] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block and the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point of the current block is derived.

[0942] Next, the inter-frame predictor 218 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 calculates the MV of each of the multiple sub-blocks as an affine MV using two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) (step Sj_13). The inter-frame predictor 218 then uses these affine MVs and the decoded reference picture to perform motion compensation on the sub-blocks (step Sj_14). When the processes in steps Sj_13 and Sj_14 are performed for each sub-block included in the current block, the inter-frame prediction using the affine merge mode of the current block ends. In other words, motion compensation of the current block is performed to generate a predicted image of the current block.

[0943] Note that the above MV candidate list can be generated in step Sj_11 as in step Sk_11.

[0944] (MV Derivation > Triangle Mode)

[0945] For example, when the information parsed from the stream indicates that the triangle mode will be applied, the inter-frame predictor 218 derives the MV in the triangle mode and performs motion compensation (prediction) using the MV.

[0946] Fig.89 is a flowchart showing an example of the process of inter-frame prediction in the triangle mode in the decoder 200.

[0947] In the triangle mode, first, the inter-frame predictor 218 divides the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as a prediction parameter, which is information related to the division. The inter-frame predictor 218 can then divide the current block into a first partition and a second partition according to the partition information.

[0948] Next, the inter-frame predictor 218 obtains a plurality of MV candidates for the current block based on information such as the MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). In other words, the inter-frame predictor 218 generates a list of MV candidates.

[0949] The inter-frame predictor 218 then respectively selects an MV candidate for the first partition and an MV candidate for the second partition as the first MV and the second MV from the plurality of MV candidates obtained in step Sx_11 (step Sx_13). At this time, the inter-frame predictor 218 may obtain MV selection information from the stream for identifying each selected MV candidate as a prediction parameter. The inter-frame predictor 218 can then select the first MV and the second MV according to the MV selection information.

[0950] Next, the inter-frame predictor 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame predictor 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).

[0951] Finally, the inter-frame predictor 218 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_16).

[0952] (MV Estimation > DMVR)

[0953] For example, if the information parsed from the stream indicates that DMVR is to be applied, the inter-frame predictor 218 performs motion estimation using DMVR.

[0954] Fig.90 is a flowchart showing an example of the process of motion estimation performed by DMVR in the decoder 200.

[0955] The inter-frame predictor 218 derives the MV of the current block according to the merge mode (step S1_11). Next, the inter-frame predictor 218 derives the final MV of the current block by searching the area around the reference picture indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined according to DMVR.

[0956] Fig.91 is a flowchart showing an example of the motion estimation process performed by DMVR in the decoder 200 and is the same as Fig.58B the same.

[0957] First, in Fig.58AIn step 1 as shown, the inter-frame predictor 218 calculates the cost between the search position (also known as the starting point) indicated by the initial MV and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 changes the target to the search position that obtains the minimum cost and executes the process in step 2 shown in FIG. 58. When the cost at the starting point is the minimum, the inter-frame predictor 218 skips Fig.58A the process in step 2 shown in

[0958] In step 2 as shown in Fig.58A , the inter-frame predictor 218 performs a search similar to the process in step 1, and regards the search position after the target change as the new starting point according to the result of the process in step 1. Then, the inter-frame predictor 218 determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 executes the process in step 4. When the cost at the starting point is the minimum, the inter-frame predictor 218 executes the process in step 3.

[0959] In step 4, the inter-frame predictor 218 regards the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0960] In Fig.58A step 3 as shown, the inter-frame predictor 218 determines the pixel position at the sub-pixel accuracy that obtains the minimum cost based on the costs at four points located at the upper, lower, left, and right positions relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position.

[0961] The pixel position at the sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), (1, 0)) of up, down, left, and right using the cost at the corresponding search position among the four search positions as the weight. The inter-frame predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0962] (Motion Compensation > BIO / OBMC / LIC)

[0963] For example, when the information parsed from the stream indicates that the predicted image needs to be corrected, when the predicted image is generated, the inter-frame predictor 218 corrects the predicted image based on the correction mode. This mode is, for example, one of the above-mentioned BIO, OBMC, and LIC.

[0964] Fig.92It is a flowchart showing an example of the process of generating a predicted image in the decoder 200.

[0965] The inter-frame predictor 218 generates a predicted image (step Sm_11) and corrects the predicted image according to any of the above patterns (step Sm_12).

[0966] Fig.93 It is a flowchart showing another example of the process of generating a predicted image in the decoder 200.

[0967] The inter-frame predictor 218 derives the MV of the current block (step Sn_11). Next, the inter-frame predictor 218 generates a predicted image using the MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, the inter-frame predictor 218 obtains the prediction parameters included in the stream and determines whether to perform a correction process based on the prediction parameters. For example, the prediction parameter is a flag indicating whether one or more of the above patterns are to be applied. Here, when it is determined to perform the correction process (Yes in step Sn_13), the inter-frame predictor 218 generates a final predicted image...

Claims

1. A non - transitory computer - readable medium storing a bitstream, the bitstream including information according to which a decoder performs a decoding process, the information including a first flag and a second flag, the first flag indicating whether a CCALF (Cross - Component Adaptive Loop Filtering) process is enabled for a first block that is adjacent to the left side of a current block, and the second flag indicating whether a CCALF process is enabled for a second block that is adjacent to the upper side of the current block, wherein, During the decoding process: Parse the first flag from the bitstream indicating whether the CCALF process is enabled for the first block; Parse the second flag from the bitstream indicating whether the CCALF process is enabled for the second block; Determine a first index associated with a color component of the current block; Derive a second index indicating a context model using an equation that is a function of the first flag, the second flag, and the first index, the equation being used to derive a third index indicating another context model for an ALF (Adaptive Loop Filter) control flag; Perform entropy decoding on a third flag using the context model indicated by the second index, the third flag indicating whether the CCALF process is enabled for the current block, and In response to the third flag indicating that the CCALF process is enabled for the current block, perform the CCALF process on the current block.