System and method for video decoding
By deciding to use luminance samples to predict chrominance sample blocks according to the splitting of virtual pipeline decoding units in video decoding, the encoding and decoding processes are optimized, the problems of coding efficiency and image quality are solved, the circuit scale and resource utilization are reduced, and the processing speed is improved.
Patent Information
- Application Number
- CN202080044292.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-06-18
AI Technical Summary
The existing video decoding technology has room for improvement in coding efficiency and image quality when processing chrominance samples, and the circuit scale is large, making it difficult to meet the demand for increasing amounts of digital video data.
By deciding whether to use luminance samples to predict chrominance sample blocks based on whether the virtual pipeline decoding unit is split into smaller blocks during the video decoding process, and adopting different encoding and decoding methods, including intra-frame prediction, inter-frame prediction, loop filtering, transform and entropy coding, the encoding and decoding process is optimized.
The coding efficiency is improved, the image quality is enhanced, the utilization of processing resources and the circuit scale are reduced, and the processing speed of encoding/decoding is increased.
Smart Images

Figure CN114026862B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding, and in particular to video encoding and decoding systems, components, and methods in video coding and decoding, for example, for performing encoding of a block using predicted chroma samples. Background Art
[0002] As video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Codec), there remains a continuing need to provide improvements and optimizations to video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in performing coding of blocks using predicted chroma samples. Summary of the Invention
[0003] In one aspect, an encoder includes: circuitry; and a memory coupled to the circuitry. The circuitry determines whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, a block of chroma samples is predicted without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, a block of chroma samples is predicted using luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, a block of chroma samples is predicted using luma samples. A block is encoded using the predicted chroma samples.
[0004] In one aspect, an encoder includes: a block splitter that splits a first image into a plurality of blocks in operation; an intra predictor that predicts a block in the first image using a reference block included in the first image in operation; an inter predictor that predicts a block in the first image using a reference block included in a second image different from the first image in operation; a loop filter that filters the block in the first image in operation; a transformer that transforms a prediction error between an original signal and a prediction signal generated by the intra predictor or the inter predictor in operation to generate a transform coefficient; a quantizer that quantizes the transform coefficient in operation to generate a quantized coefficient; and an entropy encoder that variably encodes the quantized coefficient in operation to generate an encoded bitstream, the encoded bitstream including the encoded quantized coefficient and control information. Predicting a block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the chroma sample block without using the luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the chroma sample block using the luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting the chroma sample block using the luma samples. Encoding the block using the predicted chroma samples.
[0005] In one aspect, a decoder includes: circuitry; and a memory coupled to the circuitry. The circuitry determines whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, a block of chroma samples is predicted without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, a block of chroma samples is predicted using luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, a block of chroma samples is predicted using luma samples. The block is decoded using the predicted chroma samples.
[0006] In one aspect, a decoding device includes: a decoder that decodes an encoded bitstream to output quantized coefficients; an inverse quantizer that inversely quantizes the quantized coefficients to output transform coefficients; an inverse transformer that inversely transforms the transform coefficients to output prediction errors; an intra-frame predictor that uses a reference block included in a first image to predict a block included in the first image; an inter-frame predictor that uses a reference block included in a second image different from the first image to predict a block included in the first image; a loop filter that filters the block included in the first image; and an output that outputs a picture including the first image. Predicting a block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chrominance samples without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, using luma samples to predict the block of chroma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, using luma samples to predict the block of chroma samples. The block is decoded using the predicted chroma samples.
[0007] In one aspect, an encoding method includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chroma samples without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chroma samples using luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting a block of chroma samples using luma samples. Encoding a block using the predicted chroma samples.
[0008] In one aspect, a decoding method includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chroma samples without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chroma samples using luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting a block of chroma samples using luma samples. Decoding a block using the predicted chroma samples.
[0009] In video decoding technology, it is desirable to propose new methods to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of the embodiments of the present disclosure, including the constituent elements of the embodiments of the present disclosure considered individually or in various combinations, can facilitate one or more of the following: improving coding efficiency; enhancing image quality; reducing the utilization of processing resources associated with encoding / decoding; reducing circuit size; improving encoding / decoding processing speed, etc.
[0010] Additionally, some implementations of the embodiments of the present disclosure, including constituent elements of the embodiments of the present disclosure considered individually or in various combinations, can facilitate appropriate selection of one or more elements (e.g., filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations) in encoding and decoding. Note that the present disclosure includes disclosures about configurations and methods that can provide advantages in addition to those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing increased use of processing resources.
[0011] Additional benefits and advantages of the disclosed embodiments will become apparent from the description and drawings. Benefits and / or advantages can be obtained individually from the various embodiments and features of the description and drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.
[0012] It should be noted that the general or specific embodiments may be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram showing one example of a functional configuration of a transmission system according to the embodiment.
[0014] Figure 2 is a conceptual diagram showing an example of a hierarchical structure of data in a stream.
[0015] Figure 3 is a conceptual diagram illustrating an example of a slice configuration.
[0016] Figure 4 is a conceptual diagram illustrating an example of a tile configuration.
[0017] Figure 5 is a conceptual diagram for illustrating an example of a coding structure in scalable coding.
[0018] Figure 6 is a conceptual diagram for illustrating an example of a coding structure in scalable coding.
[0019] Figure 7 is a block diagram showing a functional configuration of an encoder according to an embodiment.
[0020] Figure 8 is a functional block diagram showing an example of installation of an encoder.
[0021] Figure 9 is a flowchart indicating one example of an overall encoding process performed by an encoder.
[0022] Figure 10 is a conceptual diagram for illustrating an example of block splitting.
[0023] Figure 11 is a block diagram showing one example of a functional configuration of a splitter according to the embodiment.
[0024] Figure 12 is a conceptual diagram for illustrating an example of a split mode.
[0025] Figure 13A is a conceptual diagram showing an example of a syntax tree of a split pattern.
[0026] Figure 13B is a conceptual diagram for illustrating another example of a syntax tree of a split pattern.
[0027] Figure 14 is a chart indicating example transform basis functions for various transform types.
[0028] Figure 15 is a conceptual diagram for illustrating an example space-varying transform (SVT).
[0029] Figure 16 is a flowchart illustrating one example of a process performed by a converter.
[0030] Figure 17 is a flowchart illustrating another example of a process performed by the converter.
[0031] Figure 18 is a block diagram showing one example of a functional configuration of a quantizer according to an embodiment.
[0032] Figure 19 is a flowchart illustrating one example of a quantization process performed by a quantizer.
[0033] Figure 20 is a block diagram showing one example of a functional configuration of an entropy encoder according to an embodiment.
[0034] Figure 21 is a conceptual diagram for illustrating an example flow of a Context-Based Adaptive Binary Arithmetic Coding (CABAC) process in an entropy encoder.
[0035] Figure 22 is a block diagram showing one example of a functional configuration of a loop filter according to the embodiment.
[0036] Figure 23A is a conceptual diagram for illustrating an example of a filter shape used in an adaptive loop filter (ALF).
[0037] Figure 23B is a conceptual diagram for illustrating another example of a filter shape used in ALF.
[0038] Figure 23C is a conceptual diagram for illustrating another example of a filter shape used in ALF.
[0039] Figure 23D is a conceptual diagram for illustrating an example flow of cross-component ALF (CC-ALF).
[0040] Figure 23E is a conceptual diagram for illustrating an example of a filter shape used in CC-ALF.
[0041] Figure 23F is a conceptual diagram for illustrating an example flow of Joint Chroma CCALF (JC-CCALF).
[0042] Figure 23G is a table showing example weighted index candidates that can be employed in JC-CCALF.
[0043] Figure 24 is a block diagram indicating one example of a specific configuration of a loop filter used as a deblocking filter (DBF).
[0044] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having symmetric filtering characteristics with respect to a block boundary.
[0045] Figure 26is a conceptual diagram for illustrating a block boundary on which a deblocking filtering process is performed.
[0046] Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0047] Figure 28 is a flowchart illustrating one example of a process performed by a predictor of an encoder.
[0048] Figure 29 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0049] Figure 30 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0050] Figure 31 : is a conceptual diagram for illustrating sixty-seven intra prediction modes used in intra prediction in the embodiment.
[0051] Figure 32 is a flowchart illustrating one example of a process performed by an intra predictor.
[0052] Figure 33 is a conceptual diagram for illustrating an example of a reference picture.
[0053] Figure 34 is a conceptual diagram illustrating an example of a reference picture list.
[0054] Figure 35 is a flowchart illustrating an example basic processing flow for inter-frame prediction.
[0055] Figure 36 is a flowchart showing an example of a derivation process of a motion vector.
[0056] Figure 37 is a flowchart illustrating another example of a derivation process of a motion vector.
[0057] Figure 38A is a conceptual diagram for illustrating an example representation of a pattern for MV derivation.
[0058] Figure 38B is a conceptual diagram for illustrating an example representation of a pattern for MV derivation.
[0059] Figure 39 is a flowchart illustrating an example of a process of inter prediction in normal inter mode.
[0060] Figure 40 is a flowchart illustrating an example of a process of inter prediction in normal merge mode.
[0061] Figure 41 is a conceptual diagram for illustrating an example of a motion vector derivation process in merge mode.
[0062] Figure 42 is a conceptual diagram for illustrating an example of an MV derivation process for a current picture through the HMVP merge mode.
[0063] Figure 43 is a flow chart illustrating one example of a frame rate up-conversion (FRUC) process.
[0064] Figure 44 is a conceptual diagram for illustrating one example of pattern matching (bilateral matching) between two blocks along a motion trajectory.
[0065] Figure 45 is a conceptual diagram for illustrating one example of pattern matching (template matching) between a template in a current picture and a block in a reference picture.
[0066] Figure 46A is a conceptual diagram for illustrating an example of deriving a motion vector for each subblock based on motion vectors of a plurality of neighboring blocks.
[0067] Figure 46B is a conceptual diagram for illustrating an example of deriving a motion vector for each subblock in an affine mode in which three control points are used.
[0068] Figure 47A is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0069] Figure 47B is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0070] Figure 47C is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0071] Figure 48A is a conceptual diagram for illustrating an affine mode in which two control points are used.
[0072] Figure 48B is a conceptual diagram for illustrating an affine mode in which three control points are used.
[0073] Figure 49A is a conceptual diagram for illustrating one example of a method for MV derivation at a control point when the number of control points for an encoded block and the number of control points for a current block are different from each other.
[0074] Figure 49Bis a conceptual diagram for illustrating another example of a method for MV derivation at a control point when the number of control points for an encoded block and the number of control points for a current block are different from each other.
[0075] Figure 50 is a flowchart showing one example of a process in affine merge mode.
[0076] Figure 51 is a flowchart showing one example of a process in affine inter mode.
[0077] Figure 52A is a conceptual diagram for illustrating the generation of two triangular prediction images.
[0078] Figure 52B is a conceptual diagram illustrating an example of a first portion of a first partition overlapping with a second partition, and first and second sets of samples that may be weighted as part of a correction process.
[0079] Figure 52C is a conceptual diagram for illustrating a first portion of a first partition, which is a portion of the first partition overlapping with a portion of an adjacent partition.
[0080] Figure 53 is a flowchart illustrating one example of a process in triangle mode.
[0081] Figure 54 is a conceptual diagram for illustrating one example of an advanced temporal motion vector prediction (ATMVP) mode in which an MV is derived in units of subblocks.
[0082] Figure 55 is a flow chart illustrating the relationship between merge mode and dynamic motion vector refresh (DMVR).
[0083] Figure 56 is a conceptual diagram for illustrating an example of DMVR.
[0084] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining MV.
[0085] Figure 58A is a conceptual diagram for illustrating an example of motion estimation in DMVR.
[0086] Figure 58B is a flowchart illustrating one example of a process of motion estimation in DMVR.
[0087] Figure 59 is a flowchart illustrating an example of a process of generating a predicted image.
[0088] Figure 60 is a flowchart illustrating another example of the process of generating a predicted image.
[0089] Figure 61 is a flowchart illustrating an example of a process of correcting a predicted image through overlapped block motion compensation (OBMC).
[0090] Figure 62 is a conceptual diagram for illustrating an example of a predicted image correction process by OBMC.
[0091] Figure 63 This is a conceptual diagram illustrating a model assuming uniform linear motion.
[0092] Figure 64 is a flowchart illustrating one example of a process of inter-frame prediction according to BIO.
[0093] Figure 65 is a functional block diagram illustrating one example of a functional configuration of an inter-frame predictor that can perform inter-frame prediction according to BIO.
[0094] Figure 66A is a conceptual diagram for illustrating one example of a process of a predicted image generation method using a luminance correction process performed by the LIC.
[0095] Figure 66B is a flowchart showing one example of the process of a predicted image generation method using LIC.
[0096] Figure 67 is a block diagram showing a functional configuration of a decoder according to an embodiment.
[0097] Figure 68 is a functional block diagram showing an example of installation of a decoder.
[0098] Figure 69 is a flow chart illustrating one example of the overall decoding process performed by a decoder.
[0099] Figure 70 is a conceptual diagram for illustrating the relationship between a split determiner and other constituent elements.
[0100] Figure 71 is a block diagram showing one example of a functional configuration of an entropy decoder.
[0101] Figure 72 is a conceptual diagram for illustrating an example flow of a CABAC process in an entropy decoder.
[0102] Figure 73 is a block diagram showing one example of a functional configuration of an inverse quantizer.
[0103] Figure 74 is a flowchart illustrating one example of a process of inverse quantization performed by an inverse quantizer.
[0104] Figure 75 is a flowchart showing one example of a process performed by an inverse converter.
[0105] Figure 76 is a flowchart illustrating another example of a process performed by an inverse converter.
[0106] Figure 77 is a block diagram showing one example of a functional configuration of a loop filter.
[0107] Figure 78 is a flowchart illustrating one example of a process performed by a predictor of a decoder.
[0108] Figure 79 is a flow chart illustrating another example of a process performed by a predictor of a decoder.
[0109] Figure 80 is a flow chart illustrating another example of a process performed by a predictor of a decoder.
[0110] Figure 81 is a diagram showing one example of a process performed by an intra predictor of a decoder.
[0111] Figure 82 is a flowchart showing an example of a process of MV derivation in a decoder.
[0112] Figure 83 is a flowchart illustrating another example of a process of MV derivation in a decoder.
[0113] Figure 84 is a flowchart illustrating an example of a process of inter prediction by normal inter mode in a decoder.
[0114] Figure 85 is a flowchart illustrating an example of a process for inter prediction through normal merge mode in a decoder.
[0115] Figure 86 is a flowchart illustrating an example of a process of inter-frame prediction by FRUC mode in a decoder.
[0116] Figure 87 is a flowchart illustrating an example of a process for inter prediction by affine merge mode in a decoder.
[0117] Figure 88 is a flowchart illustrating an example of a process of inter prediction by affine inter mode in a decoder.
[0118] Figure 89 is a flow chart illustrating an example of a process for inter-frame prediction by triangular mode in a decoder.
[0119] Figure 90 is a flow chart illustrating an example of a process of motion estimation by DMVR in a decoder.
[0120] Figure 91 is a flow chart illustrating an example process for motion estimation by DMVR in a decoder.
[0121] Figure 92 is a flowchart showing an example of a process of generating a prediction image in a decoder.
[0122] Figure 93 is a flowchart illustrating another example of a process of generating a prediction image in a decoder.
[0123] Figure 94 is a flowchart illustrating an example of a process of correcting a predicted image by OBMC in a decoder.
[0124] Figure 95 is a flowchart illustrating an example of a process of correcting a predicted image through BIO in a decoder.
[0125] Figure 96 is a flowchart illustrating an example of a process of correcting a predicted image by LIC in a decoder.
[0126] Figure 97 is a flow chart illustrating an example of a process for decoding a block using predicted chroma samples.
[0127] Figure 98 is a flow chart illustrating an example of a process for decoding a block using predicted chroma samples.
[0128] Figure 99 is a conceptual diagram illustrating an example of determining whether a current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0129] Figure 100 is a conceptual diagram illustrating an example of determining whether a current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0130] Figure 101 is a conceptual diagram illustrating a virtual pipeline decoding unit (VPDU).
[0131] Figure 102 is a conceptual diagram for illustrating an example of determining whether a current VPDU can use luma samples to predict a block of chroma samples.
[0132] Figure 103 is a conceptual diagram illustrating an example approach to determining whether a luma VPDU is to be split into smaller blocks.
[0133] Figure 104 is a conceptual diagram illustrating additional considerations that may be taken into account to determine whether to use luma samples to predict chroma samples of a block.
[0134] Figure 105 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0135] Figure 106 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0136] Figure 107 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0137] Figure 108 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0138] Figure 109 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0139] Figure 110 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0140] Figure 111 is a conceptual diagram for illustrating an example of a partition having a non-rectangular shape.
[0141] Figure 112 is a diagram showing an example overall configuration of a content providing system for realizing a content distribution service.
[0142] Figure 113 is a conceptual diagram illustrating an example of a display screen for showing a web page.
[0143] Figure 114 is a conceptual diagram illustrating an example of a display screen for showing a web page.
[0144] Figure 115 is a block diagram illustrating one example of a smartphone.
[0145] Figure 116 is a block diagram illustrating an example of a functional configuration of a smartphone. DETAILED DESCRIPTION
[0146] In the drawings, like reference numerals denote like elements unless context dictates otherwise. The sizes and relative positions of elements in the drawings are not necessarily drawn to scale.
[0147] Hereinafter, (multiple) embodiments will be described with reference to the accompanying drawings. Note that the (multiple) embodiments described below each illustrate general or specific examples. The numerical values, shapes, materials, components, arrangement and connection of components, steps, relationships and order of steps, etc. indicated in the following (multiple) embodiments are merely examples and are not intended to limit the scope of the claims.
[0148] Embodiments of encoders and decoders are described below. The embodiments are examples of encoders and decoders to which the processes and / or configurations presented in the description of aspects of the present disclosure apply. The processes and / or configurations may also be implemented in encoders and decoders different from the encoders and decoders according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented:
[0149] (1) Any one of the components of the encoder or decoder according to the embodiment presented in the description of the aspects of the present disclosure may be replaced by or combined with another component presented anywhere in the description of the aspects of the present disclosure.
[0150] (2) In an encoder or decoder according to an embodiment, any changes may be made to the functions or processes performed by one or more components of the encoder or decoder, such as addition, replacement, removal, etc. of the functions or processes. For example, any function or process may be replaced by or combined with another function or process presented anywhere in the description of aspects of this disclosure.
[0151] (3) In the method implemented by the encoder or decoder according to the embodiment, any changes may be made, such as addition, replacement, and removal of one or more processes included in the method. For example, any process in the method may be replaced by or combined with another process presented anywhere in the description of the aspects of this disclosure.
[0152] (4) One or more components included in an encoder or decoder according to an embodiment may be combined with components presented anywhere in the description of aspects of this disclosure, may be combined with components including one or more functions presented anywhere in the description of aspects of this disclosure, and may be combined with components that implement one or more processes implemented by components presented in the description of aspects of this disclosure.
[0153] (5) A component including one or more functions of an encoder or decoder according to an embodiment, or a component implementing one or more processes of an encoder or decoder according to an embodiment, may be combined with or replaced by a component presented anywhere in the description of an aspect of the present disclosure, a component including one or more functions presented anywhere in the description of an aspect of the present disclosure, or a component implementing one or more processes presented anywhere in the description of an aspect of the present disclosure.
[0154] (6) In a method implemented by an encoder or decoder according to an embodiment, any one of the processes included in the method may be replaced or combined with a process presented anywhere in the description of aspects of this disclosure or by any corresponding or equivalent process.
[0155] (7) One or more processes included in the method implemented by the encoder or decoder according to the embodiment may be combined with the processes presented anywhere in the description of the aspects of the present disclosure.
[0156] (8) The implementation of the processes and / or configurations presented in the description of aspects of the present disclosure is not limited to the encoder or decoder according to the embodiments. For example, the processes and / or configurations may be implemented in a device for a purpose different from that of the moving picture encoder or moving picture decoder disclosed in the embodiments.
[0157] (Definition of terms)
[0158] The corresponding terms may be defined as indicated below as an example.
[0159] An image is a data unit configured with a set of pixels, a picture, or includes blocks smaller than pixels. Images include not only videos but also still images.
[0160] A picture is an image processing unit configured with a set of pixels and may also be referred to as a frame or field. For example, a picture may take the form of a luma sample array in a monochrome format or a luma sample array and two corresponding chroma sample arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0161] A block is a processing unit that is a collection of a certain number of pixels. A block can have any number of different shapes. For example, a block can have a rectangular shape of M×N (M columns×N rows) pixels, a square shape of M×M pixels, a triangular shape, a circular shape, and the like. Examples of blocks include slices, tiles, bricks, CTUs, super blocks, basic split units, VPDUs, processing split units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and sub-blocks. A block can take the form of an M×N array of samples or an M×N array of transform coefficients. For example, a block can be a square or rectangular pixel area that includes a luminance matrix and two chrominance matrices.
[0162] A pixel or sample is the smallest point of an image. A pixel or sample includes pixels at integer positions and pixels at sub-pixel positions, for example, pixels generated based on pixels at integer positions.
[0163] A pixel value or sample value is a characteristic value of a pixel and may include one or more of a brightness value, a chroma value, an RGB gradient level, a depth value, a binary value of zero or one, and the like.
[0164] Chroma or chrominance is the intensity of a color, typically represented by the symbols Cb and Cr, which specify that the value of an array of samples or a single sample value represents the value of one of two color difference signals associated with a primary color.
[0165] Luminance or luminance is the brightness of an image and is typically represented by the symbol or subscript Y or L, which specifies that the value of an array of samples or a single sample value represents the value of a monochromatic signal associated with a primary color.
[0166] A flag comprises one or more bits that indicate the value of, for example, a parameter or an index. A flag may be a binary flag, which indicates a binary value of the flag, or a non-binary value of a parameter.
[0167] Signals convey information, which is symbolized by or encoded into the signal. Signals include discrete digital signals and continuous analog signals.
[0168] A stream or bitstream is a sequence of digital data. It can be a single stream or multiple streams with multiple hierarchical layers. It can be sent in serial communication using a single transmission path or in packet communication using multiple transmission paths.
[0169] Difference refers to various mathematical differences, for example, simple difference (xy), absolute value of difference (|xy|), square difference (x^2-y^2), square root of difference (√(xy)), weighted difference (ax-by: a and b are constants), offset difference (x-y+a: a is an offset), etc. In the case of a scalar, simple difference may be sufficient, and difference calculations are included.
[0170] The sum refers to various mathematical sums, for example, a simple sum (x+y), the absolute value of the sum (|x+y|), a square sum (x^2+y^2), a square root of the sum (√(x+y)), a weighted difference (ax+by: a and b are constants), an offset sum (x+y+a: a is an offset), etc. In the case of scalars, a simple sum may be sufficient, and the sum calculation is included.
[0171] A frame is a combination of a top field and a bottom field, where sample rows 0, 2, 4, ... originate from the top field and sample rows 1, 3, 5, ... originate from the bottom field.
[0172] A slice is an integer number of coding tree units contained in an independent slice segment and all subsequent dependent slice segments (if any) before the next independent slice segment (if any) within the same access unit.
[0173] A tile is a rectangular area of coding tree blocks within a particular tile column and a particular tile row in a picture. A tile can be a rectangular area of a frame that is intended to be independently decodable and coded, but loop filtering across tile edges can still be applied.
[0174] A coding tree unit (CTU) can be a coding tree block of luma samples for a picture with three sample arrays, or two corresponding coding tree blocks of chroma samples. Alternatively, a CTU can be a coding tree block of samples for one of a monochrome picture and a picture coded using three separate color planes and syntax structures for coding the samples. A super block can be a square block of 64×64 pixels consisting of one or two mode information blocks, or recursively divided into four 32×32 blocks, which themselves can be further divided.
[0175] (System Configuration)
[0176] First, a transmission system according to the embodiment will be described. Figure 1 is a schematic diagram showing one example of the configuration of a transmission system 400 according to the embodiment.
[0177] The transmission system 400 is a system for transmitting a stream generated by encoding an image and decoding the transmitted stream. As shown, the transmission system 400 includes Figure 1 The encoder 100, network 300 and decoder 200 are shown in FIG.
[0178] An image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image and outputs the stream to the network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by encoding.
[0179] It should be noted that the image before being encoded by the encoder 100 is also referred to as the original image, original signal or original sample. The image can be a video or a still image. The image is a general concept of a sequence, a picture and a block, so unless otherwise specified, the image is not limited to a spatial area of a specific size and a time area of a specific size. The image is an array of pixels or pixel values, and the signal representing the image or pixel values is also referred to as a sample. The stream can be referred to as a bit stream, an encoded bit stream, a compressed bit stream or an encoded signal. In addition, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, an image encoding method or a video encoding method.
[0180] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 200 may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasting and satellite broadcasting. Alternatively, the network 300 may be replaced by a recording medium (e.g., a digital versatile disc (DVD) and a Blu-ray disc (BD)) on which the stream is recorded.
[0181] The decoder 200 generates a decoded image that is an uncompressed image by, for example, decoding a stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method adopted by the encoder 100.
[0182] It should be noted that the decoder 200 may also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 may also be referred to as a decoding method, an image decoding method, or a video decoding method.
[0183] (Data Structure)
[0184] Figure 2 is a conceptual diagram showing an example of a hierarchical structure of data in a stream. Figure 1 The transmission system 400 is described Figure 2 The stream includes, for example, a video sequence. Figure 2As shown in (a), a video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplemental enhancement information (SEI) and multiple pictures.
[0185] In a video having a plurality of layers, the VPS may include coding parameters common between some of the plurality of layers, and coding parameters related to some of the plurality of layers included in the video or to a single layer.
[0186] The SPS includes parameters for the sequence, that is, decoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the decoding parameters may indicate the width or height of the picture. It should be noted that there may be multiple SPSs.
[0187] The PPS includes parameters for a picture, i.e., decoding parameters that the decoder 200 refers to in order to decode each picture in a sequence. For example, the decoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. It should be noted that there may be multiple PPSs. Each of the SPS and PPS may be referred to simply as a parameter set.
[0188] like Figure 2 As shown in (b) of FIG. 1 , a picture may include a picture header and one or more slices. The picture header includes decoding parameters, and the decoder 200 refers to the decoding parameters to decode the one or more slices.
[0189] like Figure 2 As shown in (c) of FIG. 5 , a slice includes a slice header and one or more bricks. The slice header includes decoding parameters, and the decoder 200 refers to the decoding parameters to decode the one or more bricks.
[0190] like Figure 2 As shown in (d), a brick includes one or more coding tree units (CTUs).
[0191] It should be noted that a picture may not include any slices and may include slice groups instead of slices. In this case, a slice group includes at least one slice. Additionally, a tile may include a slice.
[0192] CTU is also called super block or basic split unit. Figure 2 As shown in (e) of FIG. 1 , a CTU includes a CTU header and at least one coding unit (CU). As shown, the CTU includes four coding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes decoding parameters, and the decoder 200 refers to the decoding parameters to decode at least one CU.
[0193] A CU can be split into multiple smaller CUs. As shown, CU (10) is not split into smaller decoding units; CU (11) is split into four smaller decoding units CU (110), CU (111), CU (112), and CU (113); CU (12) is not split into smaller decoding units; and CU (13) is split into seven smaller decoding units CU (1310), CU (1311), CU (1312), CU (1313), CU (132), CU (133), and CU (134). Figure 2 As shown in (f) of FIG, a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating the prediction residual described later. Although a CU is basically the same as a prediction unit (PU) and a transform unit (TU), it should be noted that, for example, a sub-block transform (SBT) described later may include multiple TUs smaller than a CU. Additionally, a CU may be processed for each virtual pipeline decoding unit (VPDU) included in the CU. A VPDU is, for example, a fixed unit that may be processed at one stage when pipeline processing is performed in hardware.
[0194] It should be noted that a stream may not include Figure 2 All the hierarchical layers shown in . The order of the hierarchical layers can be exchanged, or any one of the hierarchical layers can be replaced by another hierarchical layer. Here, a picture that is the target of a process to be performed by a device (e.g., the encoder 100 or the decoder 200) is referred to as a current picture. The current picture represents the current picture to be encoded when the process is an encoding process, and the current picture represents the current picture to be decoded when the process is a decoding process. Similarly, for example, a CU or CU block that is the target of a process to be performed by a device (e.g., the encoder 100 or the decoder 200) is referred to as a current block. The current block represents the current block to be encoded when the process is an encoding process, and the current block represents the current block to be decoded when the process is a decoding process.
[0195] (Image structure: slice / slice)
[0196] A picture may be configured with one or more slice units or one or more tiling units to facilitate parallel coding / decoding of the picture.
[0197] A slice is a basic coding unit included in a picture. A picture may include, for example, one or more slices. Additionally, a slice includes one or more coding tree units (CTUs).
[0198] Figure 3 is a conceptual diagram showing an example of a slice configuration. Figure 3In , a picture includes 11×8 CTUs and is split into four slices (slices 1 to 4). Slice 1 includes sixteen CTUs, slice 2 includes twenty-one CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is the shape obtained by splitting the picture horizontally. The boundary of each slice does not need to coincide with the image endpoints and can coincide with any of the boundaries between the CTUs in the image. The processing order (coding order or decoding order) of the CTUs in the slice is, for example, a raster scan order. The slice includes a slice header and encoded data. The characteristics of the slice can be written to the slice header. These characteristics may include the CTU address of the top CTU in the slice, the slice type, etc.
[0199] A tile is a unit rectangular area included in a picture. Tiles of a picture may be assigned numbers called TileIds in a raster scan order.
[0200] Figure 4 is a conceptual diagram showing an example of a sharding configuration. Figure 4 In
[15] , a picture includes 11×8 CTUs and is split into four slices (slices 1 to 4) of rectangular areas. When slicing is used, the order in which the CTUs are processed may be different from the order in which they are processed when slicing is not used. When slicing is not used, multiple CTUs in a picture are typically processed in raster scan order. When multiple slices are used, at least one CTU in each of the multiple slices is processed in raster scan order. For example, Figure 4 As shown in , the processing order of the CTUs included in slice 1 is from the left end of the first column of slice 1 toward the right end of the first column of slice 1, and then continues from the left end of the second column of slice 1 toward the right end of the second column of slice 1.
[0201] It should be noted that a shard may include one or more slices, and a slice may include one or more shards.
[0202] It should be noted that a picture can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more tiles. A picture can be configured with one of a tile set, a tile group, and a tile. For example, it is assumed that the order in which a plurality of tiles are scanned in raster scan order for each tile set is the basic coding order of the tiles. It is assumed that a set of one or more tiles that are consecutive in the basic coding order in each tile set is a tile group. Such a picture can be processed by the splitter 102 described later (see Figure 7 ) to configure.
[0203] (Scalable Coding)
[0204] Figure 5 and Figure 6 is a conceptual diagram showing an example of a scalable stream structure, and for convenience will be referred to Figure 1 Provide a description.
[0205] like Figure 5 As shown in
[15] , encoder 100 can generate a temporally and spatially scalable stream by dividing each of multiple pictures into any of multiple layers and encoding the pictures in each layer. For example, encoder 100 encodes pictures of each layer, thereby achieving scalability when an enhancement layer exists above a base layer. This encoding of each picture is also called scalable coding. In this way, decoder 200 can switch the image quality of the image displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as decoder 200's processing power and external factors such as the communication bandwidth status. As a result, decoder 200 can decode content while freely switching between low and high resolutions. For example, a user of a stream may watch half of a streaming video on a smartphone on their way home and continue watching the video at home on a device such as an internet-connected TV. It should be noted that each of the smartphones and devices described above includes a decoder 200 with the same or different capabilities. In this case, when the device decodes layers up to higher layers in the stream, the user can watch high-quality video at home. In this way, the encoder 100 does not need to generate multiple streams of different image qualities for the same content, and thus can reduce the processing load.
[0206] In addition, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 may generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the metadata. Super-resolution imaging may include, for example, an improvement in the SN ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information for identifying linear or nonlinear filter coefficients (such as used in the super-resolution process), or information for identifying parameter values in a filtering process, machine learning, or a least squares method used in super-resolution processing.
[0207] In an embodiment, a configuration may be provided in which a picture is divided into, for example, slices according to the meaning of, for example, an object in the picture. In this case, the decoder 200 can decode only a partial area in the picture by selecting the slice to be decoded. Additionally, the attributes of the object (person, car, ball, etc.) and the position of the object in the picture (coordinates in the same image) may be stored as metadata. In this case, the decoder 200 is able to identify the position of the desired object based on the metadata and determine the slice that includes the object. For example, Figure 6As shown in FIG, a data storage structure different from the image data (eg, SEI (Supplementary Enhancement Information) message in HEVC) can be used to store metadata. The metadata indicates, for example, the position, size, or color of the main object.
[0208] The metadata can be stored in units of multiple pictures (e.g., streams, sequences, random access units). In this way, the decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the picture unit information, it can identify the picture in which the object (person) exists and determine the position of the object in the picture.
[0209] (Encoder)
[0210] An encoder according to an embodiment will be described. Figure 7 1 is a block diagram illustrating a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes a video in units of blocks.
[0211] like Figure 7 As shown in FIG, the encoder 100 is an apparatus for encoding an image in units of blocks, and includes a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra-frame predictor 124 and the inter-frame predictor 126 are part of the prediction controller.
[0212] The encoder 100 is implemented as, for example, a general-purpose processor and a memory. In this case, when the software program stored in the memory is executed by the processor, the processor functions as the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128.
[0213] (Encoder installation example)
[0214] Figure 8 1 is a functional block diagram showing an example of an installation of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The multiple components of the encoder 100 shown in FIG. 1 are mounted on Figure 8 The processor a1 and the memory a2 are shown.
[0215] Processor a1 is a circuit that performs information processing and is coupled to memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes an image. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits. Alternatively, for example, processor a1 may be responsible for Figure 7 The roles of two or more constituent elements among the multiple constituent elements of the encoder 100 shown in FIG.
[0216] Memory a2 is a dedicated or general-purpose memory for storing information used by processor a1 to encode images. Memory a2 may be an electronic circuit and may be connected to processor a1. Alternatively, memory a2 may be included in processor a1. Alternatively, memory a2 may be a collection of multiple electronic circuits. Alternatively, memory a2 may be a magnetic disk, optical disk, or the like, or may be represented as a storage device, recording medium, or the like. Alternatively, memory a2 may be a non-volatile memory or a volatile memory.
[0217] For example, the memory a2 may store an image to be encoded or a bit stream corresponding to the encoded image. Alternatively, the memory a2 may store a program for causing the processor a1 to encode an image.
[0218] Alternatively, for example, memory a2 may assume Figure 7 The memory a2 may be used to store information in the plurality of components of the encoder 100 shown in FIG. Figure 7 1 and 2. The roles of the block memory 118 and the frame memory 122 are shown in FIG. More specifically, the memory a2 can store reconstructed blocks, reconstructed pictures, and the like.
[0219] It should be noted that in encoder 100, it may not be possible to implement Figure 7 All elements among the multiple constituent elements indicated in the figure, etc., and all processes described in this document may not be performed. Figure 7 A portion of the constituent elements indicated in the , etc. may be included in another device, or a portion of the process described herein may be performed by another device.
[0220] Hereinafter, the overall flow of a process performed by the encoder 100 is described, and then each of the constituent elements included in the encoder 100 will be described.
[0221] (Overall flow of the encoding process)
[0222] Figure 9 is a flowchart indicating one example of the overall encoding process performed by the encoder 100, and will be referred to for convenience. Figure 7 Provide a description.
[0223] First, the splitter 102 of the encoder 100 splits each of the pictures included in the input image into a plurality of blocks of a fixed size (e.g., 128×128 pixels) (step Sa_1). The splitter 102 then selects a splitting pattern for the fixed-size blocks (also referred to as block shape) (step Sa_2). In other words, the splitter 102 further splits the fixed-size blocks into a plurality of blocks forming the selected splitting pattern. For each of the plurality of blocks, the encoder 100 performs steps Sa_3 to Sa_9 for the block (i.e., the current block to be encoded).
[0224] The prediction controller 128 and the prediction executor (including the intra predictor 124 and the inter predictor 126) generate a prediction image for the current block (step Sa-3). The prediction image may also be referred to as a prediction signal, a prediction block, or a prediction sample.
[0225] Next, the subtractor 104 generates the difference between the current block and the predicted image as a prediction residual (step Sa_4). The prediction residual may also be referred to as a prediction error.
[0226] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients may sometimes be referred to as a coefficient block.
[0227] Next, the entropy encoder 110 encodes (specifically, performs entropy encoding) the plurality of quantized coefficients and prediction parameters related to the generation of the predicted image to generate a stream (step Sa_6). The stream may sometimes be referred to as an encoded bit stream or a compressed bit stream.
[0228] Next, the inverse quantizer 112 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 114 performs inverse transform on the result to restore the prediction residual (step Sa_7).
[0229] Next, the adder 116 adds the predicted image and the restored prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image can also be called a reconstructed block or a decoded image block.
[0230] When the reconstructed image is generated, the loop filter 120 performs filtering on the reconstructed image as needed (step Sa_9).
[0231] Then, the encoder 100 determines whether encoding of the entire picture has ended (step Sa_10). When it is determined that encoding has not ended ("No" in step Sa_10), the process starting from step Sa_2 is repeated for the next block of the picture.
[0232] Although the encoder 100 selects one splitting mode for fixed-size blocks and encodes each block according to the splitting mode in the example described above, it should be noted that each block may be encoded according to a corresponding one of a plurality of splitting modes. In this case, the encoder 100 may evaluate the cost of each of the plurality of splitting modes and, for example, may select as the output stream a stream obtained by encoding according to the splitting mode that produces the minimum cost.
[0233] As shown, the processes in steps Sa_1 to Sa_10 are performed sequentially by the encoder 100. Alternatively, two or more of the processes may be performed in parallel, the processes may be reordered, and the like.
[0234] The encoding process employed by encoder 100 is a hybrid encoding process using predictive encoding and transform encoding. Specifically, predictive encoding is performed by an encoding loop that includes a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a loop filter 120, a block memory 118, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. In other words, the prediction execution unit, which includes intra-frame predictor 124 and inter-frame predictor 126, is part of the encoding loop.
[0235] (Splitter)
[0236] The splitter 102 splits each picture included in the original image into a plurality of blocks and outputs each block to the subtractor 104. For example, the splitter 102 first splits the picture into blocks of a fixed size (e.g., 128×128 pixels). Other fixed block sizes may be used. Fixed-size blocks are also referred to as coding tree units (CTUs). The splitter 102 then splits each fixed-size block into blocks of a variable size (e.g., 64×64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a splitting mode. Variable-size blocks may also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various categories of processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or some of the blocks in the picture may be processed in units of CUs, PUs, or TUs.
[0237] Figure 10 : is a conceptual diagram for illustrating an example of block splitting according to an embodiment. Figure 10 , solid lines represent block boundaries of blocks split by quadtree block splitting, and dashed lines represent block boundaries of blocks split by binarytree block splitting.
[0238] Here, the block 10 is a square block having 128×128 pixels (128×128 block). The 128×128 block 10 is first split into four square blocks of 64×64 pixels (quadtree block splitting).
[0239] The 64×64 pixel block on the upper left is further split vertically into two rectangular 32×64 pixel blocks, and the 32×64 pixel block on the left is further split vertically into two rectangular 16×64 pixel blocks (binary tree block splitting). As a result, the 64×64 pixel block on the upper left is split into two 16×64 pixel blocks 11 and 12 and one 32×64 pixel block 13.
[0240] The upper right 64×64 pixel block is split horizontally into two rectangular 64×32 pixel blocks 14 and 15 (binary tree block splitting).
[0241] The 64×64 pixel block in the lower left square is first split into four square 32×32 pixel blocks (quadtree block splitting). The upper left block and the lower right block in the four square 32×32 pixel blocks are further split. The 32×32 pixel block in the upper left square is vertically split into two rectangular 16×32 pixel blocks, and the 16×32 pixel block on the right is further horizontally split into two 16×16 pixel blocks (binary tree block splitting). The 32×32 pixel block in the lower right is horizontally split into two 32×16 pixel blocks (binary tree block splitting). The 32×32 pixel block in the upper right square is horizontally split into two rectangular 32×16 pixel blocks (binary tree block splitting). As a result, the square 64×64 pixel block in the lower left is split into a rectangular 16×32 pixel block 16, two square 16×16 pixel blocks 17 and 18, two square 32×32 pixel blocks 19 and 20, and two rectangular 32×16 pixel blocks 21 and 22.
[0242] The lower right 64×64 pixel block 23 is not split.
[0243] As described above, in Figure 10 In FIG, based on recursive quadtree and binary tree block splitting, block 10 is split into thirteen variable-sized blocks 11 to 23. This type of splitting is also called quadtree plus binary tree (QTBT) splitting.
[0244] It should be noted that in Figure 10In the example above, a block is split into four or two blocks (quadtree or binary tree block splitting), but the splitting is not limited to these examples. For example, a block can be split into three blocks (ternary block splitting). Splitting that includes this ternary block splitting is also called multi-type tree (MBT) splitting.
[0245] Figure 11 FIG. 1 is a block diagram showing an example of a functional configuration of the splitter 102 according to an embodiment. Figure 11 As shown in FIG, the splitter 102 may include a block split determiner 102a. As an example, the block split determiner 102a may perform the following process.
[0246] For example, the block splitting determiner 102a may obtain or retrieve block information from the block memory 118 and / or the frame memory 122 and determine a splitting mode (e.g., the splitting mode described above) based on the block information. The splitter 102 splits the original image according to the splitting mode and outputs at least one block obtained by the splitting to the subtractor 104.
[0247] Additionally, for example, the block split determiner 102a outputs one or more parameters indicating the determined split mode (e.g., the split mode described above) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 may transform the prediction residual based on the one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 may generate a predicted image based on the one or more parameters. Additionally, the entropy encoder 110 may perform entropy encoding on the one or more parameters.
[0248] As indicated below, parameters related to the splitting mode may be written in the stream, as an example.
[0249] Figure 12 is a conceptual diagram illustrating examples of split modes. Examples of split modes include: split into four regions (QT), in which blocks are split horizontally into two regions and vertically into two regions; split into three regions (HT or VT), in which blocks are split in the same direction at a ratio of 1:2:1; split into two regions (HB or VB), in which blocks are split in the same direction at a ratio of 1:1; and no split (NS).
[0250] It should be noted that the split mode does not have a block splitting direction in the case of splitting into four regions and not splitting, and the split mode has splitting direction information in the case of splitting into two regions or three regions.
[0251] Figure 13A is a conceptual diagram showing an example of a syntax tree of a split pattern.
[0252] Figure 13B is a conceptual diagram for illustrating another example of a syntax tree of a split pattern.
[0253] Figure 13A and Figure 13B is a conceptual diagram showing an example of a syntax tree for a split pattern. Figure 13A In the example, first, there is information indicating whether to perform splitting (S: split flag), and next, there is information indicating whether to perform splitting into four regions (QT: QT flag). Next, there is information indicating which of splitting into three regions and splitting into two regions is to be performed (TT: TT flag, or BT: BT flag), and then, there is information indicating the direction of division (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of at least one block obtained by splitting according to such a splitting pattern can be further repeatedly split in a similar process. In other words, as an example, whether to perform splitting, whether to perform splitting into four regions, which of the horizontal direction and the vertical direction is the direction of the splitting method to be performed, which of the splitting into three regions and the splitting into two regions is to be performed can be recursively determined, and can be determined based on the splitting pattern determined by Figure 13A The encoding order disclosed by the syntax tree shown in determines the encoding result in the stream.
[0254] Additionally, although the information items indicating S, QT, TT, and Ver, respectively, are Figure 13A In the syntax tree shown in FIG, the information items indicating S, QT, Ver, and BT, respectively, are arranged in the order listed. In other words, in Figure 13B In the example of , first, there is information indicating whether to perform splitting (S: Split Flag), and next, there is information indicating whether to perform splitting into four regions (QT: QT Flag). Next, there is information indicating the splitting direction (Ver: Vertical Flag, or Hor: Horizontal Flag), and then there is information indicating which of splitting into two regions and splitting into three regions is to be performed (BT: BT Flag, or TT: TT Flag).
[0255] It should be noted that the split patterns described above are examples, and a split pattern other than the described split pattern may be used, or a part of the described split pattern may be used.
[0256] (Subtractor)
[0257] The subtractor 104 subtracts the predicted image (prediction samples input from the prediction controller 128 indicated below) from the original image input from the splitter 102 in units of blocks and split by the splitter 102. In other words, the subtractor 104 calculates the prediction residual (also referred to as error) of the current block. The subtractor 104 then outputs the calculated prediction residual to the transformer 106.
[0258] The original image may be an image that has been input into the encoder 100 as a signal (for example, a luminance signal and two chrominance signals) representing an image of each picture included in a video. The signal representing the image may also be referred to as a sample.
[0259] (Converter)
[0260] The transformer 106 transforms the prediction residual in the spatial domain into a transform coefficient in the frequency domain and outputs the transform coefficient to the quantizer 108. More specifically, the transformer 106 applies, for example, a defined discrete cosine transform (DCT) or discrete sine transform (DST) to the prediction residual in the spatial domain. The defined DCT or DST may be predefined.
[0261] It should be noted that the transformer 106 can adaptively select a transform type from a plurality of transform types and transform the prediction residual into a transform coefficient by using a transform basis function corresponding to the selected transform type. This type of transformation is also called explicit multi-kernel transform (EMT) or adaptive multiple transform (AMT). The transform basis function may also be referred to as a basis.
[0262] Transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. It should be noted that these transform types can be expressed as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 is a diagram indicating example transform basis functions for example transform types. Figure 14 In , N indicates the number of input pixels. For example, selecting a transform type from a plurality of transform types may depend on a prediction type (one of intra prediction and inter prediction) and may depend on an intra prediction mode.
[0263] Information indicating whether such EMT or AMT is applied (referred to as, for example, an EMT flag or an AMT flag) and information indicating the selected transform type are typically signaled at the CU level. It should be noted that signaling such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0264] Additionally, the transformer 106 may re-transform the transform coefficients (which are the transform results). This re-transformation is also referred to as an adaptive secondary transform (AST) or a non-separable secondary transform (NSST). For example, the transformer 106 performs re-transformation in units of sub-blocks (e.g., 4×4 pixel sub-blocks) included in the transform coefficient block corresponding to the intra-frame prediction residual. Information indicating whether NSST is applied and information related to the transform matrix used for NSST are typically signaled at the CU level. It should be noted that signaling such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0265] Transformer 106 can employ both separable and non-separable transforms. Separable transforms are methods in which transforms are performed multiple times by performing transforms separately for each of a plurality of directions according to the dimensionality of the input. Non-separable transforms are methods in which collective transforms are performed in which two or more dimensions in a multi-dimensional input are collectively treated as a single dimension.
[0266] In one example of a non-separable transform, when the input is a 4x4 pixel block, the 4x4 pixel block is considered to be a single array comprising sixteen elements, and the transform applies a 16x16 transformation matrix to the array.
[0267] In another example of a non-separable transform, an input block of 4x4 pixels is treated as a single array comprising sixteen elements, and then a transform can be performed in which a number of given rotations are applied to the array (a hypercube given transform).
[0268] In the transformation in the transformer 106, the transformation type of the transformation basis function to be transformed into the frequency domain according to the region in the CU can be switched. Examples include spatially varying transform (SVT).
[0269] Figure 15 is a conceptual diagram for illustrating an example of SVT.
[0270] In SVT, Figure 15 As shown in , the CU is split into two equal regions horizontally or vertically, and only one of the regions is transformed into the frequency domain. The transform base type can be set for each region. For example, DST7 and DST8 are used. For example, in the two regions obtained by splitting the CU vertically into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Similarly, in the two regions obtained by splitting the CU horizontally into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in the two regions, DST7 is used for the region at position 1. Although in Figure 15 In the example shown in , only one of the two regions in the CU is transformed and the other is not transformed, but each of the two regions can be transformed. Additionally, the splitting method can include not only splitting into two regions but also splitting into four regions. Additionally, the splitting method can be more flexible. For example, information indicating the splitting method can be encoded and can be signaled in the same manner as CU splitting. It should be noted that SVT can also be referred to as sub-block transform (SBT).
[0271] The AMT and EMT described above may be referred to as MTS (Multiple Transform Selection). When MTS is applied, transform types such as DST7 and DCT8 may be selected, and information indicating the selected transform type may be encoded as index information for each CU. There is another process called IMTS (Implicit MTS) as a process for selecting a transform type to be used for an orthogonal transform performed without encoding index information. When IMTS is applied, for example, when the CU has a rectangular shape, an orthogonal transform of a rectangular shape may be performed using DST7 for the short side and DST2 for the long side. Additionally, for example, when the CU has a square shape, an orthogonal transform of a rectangular shape may be performed using DCT2 when MTS is valid in the sequence and using DST7 when MTS is invalid in the sequence. DCT2 and DST7 are merely examples. Other transform types may be used, and the combination of transform types may also be changed for different combinations of transform types. IMTS may be used only for intra-prediction blocks, or may be used for both intra-prediction blocks and inter-prediction blocks.
[0272] The three processes of MTS, SBT and IMTS have been described above as selection processes for selectively switching the transform type for orthogonal transform. However, all three selection processes may be adopted, or only a part of the selection process may be selectively adopted. For example, whether to adopt one or more of the selection processes may be identified based on flag information in a header such as SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching the transform type may be a selection process different from the above three selection processes, or each of the three selection processes may be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is performed. Function [1] is a function for performing an orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing an orthogonal transform of the entire CU and determining the transform type based on a determined rule without encoding the information indicating the transform type. Function [3] is a function for performing an orthogonal transform of a partial region of the CU and encoding information indicating the transform type used in the transform. Function [4] is a function for performing orthogonal transform on a partial region of a CU and determining a transform type based on a determined rule without encoding information indicating the transform type used in the transform. The determined rule may be predetermined.
[0273] It should be noted that whether to apply MTS, IMTS and / or SBT may be determined for each processing unit. For example, whether to apply MTS, IMTS and / or SBT may be determined for each sequence, picture, tile, slice, CTU or CU.
[0274] It should be noted that the tool for selectively switching transform types in the present disclosure can be described as a method for selectively selecting a basis to be used in a transform process, a selection process, or a process for selecting a basis. Alternatively, the tool for selectively switching transform types can be described as a mode for adaptively selecting transform types.
[0275] Figure 16 is a flowchart showing one example of a process performed by the converter 106, and for convenience will be referred to Figure 7 Provide a description.
[0276] For example, the transformer 106 determines whether to perform an orthogonal transform (step St_1). Here, when it is determined that an orthogonal transform is to be performed ("Yes" in step St_1), the transformer 106 selects a transform type for the orthogonal transform from a plurality of transform types (step St_2). Next, the transformer 106 performs an orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transform type to the entropy encoder 110 to allow the entropy encoder 110 to encode the information (step St_4). On the other hand, when it is determined not to perform an orthogonal transform ("No" in step St_1), the transformer 106 outputs information indicating that an orthogonal transform is not to be performed to allow the entropy encoder 110 to encode the information (step St_5). It should be noted that whether to perform an orthogonal transform in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transform can be performed using a defined transform type without encoding information indicating the transform type used in the orthogonal transform. The defined transformation types may be pre-defined.
[0277] Figure 17 is a flowchart showing one example of a process performed by the converter 106, and for convenience will be referred to Figure 7 It should be noted that Figure 17 The example shown in FIG is a case where the transform type used in the orthogonal transform is selectively switched (as in Figure 16 An example of an orthogonal transform in the case of the example shown in ).
[0278] As an example, the first transform type group may include DCT2, DCT7, and DCT8. As another example, the second transform type group may include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group may partially overlap with each other, or may be completely different from each other.
[0279] The transformer 106 determines whether the transform size is less than or equal to the determined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the determined value ("Yes" in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform type included in the first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one transform type included in the first transform type group to the entropy encoder 110, so as to allow the entropy encoder 110 to encode the information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to the predetermined value ("No" in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The determined value may be a threshold value and may be a predetermined value.
[0280] In step Su_3, the information indicating the transform type used in the orthogonal transform may be information indicating a combination of a transform type to be applied vertically in the current block and a transform type to be applied horizontally in the current block. The first type group may include only one transform type, and the information indicating the transform type used in the orthogonal transform may not be encoded. The second transform type group may include multiple transform types, and information indicating the transform type used in the orthogonal transform among one or more transform types included in the second transform type group may be encoded.
[0281] Alternatively, the transform type may be indicated based on the transform size without encoding information indicating the transform type. It should be noted that such determination is not limited to determining whether the transform size is less than or equal to a certain value, and other processes may also be used to determine the transform type used in the orthogonal transform based on the transform size.
[0282] (Quantizer)
[0283] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in the determined scanning order and quantizes the scanned transform coefficients based on the quantization parameters (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients of the current block (hereinafter also referred to as quantized coefficients) to the entropy encoder 110 and the inverse quantizer 112. The determined scanning order may be predetermined.
[0284] The determined scanning order is the order for quantizing / inverse quantizing transform coefficients. For example, the determined scanning order may be defined as ascending order of frequency (from low frequency to high frequency) or descending order of frequency (from high frequency to low frequency).
[0285] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, as the value of the quantization parameter increases, the quantization step size also increases. In other words, as the value of the quantization parameter increases, the error in the quantized coefficients (quantization error) increases.
[0286] Additionally, a quantization matrix may be used for quantization. For example, several types of quantization matrices may be used corresponding to frequency transform sizes such as 4×4, 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luminance and chrominance pixel components. It should be noted that quantization means digitizing values sampled at a determined interval corresponding to a determined level. In the art, quantization may be referred to using other expressions (e.g., rounding and scaling), and rounding and scaling may be employed. The determined interval and the determined level may be predetermined.
[0287] Methods for using a quantization matrix include a method of using a quantization matrix that has been directly set on the encoder 100 side, and a method of using a quantization matrix that has been set as a default (default matrix). On the encoder 100 side, a quantization matrix suitable for image characteristics can be set by directly setting the quantization matrix. However, this case may have the disadvantage of increasing the amount of decoding for encoding the quantization matrix. It should be noted that the quantization matrix used for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.
[0288] There is a method for quantizing high-frequency coefficients and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be considered equivalent to a method using a quantization matrix (flat matrix) whose coefficients have the same value.
[0289] The quantization matrix can be encoded, for example, at the sequence level, picture level, slice level, tile level, or CTU level. The quantization matrix can be specified using, for example, a sequence parameter set (SPS) or a picture parameter set (PPS). The SPS includes parameters for a sequence, and the PPS includes parameters for a picture. Each of the SPS and PPS may be referred to simply as a parameter set.
[0290] When a quantization matrix is used, the quantizer 108 uses the value of the quantization matrix to scale the quantization width, which can be calculated based on, for example, a quantization parameter, for each transform coefficient. The quantization process performed without using a quantization matrix may be a process for quantizing the transform coefficients according to the quantization width calculated based on, for example, a quantization parameter. It should be noted that in the quantization process performed without using any quantization matrix, the quantization width may be multiplied by a predetermined value common to all transform coefficients in a block. This predetermined value may be predetermined.
[0291] Figure 18 1 is a block diagram showing an example of a functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a difference quantization parameter generator 108a, a predicted quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization performer 108e.
[0292] Figure 19 is a flowchart showing one example of a quantization process performed by the quantizer 108, and will be referred to for convenience. Figure 7 and Figure 18 Provide a description.
[0293] As an example, the quantizer 108 may be based on Figure 19 The flowchart shown in FIG performs quantization for each CU. More specifically, the quantization parameter generator 108 c determines whether to perform quantization (step Sv_1). Here, when it is determined that quantization is to be performed ("Yes" in step Sv_1), the quantization parameter generator 108 c generates a quantization parameter for the current block (step Sv_2) and stores the quantization parameter in the quantization parameter storage device 108 d (step Sv_3).
[0294] Next, the quantization performer 108e quantizes the transform coefficients of the current block using the quantization parameters generated in step Sv_2 (step Sv_4). The predicted quantization parameter generator 108b then obtains the quantization parameters for a processing unit different from the current block from the quantization parameter storage device 108d (step Sv_5). The predicted quantization parameter generator 108b generates a predicted quantization parameter for the current block based on the obtained quantization parameters (step Sv_6). The difference quantization parameter generator 108a calculates the difference between the quantization parameter for the current block generated by the quantization parameter generator 108c and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 108b (step Sv_7). The difference quantization parameter can be generated by calculating the difference. The difference quantization parameter generator 108a outputs the difference quantization parameter to the entropy encoder 110 to allow the entropy encoder 110 to encode the difference quantization parameter (step Sv_8).
[0295] It should be noted that the difference quantization parameter can be encoded at, for example, the sequence level, the picture level, the slice level, the tile level, or the CTU level. Additionally, the initial value of the quantization parameter can be encoded at the sequence level, the picture level, the slice level, the tile level, or the CTU level. During initialization, the quantization parameter can be generated using the initial value of the quantization parameter and the difference quantization parameter.
[0296] It should be noted that the quantizer 108 may include a plurality of quantizers and may apply dependent quantization in which a transform coefficient is quantized using a quantization method selected from a plurality of quantization methods.
[0297] (Entropy Encoder)
[0298] Figure 20 is a block diagram showing one example of the functional configuration of the entropy encoder 110 according to the embodiment, and will be referred to for convenience. Figure 7 . The entropy encoder 110 generates a stream by entropy encoding the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as entropy encoding. More specifically, the entropy encoder 110 shown includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, in which multi-level signals such as quantized coefficients and prediction parameters are converted into binary signals. Examples of binarization methods include truncated Rice binarization, exponential Golomb code, and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntax element or the surrounding state (i.e., the probability of occurrence of the binary signal). Examples of methods for deriving context values include bypassing, referring to syntax elements, referring to upper and left adjacent blocks, referring to level information, etc. The binary arithmetic encoder 110c performs arithmetic encoding on the binary signal using the derived context.
[0299] Figure 21 1 is a conceptual diagram illustrating an example flow of a CABAC process in the entropy encoder 110. First, initialization is performed in the CABAC process in the entropy encoder 110. During initialization, initialization and setting of an initial context value in the binary arithmetic encoder 110c are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c may sequentially perform binarization and arithmetic coding on a plurality of quantized coefficients in a CTU. Each time arithmetic coding is performed, the context controller 110b may update the context value. The context controller 110b may then save the context value as post-processing. For example, the saved context value may be used to initialize the context value for the next CTU.
[0300] (Inverse Quantizer)
[0301] The inverse quantizer 112 inversely quantizes the quantized coefficients input from the quantizer 108. More specifically, the inverse quantizer 112 inversely quantizes the quantized coefficients of the current block in the determined scanning order. The inverse quantizer 112 then outputs the inversely quantized transform coefficients of the current block to the inverse transformer 114. The determined scanning order may be predetermined.
[0302] (Inverse Converter)
[0303] The inverse transformer 114 restores the prediction residual by inversely transforming the transform coefficients input from the inverse quantizer 112. More specifically, the inverse transformer 114 restores the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the restored prediction residual to the adder 116.
[0304] It should be noted that since information is typically lost in quantization, the recovered prediction residual does not match the prediction residual calculated by the subtractor 104. In other words, the recovered prediction residual typically includes quantization errors.
[0305] (Adder)
[0306] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transformer 114 and the predicted image input from the prediction controller 128. Thus, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block can also be referred to as a partially decoded block.
[0307] (Block Storage)
[0308] The block memory 118 is a storage device for storing blocks in the current picture (for example, for intra prediction). More specifically, the block memory 118 stores the reconstructed image output from the adder 116.
[0309] (Frame Memory)
[0310] The frame memory 122 is a storage device for storing reference pictures used in inter-frame prediction, for example, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.
[0311] (Loop Filter)
[0312] The loop filter 120 applies a loop filter to the reconstructed image output by the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter (loop filter) used in the encoding loop. Examples of the loop filter include an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, and the like.
[0313] Figure 22 1 is a block diagram showing an example of the functional configuration of the loop filter 120 according to the embodiment. Figure 22As shown in FIG, the loop filter 120 includes a deblocking filter executor 120a, an SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs a deblocking filter process on the reconstructed image. The SAO executor 120b performs an SAO process on the reconstructed image after the deblocking filter process. The ALF executor 120c performs an ALF process on the reconstructed image after the SAO process. The ALF and deblocking filter will be described in detail later. The SAO process is a process for enhancing image quality by reducing ringing (a phenomenon in which pixel values are distorted like waves around edges) and correcting deviations in pixel values. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that in some embodiments, the loop filter 120 may not include Figure 22 All the constituent elements disclosed in, and may include some of the constituent elements, and may include additional elements. Additionally, the loop filter 120 may be configured to Figure 22 The above processes may be performed in a different processing order than that disclosed in the specification, all processes may not be performed, etc.
[0314] (Loop Filter > Adaptive Loop Filter)
[0315] In ALF, a least squares error filter is applied to remove compression artifacts. For example, a filter selected from multiple filters based on the direction and activity of the local gradient is applied to each 2×2 pixel sub-block in the current block.
[0316] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of a plurality of categories (e.g., fifteen or twenty-five categories). The classification of the sub-block can be based on, for example, gradient directionality and activity. In the example, a category index C (e.g., C=5D+A) is calculated or determined based on the gradient directionality D (e.g., 0 to 2 or 0 to 4) and the gradient activity A (e.g., 0 to 4). Then, based on the classification index C, each sub-block is classified into one of the plurality of categories.
[0317] For example, the gradient directionality D is calculated by comparing the gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). In addition, the gradient activity A is calculated by adding the gradients in multiple directions and quantifying the added result.
[0318] A filter to be used for each sub-block may be determined from among a plurality of filters based on a result of such classification.
[0319] The filter shape to be used in the ALF is, for example, a circularly symmetric filter shape. Figures 23A to 23C is a conceptual diagram for illustrating an example of a filter shape used in ALF. Figure 23Ashows a 5×5 diamond filter, Figure 23B A 7×7 diamond filter is shown, and Figure 23C A 9×9 diamond filter is shown. Information indicating the filter shape is typically signaled at the picture level. It should be noted that signaling such information indicating the filter shape does not necessarily need to be performed at the picture level, but can also be performed at another level (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0320] For example, the turning on or off of ALF can be determined at the picture level or the CU level. For example, whether to apply ALF to luma can be determined at the CU level, and whether to apply ALF to chroma can be determined at the picture level. Information indicating whether ALF is turned on or off is typically signaled at the picture level or the CU level. It should be noted that signaling information indicating whether ALF is turned on or off does not necessarily need to be performed at the picture level or the CU level, and can be performed at another level (e.g., sequence level, slice level, tile level, or CTU level).
[0321] Alternatively, as described above, a filter is selected from a plurality of filters and the ALF process for the sub-block is performed. A coefficient set for each of the plurality of filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that signaling the coefficient set does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or sub-block level).
[0322] (Loop Filter > Cross-Component Adaptive Loop Filter)
[0323] Figure 23D is a conceptual diagram for illustrating an example flow of cross-component ALF (CC-ALF). Figure 23E is used to show the CC-ALF (e.g. Figure 23D Conceptual diagram of an example of the filter shape used in CC-ALF). Figure 23D and Figure 23E The CC-ALF operates by applying a linear diamond filter to the luma channel of each chroma component. For example, the filter coefficients can be sent in APS, scaled by a factor of 2^10, and rounded for fixed point representation. For example, in Figure 23D In , the Y sample (first component) is used for CCALF for Cb and CCALF for Cr (a component different from the first component).
[0324] The application of the filter can be controlled over a variable block size and signaled by a context coded flag received for each sample block. The block size along with the CC-ALF enable flag can be received at the slice level for each chroma component. CC-ALF can support various block sizes, for example, (in chroma samples) 16×16 pixels, 32×32 pixels, 64×64 pixels, 128×128 pixels.
[0325] (Loop Filter > Joint Chroma Cross-Component Adaptive Loop Filter)
[0326] An example of joint chroma-CCALF is given in Figure 23F and Figure 23G Shown in. Figure 23F is a conceptual diagram illustrating an example flow of joint chroma CCALF. Figure 23G is a table showing example weighted index candidates. As shown, a CCALF filter is used to generate a CCALF-filtered output as a chroma refinement signal for one color component, while a weighted version of the same chroma refinement signal is applied to the other color component. In this way, the complexity of the existing CCALF is reduced by approximately half. The weighted value can be decoded into a sign flag and a weighted index. The weighted index (denoted as weight_index) can be decoded into 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight, which is a non-zero magnitude. For example, the magnitude of JcCcWeight can be determined as follows:
[0327] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index>>2;
[0328] Otherwise, JcCcWeight is equal to 4 / (weight_index-4).
[0329] Block-level on / off control for ALF filtering can be separate for Cb and Cr. This is the same as in CCALF, and two separate sets of block-level on / off control flags can be decoded. Unlike CCALF, the Cb and Cr on / off control blocks in this paper are the same size, so only one block size variable can be decoded.
[0330] (Loop Filter > Deblocking Filter)
[0331] In the deblocking filtering process, the loop filter 120 performs a filtering process on block boundaries in the reconstructed image in order to reduce distortion occurring at the block boundaries.
[0332] Figure 24 FIG. 1 is a diagram showing a loop filter 120 (see FIG. 1 ) used as a deblocking filter. Figure 7 and Figure 22 ) is a block diagram of an example of a specific configuration of the deblocking filter executor 120a.
[0333] The deblocking filter executor 120 a includes: a boundary determiner 1201 ; a filter determiner 1203 ; a filter executor 1205 ; a process determiner 1208 ; a filter characteristic determiner 1207 ; and switches 1202 , 1204 , and 1206 .
[0334] The boundary determiner 1201 determines whether a pixel to be deblocked (ie, a current pixel) exists around a block boundary, and then outputs the determination result to the switch 1202 and the processing determiner 1208.
[0335] In the case where the boundary determiner 1201 has determined that the current pixel exists around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1204. In the case where the boundary determiner 1201 has determined that the current pixel does not exist around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1206. It should be noted that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.
[0336] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel value of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0337] In the case where the filter determiner 1203 has determined that deblocking filtering is to be performed on the current pixel, the switch 1204 outputs the unfiltered image obtained by the switch 1202 to the filter executor 1205. In the opposite case where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained by the switch 1202 to the switch 1206.
[0338] When an unfiltered image is obtained through switches 1202 and 1204 , filter executor 1205 performs deblocking filtering on the current pixel using the filter characteristics determined by filter characteristic determiner 1207 . Filter executor 1205 then outputs the filtered pixel to switch 1206 .
[0339] Under the control of the processing determiner 1208 , the switch 1206 selectively outputs one of pixels that have not been subjected to deblocking filtering and pixels that have been subjected to deblocking filtering by the filtering executor 1205 .
[0340] The processing determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary, and when the filter determiner 1203 has determined that deblocking filtering is to be performed on the current pixel, the processing determiner 1208 causes the switch 1206 to output the pixel that has been deblocking filtered. Alternatively, in addition to the above-mentioned case, the processing determiner 1208 causes the switch 1206 to output the pixel that has not been deblocking filtered. By repeatedly outputting pixels in this manner, a filtered image is output from the switch 1206. It should be noted that Figure 24 The configuration shown in FIG. 1 is one example of a configuration in the deblocking filter executor 120 a. The deblocking filter executor 120 a may have various configurations.
[0341] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having symmetric filtering characteristics with respect to a block boundary.
[0342] In the deblocking filtering process, pixel values and quantization parameters can be used to select one of two deblocking filters (i.e., a strong filter and a weak filter) with different characteristics. In the case of a strong filter, when pixels p0 to p2 and pixels q0 to q2 exist across block boundaries, as shown in FIG. Figure 25 As shown in , by performing calculation according to, for example, the following expressions, the pixel values of the corresponding pixels q0 to q2 are changed to pixel values q′0 to q′2.
[0343] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0344] q'1=(p0+q0+q1+q2+2) / 4
[0345] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0346] It should be noted that in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the corresponding pixels p0 to p2 and pixels q0 to q2. In addition, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of the block boundary to the pixel q2. In addition, on the right side of each of the expressions, the coefficient multiplied by the corresponding pixel value of the pixel to be used for deblocking filtering is the filter coefficient.
[0347] Furthermore, during deblocking filtering, clipping can be performed so that the calculated pixel value does not change by more than a threshold value. For example, during clipping, the pixel value calculated according to the above expression can be clipped to a value obtained by "calculated pixel value ± 2 × threshold value" using a threshold value determined based on a quantization parameter. This prevents oversmoothing.
[0348] Figure 26 is a conceptual diagram for illustrating a block boundary on which a deblocking filtering process is performed. Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0349] The block boundary on which the deblocking filtering process is performed is, for example, a boundary between CUs, Pus, or TUs having 8×8 pixel blocks, such as Figure 26 As shown in FIG. The deblocking filtering process can be performed in units of, for example, four rows or four columns. First, for Figure 26 The block P and block Q shown in FIG determine the boundary strength (Bs) value, as shown in FIG. Figure 27 Indicated in.
[0350] according to Figure 27 The Bs value in the image can determine whether to perform a deblocking filtering process on block boundaries belonging to the same image with different intensities. When the Bs value is 2, a deblocking filtering process for a chrominance signal is performed. When the Bs value is 1 or greater and a determined condition is satisfied, a deblocking filtering process for a luminance signal is performed. The determined condition may be predetermined. It should be noted that the condition for determining the Bs value is not limited to Figure 27 Those indicated in , and the Bs value can be determined based on another parameter.
[0351] (Predictor (Intra Predictor, Inter Predictor, Prediction Controller))
[0352] Figure 28 1 is a flowchart showing an example of a process performed by the predictor of the encoder 100. It should be noted that the predictor includes all or part of the following constituent elements: an intra-frame predictor 124; an inter-frame predictor 126; and a prediction controller 128. The prediction performer includes, for example, the intra-frame predictor 124 and the inter-frame predictor 126.
[0353] The predictor generates a predicted image for the current block (step Sb_1). This predicted image may also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra-frame predicted image (image prediction signal) or an inter-frame predicted image (inter-frame prediction signal). The predictor generates a predicted image for the current block using a reconstructed image obtained from another block by generating a predicted image, generating a prediction residual, generating quantized coefficients, restoring the prediction residual, and adding the predicted image.
[0354] The reconstructed image may be, for example, an image in a reference picture, or an image of a coded block (i.e., other blocks described above) in a current picture that is a picture including the current block. The coded block in the current picture may be, for example, a neighboring block of the current block.
[0355] Figure 29is a flowchart illustrating another example of a process performed by the predictor of the encoder 100 .
[0356] The predictor generates a predicted image using the first method (step Sc_1a), generates a predicted image using the second method (step Sc_1b), and generates a predicted image using the third method (step Sc_1c). The first method, the second method, and the third method may be mutually different methods for generating a predicted image. Each of the first method to the third method may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0357] Next, the prediction processor evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates a cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the predicted images by comparing the costs C of the predicted images. Note that the cost C can be calculated, for example, according to an expression of an RD optimization model (e.g., C=D+λ×R). In this expression, D represents a compression artifact of the predicted image and is expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. Additionally, R represents the bit rate of the stream. Additionally, λ represents a multiplier, for example, according to a method of Lagrange multipliers.
[0358] Then, the predictor selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects a method or mode for obtaining the final predicted image. For example, the predictor selects a predicted image with the minimum cost C based on the cost C calculated for the predicted image. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be performed based on parameters used in the encoding process. The encoder 100 may transform information for identifying the selected predicted image, method, or mode into a stream. The information may be, for example, a flag, etc. In this way, the decoder 200 is able to generate a predicted image according to the method or mode selected by the encoder 100 based on the information. Note that in Figure 29 In the example shown in , after generating the predicted image using the corresponding method, the predictor selects any one of the predicted images. However, the predictor may select a method or mode based on the parameters used in the encoding process described above before generating the predicted image, and may generate the predicted image according to the selected method or mode.
[0359] For example, the first method and the second method may be intra prediction and inter prediction, respectively, and the predictor may select a final prediction image for the current block from prediction images generated according to the prediction methods.
[0360] Figure 30 is a flowchart illustrating another example of a process performed by the predictor of the encoder 100 .
[0361] First, the predictor generates a predicted image using intra prediction (step Sd_1a) and generates a predicted image using inter prediction (step Sd_1b). Note that a predicted image generated by intra prediction is also referred to as an intra-predicted image, and a predicted image generated by inter prediction is also referred to as an inter-predicted image.
[0362] Next, the predictor evaluates each of the intra-frame prediction image and the inter-frame prediction image (step Sd_2). The cost C described above can be used in the evaluation. The predictor can then select the prediction image for which the minimum cost C has been calculated from the intra-frame prediction image and the inter-frame prediction image as the final prediction image for the current block (step Sd_3). In other words, the prediction method or mode for generating the prediction image for the current block is selected.
[0363] (Intra-frame predictor)
[0364] The intra predictor 124 generates a prediction signal (i.e., an intra-predicted image) by performing intra-prediction (also referred to as intra-prediction) of the current block by referring to one or more blocks in the current picture and stored in the block memory 118. More specifically, the intra predictor 124 generates an intra-predicted image by performing intra-prediction by referring to pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block, and then outputs the intra-predicted image to the prediction controller 128.
[0365] For example, the intra-frame predictor 124 performs intra-frame prediction by using one of a plurality of defined intra-frame prediction modes. The intra-frame prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes. The defined modes may be predefined.
[0366] The one or more non-directional prediction modes include planar prediction mode and DC prediction mode, such as defined in the H.265 / High Efficiency Video Coding (HEVC) standard.
[0367] The plurality of directional prediction modes include, for example, the thirty-three directional prediction modes defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes may include thirty-two directional prediction modes in addition to the thirty-three directional prediction modes (a total of sixty-five directional prediction modes). Figure 31: is a conceptual diagram showing a total of sixty-seven intra prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra prediction. Solid arrows represent thirty-three directions defined in the H.265 / HEVC standard, and dotted arrows represent additional thirty-two directions ( Figure 31 The two non-directional prediction modes are not shown).
[0368] In various types of processing examples, a luma block can be referenced in intra prediction of a chroma block. In other words, the chroma component of the current block can be predicted based on the luma component of the current block. This intra prediction is also called cross-component linear model (CCLM) prediction. An intra prediction mode for a chroma block that references this luma block (also called, for example, a CCLM mode) can be added as one of the intra prediction modes for the chroma block.
[0369] The intra predictor 124 can correct the pixel values of the intra prediction based on the horizontal / vertical reference pixel gradients. Intra prediction with such correction is also called position-dependent intra prediction combining (PDPC). Information indicating whether PDPC is applied (e.g., called a PDPC flag) is typically signaled at the CU level. Note that signaling such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0370] Figure 32 is a flowchart illustrating one example of a process performed by the intra predictor 124 .
[0371] The intra-frame predictor 124 selects an intra-frame prediction mode from a plurality of intra-frame prediction modes (step Sw_1). The intra-frame predictor 124 then generates a predicted image based on the selected intra-frame prediction mode (step Sw_2). Next, the intra-frame predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra-frame prediction modes. For example, two of the six intra-frame prediction modes may be a planar mode and a DC prediction mode, and the other four modes may be directional prediction modes. The intra-frame predictor 124 determines whether the intra-frame prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0372] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPM ("Yes" in step Sw_4), the intra predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). Note that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110.
[0373] When it is determined that the selected intra prediction mode is not included in the MPM ("No" in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flag. The intra predictor 124 then generates information indicating the selected intra prediction mode that is not included in the MPM among at least one intra prediction mode (step Sw_8). Note that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110. The information indicating the intra prediction mode indicates, for example, any one of 0 to 60.
[0374] (Inter-frame predictor)
[0375] The inter-frame predictor 126 generates a predicted image (inter-frame prediction image) by performing inter-frame prediction (also referred to as inter-frame prediction) on the current block by referring to one or more blocks in a reference picture, which is different from the current picture and stored in the frame memory 122. Inter-frame prediction is performed in units of the current block or the current sub-block (for example, a 4×4 block) in the current block. A sub-block is included in a block and is a unit smaller than a block. The size of the sub-block can be in the form of a slice, a brick, a picture, etc.
[0376] For example, the inter-frame predictor 126 performs motion estimation on the current block or current sub-block in the reference picture and finds the reference block or reference sub-block that best matches the current block or current sub-block. The inter-frame predictor 126 then obtains motion information (e.g., a motion vector) that compensates for the motion or change from the reference block or reference sub-block to the current block or sub-block. The inter-frame predictor 126 generates an inter-frame predicted image for the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter-frame predictor 126 outputs the generated inter-frame predicted image to the prediction controller 128.
[0377] The motion information used in motion compensation can be signaled in various forms as an inter-frame prediction signal. For example, a motion vector can be signaled. As another example, the difference between a motion vector and a motion vector predictor can be signaled.
[0378] (Reference picture list)
[0379] Figure 33 is a conceptual diagram for illustrating an example of a reference picture. Figure 34 1 is a conceptual diagram for illustrating an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. Note that Figure 33In FIG, each rectangle indicates a picture, each arrow indicates a picture reference relationship, the horizontal axis indicates time, I, P, and B in the rectangle indicate an intra-frame prediction picture, a single prediction picture, and a bi-prediction picture, respectively, and the numbers in the rectangle indicate the decoding order. Figure 33 As shown in FIG, the decoding order of pictures is the order of I0, P1, B2, B3 and B4, and the display order of pictures is the order of I0, B3, B2, B4 and P1. Figure 34 As shown in , a reference picture list is a list indicating reference picture candidates. For example, a picture (or slice) may include at least one reference picture list. For example, when the current picture is a uni-predicted picture, one reference picture list is used, and when the current picture is a bi-predicted picture, two reference picture lists are used. Figure 33 and Figure 34 In the example of , picture B3, which is the current picture currPic, has two reference picture lists, namely, the L0 list and the L1 list. When the current picture currPic is picture B3, the reference picture candidates for the current picture currPic are I0, P1, and B2, and the reference picture lists (i.e., the L0 list and the L1 list) indicate these pictures. The inter predictor 126 or the prediction controller 128 specifies which picture in each reference picture list to actually reference in the form of a reference picture index refidxLx. Figure 34 , reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.
[0380] Such a reference picture list may be generated for each unit such as a sequence, a picture, a slice, a tile, a CTU, or a CU. Additionally, among the reference pictures indicated in the reference picture list, a reference picture index indicating a reference picture to be referenced in inter prediction may be signaled at the sequence level, the picture level, the slice level, the tile level, the CTU level, or the CU level. Additionally, a common reference picture list may be used in a variety of inter prediction modes.
[0381] (Basic process of inter-frame prediction)
[0382] Figure 35 is a flowchart illustrating an example basic processing flow of the process of inter-frame prediction.
[0383] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates a difference between the current block and the predicted image as a prediction residual (step Se_4).
[0384] Here, when generating a predicted image, the inter-frame predictor 126 determines the motion vector (MV) of the current block (steps Se_1 and Se_2) and performs motion compensation to generate the predicted image (step Se_3). Furthermore, when determining the MV, the inter-frame predictor 126 determines the MV by selecting motion vector candidates (MV candidates) (step Se_1) and deriving the MV (step Se_2). The selection of the MV candidate is performed, for example, by the inter-frame predictor 126 generating an MV candidate list and selecting at least one MV candidate from the MV candidate list. Note that previously derived MVs may be added to the MV candidate list. Alternatively, when deriving the MV, the inter-frame predictor 126 may select at least one MV candidate from the at least one MV candidate and determine the selected at least one MV candidate as the MV for the current block. Alternatively, the inter-frame predictor 126 may determine the MV for the current block by performing estimation in a reference picture region specified by each of the at least one selected MV candidate. Note that estimation in a reference picture region may be referred to as motion estimation.
[0385] Additionally, although in the example described above, steps Se_1 to Se_3 are performed by the inter predictor 126 , processes such as step Se_1 , step Se_2 , and the like may be performed by another constituent element included in the encoder 100 .
[0386] Note that the MV candidate list can be generated for each process in the inter prediction mode, or a common MV candidate list can be used in multiple inter prediction modes. The processes in steps Se_3 and Se_4 correspond to Figure 9 The process in step Se_3 corresponds to Figure 30 The process in step Sd_1b in .
[0387] (Motion vector derivation process)
[0388] Figure 36 is a flowchart showing an example of a process of derivation of a motion vector.
[0389] The inter-frame predictor 126 can derive the MV of the current block in a mode for encoding motion information (e.g., MV). In this case, for example, the motion information can be encoded as a prediction parameter and can be signaled. In other words, the encoded motion information is included in the stream.
[0390] Alternatively, the inter predictor 126 may derive the MV in a mode where motion information is not encoded. In this case, motion information is not included in the stream.
[0391] Here, the MV derivation mode may include the normal inter mode, normal merge mode, FRUC mode, affine mode, and the like, which will be described later. Among the modes, the modes in which motion information is encoded include the normal inter mode, normal merge mode, affine mode (specifically, affine inter mode and affine merge mode), and the like. Note that the motion information may include not only the MV but also the motion vector predictor selection information described later. Modes in which motion information is not encoded include the FRUC mode, and the like. The inter predictor 126 selects a mode for deriving the MV of the current block from a plurality of modes and uses the selected mode to derive the MV of the current block.
[0392] Figure 37 is a flowchart illustrating another example of derivation of motion vectors.
[0393] The inter-frame predictor 126 can derive the MV for the current block in a mode where the MV difference is encoded. In this case, for example, the MV difference can be encoded as a prediction parameter and can be signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. Note that the MV predictor is a motion vector predictor.
[0394] Alternatively, the inter-frame predictor 126 may derive the MV in a mode where the MV difference is not encoded. In this case, the encoded MV difference is not included in the stream.
[0395] Here, as described above, the MV derivation mode includes the normal inter mode, normal merge mode, FRUC mode, affine mode, etc., which will be described later. Among the modes, the mode in which the MV difference is encoded includes the normal inter mode, the affine mode (specifically, the affine inter mode), etc. The mode in which the MV difference is not encoded includes the FRUC mode, the normal merge mode, the affine mode (specifically, the affine merge mode), etc. The inter predictor 126 selects a mode for deriving the MV of the current block from the plurality of modes, and uses the selected mode to derive the MV of the current block.
[0396] (Motion vector derivation mode)
[0397] Figure 38A and Figure 38B is a conceptual diagram for illustrating an example classification of patterns for MV derivation. Figure 38A As shown in FIG, the MV derivation mode is roughly divided into three modes according to whether motion information is encoded and whether MV difference is encoded. These three modes are inter mode, merge mode and frame rate up conversion (FRUC) mode. Inter mode is a mode in which motion estimation is performed and motion information and MV difference are encoded. For example, Figure 38BAs shown in , inter mode includes affine inter mode and normal inter mode. Merge mode is a mode in which motion estimation is not performed, and in which MV is selected from coded surrounding blocks and MV for the current block is derived using the MV. Merge mode is a mode in which motion information is basically encoded and MV difference is not encoded. For example, Figure 38B As shown in , the merge mode includes a normal merge mode (also referred to as a normal merge mode or a regular merge mode), a motion vector difference merge (MMVD) mode, a combined inter merge / intra prediction (CIIP) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, among the modes included in the merge mode, the MV difference is encoded abnormally in the MMVD mode. Note that the affine merge mode and the affine inter mode are modes included in the affine mode. The affine mode is a mode for deriving the MV of each of a plurality of sub-blocks included in the current block as the MV of the current block assuming an affine transformation. The FRUC mode is a mode for deriving the MV of the current block by performing estimation between the encoded regions, and in which neither motion information nor any MV difference is encoded. Note that the corresponding modes will be described in more detail later.
[0398] Notice, Figure 38A and Figure 38B The classification of the modes shown in is an example, and the classification is not limited thereto. For example, when the MV difference is encoded in the CIIP mode, the CIIP mode is classified as the inter mode.
[0399] (MV derivation > normal inter-frame mode)
[0400] Normal inter mode is an inter prediction mode for deriving the MV of the current block from a reference picture region specified by an MV candidate based on a block similar to an image of the current block. In this normal inter mode, the MV difference is encoded.
[0401] Figure 39 is a flowchart illustrating an example of a process of inter prediction in normal inter mode.
[0402] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (eg, MVs of multiple encoded blocks temporally or spatially surrounding the current block) (step Sg_1). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0403] Next, the inter-frame predictor 126 extracts N (an integer of 2 or greater) MV candidates from the plurality of MV candidates obtained in step Sg_1 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the determined priority order (step Sg_2). Note that the priority order may be predetermined for each of the N MV candidates.
[0404] Next, the inter-frame predictor 126 selects a motion vector predictor candidate from the N predicted motion vector candidates as the motion vector predictor (also called MV predictor) for the current block (step Sg_3). At this time, the inter-frame predictor 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor in the stream. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.
[0405] Next, the inter-frame predictor 126 derives the MV of the current block by referring to the coded reference picture (step Sg_4). At this time, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as an MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130. Note that the coded reference picture is a picture that includes a plurality of blocks that have been reconstructed after encoding.
[0406] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The process in steps Sg_1 to Sg_5 is performed on each block. For example, when the process in steps Sg_1 to Sg_5 is performed on all blocks in the slice, inter-frame prediction of the slice using the normal inter-frame mode is completed. For example, when the process in steps Sg_1 to Sg_5 is performed on all blocks in the picture, inter-frame prediction of the picture using the normal inter-frame mode is completed. Note that not all blocks included in the slice may undergo the process in steps Sg_1 to Sg_5, and when a portion of the blocks undergo the process, inter-frame prediction of the slice using the normal inter-frame mode may be completed. This also applies to the process in steps Sg_1 to Sg_5. When the process is performed on a portion of the blocks in the picture, inter-frame prediction of the picture using the normal inter-frame mode may be completed.
[0407] Note that the predicted image is an inter prediction signal as described above. Additionally, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the predicted image is encoded as a prediction parameter in the encoded signal, for example.
[0408] Note that the MV candidate list can also be used as a list for use in other modes. Additionally, processes related to the MV candidate list can be applied to processes related to lists used in other modes. Processes related to the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering MV candidates, or deleting MV candidates.
[0409] (MV derivation > normal merge mode)
[0410] Normal merge mode is used to select an MV candidate from the MV candidate list as the MV of the current block, thereby deriving the inter-frame prediction mode of the MV. Note that normal merge mode is a type of merge mode and can be simply referred to as merge mode. In this embodiment, normal merge mode and merge mode are distinguished, and merge mode is used in a broader sense.
[0411] Figure 40 is a flowchart illustrating an example of inter prediction in normal merge mode.
[0412] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (eg, MVs of multiple encoded blocks temporally or spatially surrounding the current block) (step Sh_1). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0413] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter-frame predictor 126 encodes MV selection information used to identify the selected MV candidate in the stream. In other words, the inter-frame predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.
[0414] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). For example, the process in steps Sh_1 to Sh_3 is performed on each block. For example, when the process in steps Sh_1 to Sh_3 is performed on all blocks in the slice, inter-frame prediction of the slice using normal merge mode is completed. Alternatively, when the process in steps Sh_1 to Sh_3 is performed on all blocks in the picture, inter-frame prediction of the picture using normal merge mode is completed. Note that not all blocks included in the slice can undergo the process in steps Sh_1 to Sh_3, and when a portion of the blocks undergo the process, inter-frame prediction of the slice using normal merge mode may be completed. This also applies to the process in steps Sh_1 to Sh_3. When the process is performed on a portion of the blocks in the picture, inter-frame prediction of the picture using normal merge mode may be completed.
[0415] Additionally, for example, information indicating an inter prediction mode (normal merge mode in the above example) used to generate a predicted image and included in an encoded signal is encoded as a prediction parameter in a stream.
[0416] Figure 41 is a conceptual diagram for illustrating an example of a motion vector derivation process for a current picture through a normal merge mode.
[0417] First, the inter-frame predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of multiple coded blocks spatially located around the current block; temporally adjacent MV candidate lists, which are MVs of surrounding blocks on which the position of the current block in the coded reference picture is projected; combined MV candidates, which are MVs generated by combining the MV values of spatially adjacent MV predictors and the MV values of temporally adjacent MV predictors; and zero MV candidates, which are MVs with a value of zero.
[0418] Next, the inter predictor 126 selects one MV candidate from a plurality of MV candidates registered in the MV candidate list, and determines the MV candidate as the MV of the current block.
[0419] Furthermore, the entropy encoder 110 writes and encodes merge_idx, which is a signal indicating which MV candidate has been selected, in the stream.
[0420] Note that in Figure 41 The MV candidates registered in the MV candidate list described in the figure are examples. The number of MV candidates may be different from the number of MV candidates in the figure, and the MV candidate list may be configured in such a way that some of the categories of MV candidates in the figure may not be included, or one or more MV candidates other than the categories of MV candidates in the figure may be included.
[0421] The final MV can be determined by performing dynamic motion vector refresh (DMVR) described later using the MV of the current block derived by the normal merge mode. Note that in the normal merge mode, motion information is encoded and the MV difference is not encoded. In the MMVD mode, an MV candidate is selected from the MV candidate list, and the MV difference is encoded as in the case of the normal merge mode. Figure 38B As shown in
[15] , MMVD can be classified as a merge mode along with the normal merge mode. Note that the MV difference in MMVD mode does not always need to be the same as the MV difference used in inter mode. For example, MV difference derivation in MMVD mode can be a process that requires less processing than MV difference derivation in inter mode.
[0422] Additionally, a combined inter-merging / intra-prediction (CIIP) mode may be performed, which is used to overlap a prediction image generated in inter-prediction and a prediction image generated in intra-prediction to generate a prediction image for a current block.
[0423] Note that the MV candidate list may be referred to as a candidate list. Additionally, merge_idx is MV selection information.
[0424] (MV derivation > HMVP mode)
[0425] Figure 42 is a conceptual diagram illustrating an example of an MV derivation process for a current picture using the HMVP merge mode.
[0426] In normal merge mode, the MV for a CU, such as the current block, is determined by selecting an MV candidate from an MV list generated with reference to a coded block (e.g., a CU). Here, another MV candidate can be registered in the MV candidate list. The mode in which such another MV candidate is registered is called HMVP mode.
[0427] In HMVP mode, the MV candidates are managed using the HMVP's First-In-First-Out (FIFO) server, separate from the MV candidate list used for normal merge mode.
[0428] In the FIFO buffer, the latest motion information such as the MV of the block processed in the past is stored first. When managing the FIFO buffer, each time a block is processed, the MV for the latest block (i.e., the CU that was just processed previously) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the earliest processed CU) is deleted from the FIFO buffer. Figure 42 In the example shown in , HMVP1 is the MV for the newest block, and HMVP5 is the MV for the oldest MV.
[0429] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is an MV different from all MV candidates already registered in the MV candidate list for normal merge mode, starting from HMVP1. If it is determined that the MV is different from all MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for normal merge mode as an MV candidate. At this time, one or more of the MV candidates in the FIFO buffer can be registered (added to the MV candidate list).
[0430] By using the HMVP mode in this way, not only can the MVs of blocks adjacent to the current block in space or time be added, but also the MVs of blocks processed in the past can be added. As a result, the variety of MV candidates for normal merge mode is expanded, which increases the possibility of improving decoding efficiency.
[0431] Note that the MV may be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer may include not only the MV value but also reference picture information, reference direction, number of pictures, etc. Alternatively, the block may be, for example, a CU.
[0432] Notice, Figure 42 The MV candidate list and FIFO buffer shown in FIG are examples. The size of the MV candidate list and FIFO buffer can be Figure 42 or can be configured to be different from Figure 42 MV candidates are registered in an order different from the order in . Additionally, the process described herein may be common between the encoder 100 and the decoder 200 .
[0433] Note that the HMVP mode can be applied to modes other than the normal merge mode. For example, motion information such as the MV of a block processed in the past in the affine mode can also be stored first and used as an MV candidate, which can further promote efficiency. The mode obtained by applying the HMVP mode to the affine mode can be called the historical affine mode.
[0434] (MV derivation > FRUC mode)
[0435] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, motion estimation is performed on the decoder side without using any pixel values in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values in the current block include frame rate up conversion (FRUC) mode, pattern matching motion vector derivation (PMMVD) mode, and the like.
[0436] Figure 43 An example of a FRUC process in the form of a flowchart is shown in FIG. First, a list indicates MVs for coded blocks (each of the coded blocks is spatially or temporally adjacent to the current block) as MV candidates by referring to the MV (the list can be an MV candidate list and can also be used as an MV candidate list for normal merge mode) (step Si_1).
[0437] Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation values of the corresponding MV candidates included in the MV candidate list are calculated, and an MV candidate is selected based on the evaluation values. Based on the selected motion vector candidate, the motion vector for the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector for the current block. Alternatively, for example, the motion vector for the current block can be derived using pattern matching in a surrounding area of a position in a reference picture, where the position in the reference picture corresponds to the selected motion vector candidate. In other words, estimation using pattern matching and evaluation values can be performed in a surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to an MV that produces a better evaluation value, and the updated MV can be determined as the final MV for the current block. In some embodiments, updating of the motion vector that produces a better evaluation value may not be performed.
[0438] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). For example, the process in steps Si_1 to Si_5 is performed on each block. For example, when the process in steps Si_1 to Si_5 is performed on all blocks in the slice, the inter-frame prediction of the slice using the FRUC mode is completed. For example, when the process in steps Si_1 to Si_5 is performed on all blocks in the picture, the inter-frame prediction of the picture using the FRUC mode is completed. Note that not all blocks included in the slice can undergo the process of steps Si_1 to Si_5, and when a part of the blocks undergoes the process, the inter-frame prediction of the slice using the FRUC mode may be completed. When the process in steps Si_1 to Si_5 is performed on a part of the blocks included in the picture in a similar manner, the inter-frame prediction of the picture using the FRUC mode may be completed.
[0439] A similar process may be performed in units of sub-blocks.
[0440] The evaluation value can be calculated using various methods. For example, a comparison can be made between a reconstructed image of a region in a reference picture corresponding to a motion vector and a reconstructed image in a determined region (which may be, for example, a region in another reference picture or a region in an adjacent block of the current picture, as described below). The determined region may be predetermined.
[0441] The difference between the pixel values of the two reconstructed images can be used for the evaluation value of the motion vector. Note that the evaluation value can be calculated using information other than the value of the difference.
[0442] Next, an example of pattern matching is described in detail. First, an MV candidate included in an MV candidate list (e.g., a merge list) is selected as a starting point for estimation through pattern matching. For example, first pattern matching or second pattern matching can be used as pattern matching. First pattern matching and second pattern matching can be referred to as bilateral matching and template matching, respectively.
[0443] (MV derivation > FRUC > bilateral matching)
[0444] In the first pattern matching, pattern matching is performed between two blocks located along the motion trajectory of the current block and included in two different reference pictures. Therefore, in the first pattern matching, the area along the motion trajectory of the current block in the other reference picture is used as the determination area for calculating the candidate evaluation value described above. The determination area can be predetermined.
[0445] Figure 44 : is a conceptual diagram for illustrating an example of first pattern matching (bilateral matching) between two blocks along a motion trajectory in two reference pictures. Figure 44 As shown in , in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the best matching pair between the pairs in two blocks included in two different reference pictures (Ref0, Ref1) and located along the motion trajectory of the current block (Cur block). More specifically, the difference between the reconstructed image at a specified position in the first coded reference picture (Ref0) specified by the MV candidate and the reconstructed image at a specified position in the second coded reference picture (Ref1) specified by the symmetrical MV obtained by scanning the MV candidates at the display time interval is derived for the current block, and the obtained difference value is used to calculate the evaluation value. The MV candidate that produces the best evaluation value and is likely to produce good results can be selected as the final MV among multiple MV candidates.
[0446] Under the assumption of a continuous motion trajectory, the motion vectors (MV0, MV1) of the two reference blocks are specified to be proportional to the temporal distance (TD0, TD1) between the current picture (CurPic) and the two reference pictures (Ref0, Ref1). For example, when the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, a mirror-symmetric bidirectional motion vector is derived in the first pattern matching.
[0447] (MV derivation > FRUC > template matching)
[0448] In the second pattern matching (template matching), pattern matching is performed between a block in a reference picture and a template in the current picture (the template is a block adjacent to the current block in the current picture (the adjacent block is, for example, an upper and / or left adjacent block)). Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a determination area for calculating the evaluation value of the MV candidate described above.
[0449] Figure 45 is a conceptual diagram for illustrating an example of pattern matching (template matching) between a template in a current picture and a block in a reference picture. Figure 45 As shown in , in the second mode matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference picture (Ref0) that best matches the block adjacent to the current block in the current picture (Cur Pic). More specifically, the difference between the reconstructed image in the coded area (which is adjacent to the upper left or adjacent to the left or above) and the reconstructed image in the corresponding area in the coded reference picture (Ref0) and specified by the MV candidate is derived, and the obtained difference value is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.
[0450] Such information indicating whether the FRUC mode is applied (e.g., referred to as a FRUC flag) may be signaled at the CU level. Alternatively, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (e.g., first pattern matching or second pattern matching) may be signaled at the CU level. Note that signaling such information does not necessarily need to be performed at the CU level and may be performed at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0451] (MV Derivation > Affine Mode)
[0452] Affine mode is a mode for generating MVs using affine transformation. For example, MVs can be derived in sub-block units based on motion vectors of multiple adjacent blocks. This mode is also called affine motion compensation prediction mode.
[0453] Figure 46A : is a conceptual diagram for illustrating an example of MV derivation in sub-block units based on motion vectors of multiple adjacent blocks. Figure 46AIn the example, the current block includes, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the upper left control point in the current block is derived based on the motion vectors of the adjacent blocks, and similarly, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of the adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (v x ,v y ).
[0454] [Mathematics 1]
[0455]
[0456] Here, x and y indicate the horizontal position and vertical position of the sub-block, respectively, and w indicates a determined weighting coefficient. The determined weighting coefficient may be predetermined.
[0457] Such information indicating the affine mode (e.g., referred to as an affine flag) may be signaled at the CU level. It should be noted that signaling the information indicating the affine mode does not necessarily need to be performed at the CU level and may also be performed at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0458] Additionally, the affine mode may include several modes of different methods for deriving motion vectors at the upper left control point and the upper right control point. For example, the affine mode includes two modes: an affine inter mode (also referred to as an affine normal inter mode) and an affine merge mode.
[0459] (MV Derivation > Affine Mode)
[0460] Figure 46B is a conceptual diagram for illustrating an example of MV derivation in units of sub-blocks in an affine mode in which three control points are used. Figure 46B In the example, the current block includes sixteen 4×4 blocks. Here, the motion vector V0 at the upper left control point in the current block is derived based on the motion vectors of the adjacent blocks. Here, the motion vector V1 at the upper right control point in the current block is derived based on the motion vectors of the adjacent blocks, and similarly, the motion vector V2 at the lower left control point of the current block is derived based on the motion vectors of the adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v x ,v y ).
[0461] [Mathematics 2]
[0462]
[0463] Here, x and y indicate the horizontal position and vertical position of the sub-block, respectively, and w and h may be weighting coefficients, which may be predetermined weighting coefficients. In an embodiment, w may indicate the width of the current block, and h may indicate the height of the current block.
[0464] Affine modes in which different numbers of control points (e.g., two and three control points) are used can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0465] Additionally, the affine mode using three control points may include different methods for deriving motion vectors at the upper left control point, the upper right control point, and the lower left control point. For example, as in the case of the affine mode using two control points, the affine mode using three control points may include two modes: the affine inter mode and the affine merge mode.
[0466] Note that in the affine mode, the size of each sub-block included in the current block may not be limited to 4×4 pixels, and may also be another size. For example, the size of each sub-block may be 8×8 pixels.
[0467] (MV Derivation > Affine Mode > Control Points)
[0468] Figure 47A 、 Figure 47B and Figure 47C is a conceptual diagram for illustrating an example of MV derivation at a control point in the affine mode.
[0469] like Figure 47A As shown in FIG, in affine mode, for example, a motion vector predictor at a corresponding control point of a current block is calculated based on a plurality of motion vectors corresponding to blocks coded according to the affine mode between coded blocks A (left), B (above), C (top right), D (bottom left), and E (top left) adjacent to the current block. More specifically, coded blocks A (left), B (above), C (top right), D (bottom left), and E (top left) are examined in the order listed, and the first valid block coded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on a plurality of motion vectors corresponding to the identified blocks.
[0470] For example, Figure 47B, when a block A adjacent to the current block on the left has been encoded according to an affine mode in which two control points are used, motion vectors v3 and v4 projected at the upper left corner position and the upper right corner position of the encoded block including the block A are derived. A motion vector v0 at the upper left control point of the current block and a motion vector v1 at the upper right control point of the current block are then calculated based on the derived motion vectors v3 and v4.
[0471] For example, Figure 47C , when a block A adjacent to the current block on the left has been encoded according to an affine mode in which three control points are used, motion vectors v3, v4, and v5 projected at the upper left corner position, the upper right corner position, and the lower left corner position of the encoded block including the block A are derived. Then, a motion vector v0 at the upper left control point of the current block, a motion vector v1 at the upper right control point of the current block, and a motion vector v2 at the lower left control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0472] Figures 47A to 47C The MV derivation method shown in can be used to Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in FIG, or can be used for the later described Figure 51 The MV predictor at each control point of the current block in step Sj_1 is deduced as shown in FIG.
[0473] Figure 48A and Figure 48B is a conceptual diagram for illustrating an example of MV derivation at a control point in the affine mode.
[0474] Figure 48A is a conceptual diagram for illustrating an example affine pattern in which two control points are used.
[0475] In affine mode, such as Figure 48A As shown in FIG, an MV selected from the MVs at the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, an MV selected from the MVs at the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block.
[0476] Figure 48B is a conceptual diagram for illustrating an example affine pattern in which three control points are used.
[0477] In affine mode, such as Figure 48BAs shown in FIG, an MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the upper left corner control point of the current block. Similarly, an MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the upper right corner control point of the current block. Furthermore, an MV selected from the MVs of the coded blocks F and G adjacent to the current block is used as the motion vector v2 at the lower left corner control point of the current block.
[0478] Notice, Figure 48A and Figure 48B The MV derivation method shown in FIG can be used for the MV derivation method described later. Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in FIG, or can be used for the later described Figure 51 The MV predictor at each control point of the current block in step Sj_1 is deduced as shown in FIG.
[0479] Here, when affine modes in which different numbers of control points (e.g., two and three control points) are used can be switched and signaled at the CU level, the number of control points of the encoded block and the number of control points of the current block can be different from each other.
[0480] Figure 49A and Figure 49B is a conceptual diagram for illustrating an example of a method for MV derivation at a control point when the number of control points of an encoded block and the number of control points of a current block are different from each other.
[0481] For example, Figure 49A As shown in , the current block has three control points at the upper left corner, upper right corner, and lower left corner, and the block A adjacent to the current block on the left has been encoded according to an affine mode in which two control points are used. In this case, motion vectors v3 and v4 projected at the upper left corner position and the upper right corner position in the encoded block including block A are derived. Then, the motion vector v0 at the upper left control point and the motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3 and v4. In addition, the motion vector v2 at the lower left control point is calculated based on the derived motion vectors v0 and v1.
[0482] For example, Figure 49BAs shown in , the current block has two control points at the upper left and upper right corners, and block A adjacent to the current block on the left has been encoded according to an affine mode in which three control points are used. In this case, motion vectors v3, v4, and v5 projected at the upper left corner position in the encoded block including block A, the upper right corner position in the encoded block, and the lower left corner position in the encoded block are derived. Then, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0483] Notice, Figure 49A and Figure 49B The MV derivation method shown in FIG can be used in the later described Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in FIG, or can be used for the later described Figure 51 The MV predictor at each control point of the current block in step Sj_1 is deduced as shown in FIG.
[0484] (MV derivation > affine mode > affine merge mode)
[0485] Figure 50 is a flowchart showing one example of a process in affine merge mode.
[0486] In the affine merge mode as shown, first, the inter-frame predictor 126 derives the MV at the corresponding control point of the current block (step Sk_1). Figure 46A As shown in , the control points are the upper left corner point of the current block and the upper right corner point of the current block, or as Figure 46B As shown in , the control points are the top left corner point of the current block, the top right corner point of the current block, and the bottom left corner point of the current block. The inter predictor 126 may encode MV selection information to identify two or three derived MVs in the stream.
[0487] For example, when using Figures 47A to 47C When the MV derivation method shown in Figure 47A As shown in FIG, the inter-frame predictor 126 examines the encoded blocks A (left), B (above), C (above right), D (below left), and E (above left) in the order listed and identifies the first valid block that is encoded according to the affine mode.
[0488] The inter-frame predictor 126 uses the first identified valid block that is encoded according to the identified affine mode to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as Figure 47B, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left corner of the coded block and the upper right corner of the coded block including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block and the motion vector v1 at the upper right control point of the current block by projecting the motion vectors v3 and v4 at the upper left corner and the upper right corner of the coded block onto the current block.
[0489] Alternatively, when block A is identified and block A has three control points, as Figure 47C , the inter-frame predictor 126 calculates a motion vector v0 at the upper left control point of the current block, a motion vector v1 at the upper right control point of the current block, and a motion vector v2 at the lower left control point of the current block based on the motion vectors v3, v4, and v5 of the upper left corner of the coded block, the upper right corner of the coded block, and the lower left corner of the coded block including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the upper left control point of the current block, the motion vector v1 at the upper right control point of the current block, and the motion vector v2 at the lower left control point of the current block by projecting the motion vectors v3, v4, and v5 at the upper left corner, the upper right corner, and the lower left corner of the coded block onto the current block.
[0490] Note that, as described above Figure 49A As shown in FIG, when block A is identified and block A has two control points, the MV at three control points can be calculated, and as described above Figure 49B As shown in , when block A is identified and block A has three control points, MVs at two control points can be calculated.
[0491] Next, the inter-frame predictor 126 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the MV for each of the multiple sub-blocks as an affine MV using, for example, two motion vectors v0 and v1 and the above expression (1A) or using three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-blocks (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all sub-blocks included in the current block, the process of generating a predicted image using the affine merge mode for the current block ends. In other words, motion compensation is performed on the current block to generate a predicted image for the current block.
[0492] Note that the MV candidate list described above may be generated in step Sk_1. The MV candidate list may be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown in Figure 49A and Figure 49B Any combination of the MV derivation method shown in and other MV derivation methods.
[0493] Note that the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks other than the affine mode.
[0494] Note that, for example, an MV candidate list including MV candidates in an affine merge mode in which two control points are used and an affine merge mode in which three control points are used may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in an affine merge mode in which two control points are used and an MV candidate list including MV candidates in an affine merge mode in which three control points are used may be generated separately. Alternatively, an MV candidate list including MV candidates in one of an affine merge mode in which two control points are used and an affine merge mode in which three control points are used may be generated. The (multiple) MV candidates may be, for example, MVs for the encoded block A (left side), block B (above), block C (top right), block D (bottom left), and block E (top left), or MVs for valid blocks in the block.
[0495] Note that an index indicating one of the MVs in the MV candidate list may be transmitted as the MV selection information.
[0496] (MV derivation > affine mode > affine inter-frame mode)
[0497] Figure 51 is a flowchart showing one example of a process in affine inter mode.
[0498] In the affine inter mode, first, the inter predictor 126 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block, such as Figure 46A or Figure 46B As shown in .
[0499] For example, when using Figure 48A and Figure 48B When the MV derivation method shown in FIG. 1 is used, the inter-frame predictor 126 selects Figure 48A or Figure 48B The inter-frame predictor 126 derives the MV predictor (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block by using the MV of any one of the coded blocks near the corresponding control point of the current block shown in FIG. In this case, the inter-frame predictor 126 encodes MV predictor selection information for identifying the selected two or three MV predictors in the stream.
[0500] For example, the inter-frame predictor 126 may determine, using cost evaluation or the like, which block to select as the MV predictor at the control point from the coded blocks adjacent to the current block, and may write a flag indicating which MV predictor has been selected in the bitstream. In other words, the inter-frame predictor 126 outputs MV predictor selection information such as the flag as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.
[0501] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while simultaneously updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 uses Expression (1A) or Expression (1B) described above to calculate the MV for each subblock corresponding to the updated MV predictor as an affine MV (step Sj_3). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation on the subblocks (step Sj_4). When the MV predictor is updated in step Sj_2, the process in steps Sj_3 and Sj_4 is performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor in the stream as an MV difference. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130 .
[0502] Finally, the inter predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0503] Note that the MV candidate list described above may be generated in step Sj_1. The MV candidate list may be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown in Figure 49A and Figure 49BAny combination of the MV derivation method shown in and other MV derivation methods.
[0504] Note that the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks other than the affine mode.
[0505] Note that, for example, an MV candidate list including MV candidates in an affine inter mode in which two control points are used and an affine inter mode in which three control points are used may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in an affine inter mode in which two control points are used and an MV candidate list including MV candidates in an affine inter mode in which three control points are used may be generated separately. Alternatively, an MV candidate list including MV candidates in one of an affine inter mode in which two control points are used and an affine inter mode in which three control points are used may be generated. The (multiple) MV candidates may be, for example, MVs for the encoded block A (left side), block B (above), block C (top right), block D (bottom left), and block E (top left), or MVs for valid blocks in the block.
[0506] Note that an index indicating one of the MVs in the MV candidate list may be transmitted as MV predictor selection information.
[0507] (MV derivation > triangle pattern)
[0508] In the above example, the inter-frame predictor 126 generates a single rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 may generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and may combine the multiple prediction images to generate a final rectangular prediction image. A shape other than a rectangle may be, for example, a triangle.
[0509] Figure 52A It is a conceptual diagram for illustrating the generation of two triangular prediction images.
[0510] The inter-frame predictor 126 performs motion compensation on the first partition having a triangular shape in the current block using the first MV of the first partition to generate a triangular predicted image. Similarly, the inter-frame predictor 126 performs motion compensation on the second partition having a triangular shape in the current block using the second MV of the second partition to generate a triangular predicted image. Then, the inter-frame predictor 126 combines these predicted images to generate a predicted image having the same rectangular shape as the current block.
[0511] Note that a first predicted image having a rectangular shape corresponding to the current block can be generated using the first MV as the predicted image for the first partition. Alternatively, a second predicted image having a rectangular shape corresponding to the current block can be generated using the second MV as the predicted image for the second partition. The predicted image for the current block can be generated by performing a weighted addition of the first predicted image and the second predicted image. Note that the portion subjected to the weighted addition may be a portion of the region that straddles the boundary between the first and second partitions.
[0512] Figure 52B is a conceptual diagram illustrating an example of a first portion of a first partition that overlaps with a second partition, and first and second sets of samples that may be weighted as part of a correction process. The first portion may have, for example, one-quarter the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to an edge of the first partition, where N is an integer greater than zero, for example, N may be the integer 2. As shown, Figure 52B The left example of shows a rectangular partition having a rectangular portion whose width is one quarter the width of the first partition, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion. Figure 52B The center example shows a rectangular partition having a rectangular portion whose height is one quarter of the height of the first partition, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion. Figure 52B The right example of shows a triangular partition with a polygonal portion whose height corresponds to two samples, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion.
[0513] The first portion may be a portion of the first partition that overlaps with an adjacent partition. Figure 52C This is a conceptual diagram illustrating a first portion of a first partition, which is a portion of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition is shown with an overlapping portion with a spatially adjacent rectangular partition. Partitions of other shapes (e.g., triangular partitions) may be used, and the overlapping portion may overlap with spatially or temporally adjacent partitions.
[0514] Additionally, although an example has been given in which a predicted image for each of two partitions is generated using inter prediction, a predicted image for at least one partition may be generated using intra prediction.
[0515] Figure 53is a flowchart illustrating one example of a process in triangle mode.
[0516] In triangular mode, the inter-frame predictor 126 first splits the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode the partition information, which is information related to the partitioning, as a prediction parameter in the stream. In other words, the inter-frame predictor 126 can output the partition information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.
[0517] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (eg, MVs of multiple encoded blocks temporally or spatially surrounding the current block) (step Sx_2). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0518] The inter-frame predictor 126 then selects an MV candidate for the first partition and an MV candidate for the second partition from the multiple MV candidates obtained in step Sx_1 as the first MV and the second MV, respectively (step Sx_3). At this time, the inter-frame predictor 126 encodes MV selection information used to identify the selected MV candidate in the stream as a prediction parameter. In other words, the inter-frame predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130.
[0519] Next, the inter-frame predictor 126 generates a first predicted image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Similarly, the inter-frame predictor 126 generates a second predicted image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).
[0520] Finally, the inter predictor 126 generates a predicted image for the current block by performing weighted addition of the first predicted image and the second predicted image (step Sx_6).
[0521] Note that although Figure 52A In the example shown in FIG, the first partition and the second partition are triangular, but the first partition and the second partition may be trapezoidal or other shapes different from each other. In addition, although the current block is Figure 52A and Figure 52C The example shown in includes two partitions, but the current block may include three or more partitions.
[0522] Additionally, the first partition and the second partition may overlap each other. In other words, the first partition and the second partition may include the same pixel region. In this case, the predicted image in the first partition and the predicted image in the second partition may be used to generate a predicted image for the current block.
[0523] Additionally, although an example has been shown in which a predicted image for each of two partitions is generated using inter prediction, a predicted image for at least one partition may be generated using intra prediction.
[0524] Note that the MV candidate list used to select the first MV and the MV candidate list used to select the second MV may be different from each other, or the MV candidate list used to select the first MV may be used as the MV candidate list used to select the second MV.
[0525] Note that the partition information may include an index indicating the split direction in which at least the current block is split into multiple partitions. The MV selection information may include an index indicating a selected first MV and an index indicating a selected second MV. One index may indicate multiple pieces of information. For example, one index may be encoded that collectively indicates part or all of the partition information and part or all of the MV selection information.
[0526] (MV derivation > ATMVP mode)
[0527] Figure 54 is a conceptual diagram for illustrating one example of an advanced temporal motion vector prediction (ATMVP) mode in which an MV is derived in units of subblocks.
[0528] The ATMVP mode is a mode classified as a merge mode. For example, in the ATMVP mode, an MV candidate of each subblock is registered in an MV candidate list for use in a normal merge mode.
[0529] More specifically, in the ATMVP mode, first, as Figure 54 As shown in , a temporal MV reference block associated with the current block is identified in a coded reference picture specified by the MV (MV0) of a neighboring block located at a lower left position relative to the current block. Next, in each subblock in the current block, an MV for encoding the area corresponding to the subblock in the temporal MV reference block is identified. The MVs identified in this manner are included as MV candidates in an MV candidate list for the subblocks in the current block. When an MV candidate for each subblock is selected from the MV candidate list, the subblock undergoes motion compensation, where the MV candidate is used as the MV for the subblock. In this way, a predicted image is generated for each subblock.
[0530] Although Figure 54 In the example shown in , a block located at the lower left position relative to the current block is used as the surrounding MV reference block, but it should be noted that another block can be used. Additionally, the size of the sub-block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub-block can be switched for units such as slices, bricks, pictures, etc.
[0531] (Motion Estimation > DMVR)
[0532] Figure 55 is a flow chart illustrating the relationship between merge mode and decoder motion vector refinement (DMVR).
[0533] The inter-frame predictor 126 derives a motion vector for the current block according to the merge mode (step S1_1). Next, the inter-frame predictor 126 determines whether to perform estimation of the motion vector, that is, motion estimation (step S1_2). Here, when it is determined not to perform motion estimation ("No" in step S1_2), the inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). In other words, in this case, the motion vector for the current block is determined according to the merge mode.
[0534] When it is determined in step S1_1 that motion estimation is to be performed ("Yes" in step S1_2), the inter-frame predictor 126 derives a final motion vector for the current block by estimating the surrounding area of the reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to DMVR.
[0535] Figure 56 is a conceptual diagram for illustrating one example of a DMVR process for determining an MV.
[0536] First, for example, in merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference picture (L0), which is a coded picture in the L0 list, based on the MV candidate (L0). Similarly, reference pixels are identified from the second reference picture (L1), which is a coded picture in the L1 list, based on the MV candidate (L1). A template is generated by averaging these reference pixels.
[0537] Next, the template is used to estimate each of the surrounding areas of the MV candidates for the first reference picture (L0) and the second reference picture (L1), and the MV that produces the minimum cost is determined as the final MV. Note that the cost can be calculated using, for example, the difference between each pixel value in the template and the corresponding one of the pixel values in the estimation area, the value of the MV candidate, etc.
[0538] It is not always necessary to perform exactly the same procedure described here. Other procedures that enable derivation of the final MV by estimation in the surrounding area of the MV candidate may be used.
[0539] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining MV. Figure 56 An example of DMVR is shown in Figure 57 In the example shown in , the cost is calculated without generating a template.
[0540] First, the inter-frame predictor 126 estimates the surrounding area of the reference block in each of the reference pictures included in the L0 list and the L1 list based on the initial MV as the MV candidate obtained from each MV candidate list. Figure 57 As shown in , the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets a search position for the reference picture in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set, specifically the initial MV (i.e., InitMV_L0), the vector difference from the search position is MVd_L0. The inter-frame predictor 126 then determines an estimated position in the reference picture in the L1 list. The search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference as MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines a position symmetrical with respect to the position indicated by the initial MV as the search position in each reference picture in the L0 list and the L1 list. The inter predictor 126 calculates the sum of absolute differences (SAD) between the values of pixels at the search positions in the block as a cost for each search position, and finds a search position that results in the minimum cost.
[0541] Figure 58A is a conceptual diagram for illustrating an example of motion estimation in DMVR, and Figure 58B is a flowchart for illustrating one example of a process of motion estimation.
[0542] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and the eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each of the search positions other than the starting point is minimized. Here, if the cost at a search position other than the starting point is determined to be the smallest, the inter-frame predictor 126 changes the target to the search position at which the minimum cost is obtained and executes the process in step 2. If the cost at the starting point is the smallest, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.
[0543] In step 2, the inter-frame predictor 126 performs a search similar to the process in step 1, treating the search position after the target change as a new starting point based on the result of the process in step 1. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, if the cost at the search position other than the starting point is determined to be minimized, the inter-frame predictor 126 performs the process in step 4. If the cost at the starting point is minimized, the inter-frame predictor 126 performs the process in step 3.
[0544] In step 4, the inter predictor 126 regards the search position at the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0545] In step 3, the inter-frame predictor 126 determines the pixel position with sub-pixel accuracy that achieves the minimum cost based on the costs at four points located above, below, to the left, and to the right relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position. The pixel position with sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) above, below, to the left, and to the right, using the cost at a corresponding one of the four search positions as a weight. The inter-frame predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0546] (Motion Compensation > BIO / OBMC / LIC)
[0547] Motion compensation involves a mode for generating a predicted image and correcting the predicted image, such as bidirectional optical flow (BIO), overlapped block motion compensation (OBMC), and local illumination compensation (LIC), which will be described later.
[0548] Figure 59 is a flowchart illustrating an example of a process of generating a predicted image.
[0549] The inter predictor 126 generates a predicted image (step Sm_1 ), and corrects the predicted image according to, for example, any of the modes described above (step Sm_2 ).
[0550] Figure 60 is a flowchart illustrating another example of the process of generating a predicted image.
[0551] The inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, the inter-frame predictor 126 uses the motion vector to generate a predicted image (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, if it is determined that the correction process is to be performed ("Yes" in step Sn_3), the inter-frame predictor 126 generates a final predicted image by correcting the predicted image (step Sn_4). Note that in the LIC described later, luminance and chroma can be corrected in step Sn_4. If it is determined that the correction process is not to be performed ("No" in step Sn_3), the inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).
[0552] (Motion Compensation > OBMC)
[0553] Note that in addition to the motion information for the current block obtained through motion estimation, the motion information for the neighboring blocks can also be used to generate an inter-frame prediction image. More specifically, by performing a weighted addition of the prediction image based on the motion information obtained through motion estimation (in the reference picture) and the prediction image based on the motion information of the neighboring blocks (in the current picture), an inter-frame prediction image can be generated for each subblock in the current block. This inter-frame prediction (motion compensation) is also called overlapped block motion compensation (OBMC) or OBMC mode.
[0554] In OBMC mode, information indicating the sub-block size used for OBMC (e.g., referred to as OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether OBMC mode is applied (e.g., referred to as OBMC flag) may be signaled at the CU level. Note that signaling of such information does not necessarily need to be performed at the sequence level and the CU level, and may be performed at another level (e.g., picture level, slice level, tile level, CTU level, or sub-block level).
[0555] The OBMC mode will be described in more detail. Figure 61 and Figure 62 1 is a flowchart and a conceptual diagram for illustrating an outline of a predicted image correction process performed by OBMC.
[0556] First, if Figure 62 As shown in , the MV assigned to the current block is used to obtain a predicted image (Pred) through normal motion compensation. Figure 62 In FIG, the arrow “MV” points to the reference picture and indicates what the current block of the current picture refers to in order to obtain the predicted image.
[0557] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector for the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which indicates the reference picture from the current block. The first correction of the predicted image is performed by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0558] Similarly, a predicted image (Pred_U) is obtained by applying the MV (MV_U) derived for the coded block adjacent above the current block to the current block (reusing the MV for the current block). The MV (MV_U) is indicated by the arrow "MV_U," which indicates a reference picture from the current block. A second correction is performed on the predicted image by overlapping the predicted image Pred_U with the predicted image (e.g., Pred and Pred_L) on which the first correction has been performed. This provides an effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.
[0559] Although the above example is a two-path correction method using left and upper neighboring blocks, it should be noted that the correction method may be a three-path or more correction method also using right and / or lower neighboring blocks.
[0560] Note that the area where such overlapping is performed may be only a portion of the area near the block boundary, rather than the entire pixel area of the block.
[0561] Note that the above description describes a predicted image correction process based on OBMC, which is used to obtain a predicted image Pred from a single reference picture by superimposing the additional predicted images Pred_L and Pred_U. However, when correcting a predicted image based on multiple reference pictures, a similar process can be applied to each of the multiple reference pictures. In this case, after performing OBMC image correction based on the multiple reference pictures to obtain corrected predicted images from the corresponding reference pictures, the corrected predicted images are further superimposed to obtain the final predicted image.
[0562] Note that in OBMC, the current block unit can be a PU, or a sub-block unit obtained by further splitting the PU.
[0563] One example of a method for determining whether to apply OBMC is a method using obmc_flag, which serves as a signal indicating whether OBMC is applied. As a specific example, the encoder 100 may determine whether the current block belongs to an area with complex motion. When the block belongs to an area with complex motion, the encoder 100 sets obmc_flag to a value of "1" and applies OBMC during encoding. When the block does not belong to an area with complex motion, the encoder 100 sets obmc_flag to a value of "0" and encodes the block without applying OBMC. The decoder 200 switches between applying and not applying OBMC by decoding obmc_flag in the write stream.
[0564] (Motion Compensation > BIO)
[0565] Next, we describe the MV derivation method. First, we describe a method for derivation of MVs based on a model assuming uniform linear motion. This method is also known as the bidirectional optical flow (BIO) method. Alternatively, this bidirectional optical flow can be written as BDOF instead of BIO.
[0566] Figure 63 This is a conceptual diagram showing a model assuming uniform linear motion. Figure 63 In (v x , v y ) indicates a velocity vector, and τ0 and τ1 indicate the temporal distance between the current picture (Cur Pic) and two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) indicates the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) represents the MV corresponding to the reference picture Ref1.
[0567] Here, it is assumed that the velocity vector (v x , v y ) shows uniform linear motion, (MV x0 , MV y0 ) and (MV x1 , MV y1 ) are respectively expressed as (v xτ0 , v yτ0 ) and (-v xτ1 , -v yτ1 ), and gives the following optical flow equation (2).
[0568] [Mathematics 3]
[0569]
[0570] Here, I(k) indicates the motion-compensated luminance value k (k = 0, 1) of the reference picture after motion compensation. This optical flow equation indicates that the sum of the following items is equal to zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected on a pixel-by-pixel basis.
[0571] Note that a method other than deriving a motion vector based on a model assuming uniform linear motion may be used to derive a motion vector at the decoder side 200. For example, a motion vector may be derived in units of subblocks based on motion vectors of a plurality of adjacent blocks.
[0572] Figure 64 is a flowchart illustrating one example of a process of inter-frame prediction according to BIO. Figure 65 is a functional block diagram showing one example of a functional configuration of the inter-frame predictor 126 that can perform inter-frame prediction according to BIO.
[0573] like Figure 65 As shown in FIG, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolation image deriver 126b, a gradient image deriver 126c, an optical flow deriver 126d, a correction value deriver 126e, and a predicted image corrector 126f. Note that the memory 126a may be the frame memory 122.
[0574] The inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter-frame predictor 126 derives a predicted image for the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ), and the motion vector M1 is the motion vector (MV x1 , MV y1 ).
[0575] Next, the interpolation image deriver 126b derivates an interpolation image I for the current block using the motion vector M0 and the reference picture L0 through the reference memory 126a. 0 Next, the interpolation image deriver 126b derivates an interpolation image I for the current block using the motion vector M1 and the reference picture L1 through the reference memory 126a. 1 (Step Sy_2). Here, the interpolated image I 0is an image included in the reference picture Ref0 and derived for the current block, and the interpolated image I 1 Is an image included in the reference picture Ref1 and derived for the current block. Interpolated image I 0 and interpolated image I 1 Each of the interpolated images I may be the same size as the current block. 0 and interpolated image I 1 Each of the images can be larger than the current block. 0 and interpolated image I 1 A predicted image obtained by using a motion vector (M0, M1) and a reference picture (L0, L1) and applying a motion compensation filter may be included.
[0576] Additionally, the gradient image deriver 126c generates the image according to the interpolation image I 0 and interpolated image I 1 Derive the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1 )(Step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The gradient image deriver 126c may derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may indicate the amount of spatial change in pixel values in the horizontal direction, in the vertical direction, or in both directions.
[0577] Next, the optical flow deriver 126d uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) An optical flow (vx, vy) as a velocity vector is derived for each sub-block of the current block (step Sy_4). The optical flow indicates a coefficient for correcting the amount of spatial pixel movement and may be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block may be a 4×4 pixel sub-CU. Note that optical flow derivation may be performed for each pixel unit, etc., rather than for each sub-block.
[0578] Next, the inter-frame predictor 126 uses the optical flow (vx, vy) to correct the predicted image for the current block. For example, the correction value deriver 126e uses the optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_5). The predicted image corrector 126f can then use the correction values to correct the predicted image for the current block (step Sy_6). Note that the correction values can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.
[0579] Note that the BIO process flow is not limited to Figure 64 For example, you can just execute Figure 64 A part of the process disclosed in the specification may be used, or a different process may be added or used instead, or the processes may be performed in a different processing order, etc.
[0580] (Motion Compensation > LIC)
[0581] Next, one example of a mode for generating a predicted image (Prediction) using a Local Illumination Compensation (LIC) process is described.
[0582] Figure 66A is a conceptual diagram for illustrating one example of a process of a predicted image generation method using an illumination correction process performed by the LIC. Figure 66B is a flowchart showing one example of the process of a predicted image generation method using LIC.
[0583] First, the inter predictor 126 derives MV from the encoded reference picture and obtains a reference image corresponding to the current block (step Sz_1 ).
[0584] Next, the inter-frame predictor 126 extracts information indicating how the luma value of the current block changes between the current block and the reference picture (step Sz_2). This extraction is performed based on the luma pixel values of the encoded left-neighboring reference region (surrounding reference region) and the encoded upper-neighboring reference region (surrounding reference region) in the current picture, as well as the luma pixel values at corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 126 uses this information indicating how the luma value changes to calculate illumination correction parameters (step Sz_3).
[0585] The inter-frame predictor 126 generates a predicted image for the current block by performing an illumination correction process in which illumination correction parameters are applied to the reference image in the reference picture specified by the MV (step Sz_4). In other words, the predicted image is corrected based on the illumination correction parameters, which serve as the reference image in the reference picture specified by the MV. In this correction, the illumination, chroma, or both can be corrected. In other words, the chroma correction parameters can be calculated using information indicating how the chroma has changed, and the chroma correction process can be performed.
[0586] Notice, Figure 66A The shape of the surrounding reference area shown in is an example; another shape may be used.
[0587] In addition, although the process of generating a predicted image based on a single reference picture has been described here, the case of generating a predicted image based on multiple reference pictures can be described in the same manner. The predicted image can be generated after performing the illumination correction process on the reference image obtained from the reference picture in the same manner as described above.
[0588] One example of a method for determining whether to apply LIC is a method using lic_flag, which is a signal indicating whether LIC is applied. As a specific example, the encoder 100 determines whether the current block belongs to an area with illumination changes. When the block belongs to an area with illumination changes, the encoder 100 sets the lic_flag value to "1" and applies LIC during encoding. When the block does not belong to an area with illumination changes, the encoder 100 sets the lic_flag value to "0" and performs encoding without applying LIC. The decoder 200 can decode the lic_flag written in the stream and decode the current block by switching between applying and not applying LIC according to the flag value.
[0589] One example of a different method for determining whether to apply the LIC process is based on whether the LIC process has already been applied to surrounding blocks. As a specific example, when the current block is already processed in merge mode, the inter-frame predictor 126 determines whether the coded surrounding blocks selected in MV derivation in merge mode have already been coded using LIC. The inter-frame predictor 126 performs encoding by switching between applying and not applying LIC based on the result. Note that in this example, the same process is also applied to the decoder 200.
[0590] The illumination correction (LIC) process has been referenced Figure 66A and Figure 66B are described and are further described below.
[0591] First, the inter predictor 126 derives an MV for obtaining a reference image corresponding to a current block to be encoded from a reference picture that is an encoded picture.
[0592] Next, the inter-frame predictor 126 uses the luminance pixel values of the encoded surrounding reference areas adjacent to the left and above the current block and the luminance values in the corresponding positions of the reference picture specified by the MV to extract information indicating how the luminance value of the reference picture changes to the luminance value of the current picture and calculate the illumination correction parameters. For example, assume that the luminance pixel value of a given pixel in the surrounding reference area in the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference area in the reference picture is p1. The inter-frame predictor 126 calculates coefficients A and B for optimizing A×p1+B=p0 as illumination correction parameters for multiple pixels in the surrounding reference area.
[0593] Next, the inter-frame predictor 126 performs illumination correction processing using the illumination correction parameters for the reference image in the reference picture specified by MV to generate a predicted image for the current block. For example, assume that the luminance pixel value in the reference image is p2, and the luminance pixel value of the predicted image after illumination correction is p3. The inter-frame predictor 126 calculates A×p2+B=p3 for each pixel in the reference image, and generates a predicted image after undergoing the illumination correction process.
[0594] For example, a region having a certain number of pixels extracted from each of the upper adjacent pixels and the left adjacent pixels may be used as a surrounding reference region. Additionally, the surrounding reference region is not limited to a region adjacent to the current block, and may also be a region not adjacent to the current block. Figure 66A In the example shown in , the surrounding reference region in the reference picture may be a region specified by another MV in the current picture from among the surrounding reference regions in the current picture. For example, the other MV may be an MV in the surrounding reference region in the current picture.
[0595] Although operations performed by encoder 100 have been described herein, it should be noted that decoder 200 performs similar operations.
[0596] Note that LIC can be applied not only to luminance but also to chrominance. In this case, correction parameters can be derived separately for each of Y, Cb, and Cr, or common correction parameters can be used for any one of Y, Cb, and Cr.
[0597] Alternatively, the LIC process may be applied in units of subblocks. For example, correction parameters may be derived using surrounding reference regions in the current subblock and surrounding reference regions in reference subblocks in a reference picture specified by the MV of the current subblock.
[0598] (Predictive Controller)
[0599] The prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from the intra-frame predictor 124) and the inter-frame prediction signal (the image or signal output from the inter-frame predictor 126), and outputs the selected prediction image to the subtractor 104 and the adder 116 as a prediction signal.
[0600] (Prediction Parameter Generator)
[0601] The prediction parameter generator 130 may output information related to intra prediction, inter prediction, selection of a prediction image in the prediction controller 128, and the like as prediction parameters to the entropy encoder 110. The entropy encoder 110 may generate a stream based on the prediction parameters input from the prediction parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameters may be used in the decoder 200. The decoder 200 may receive and decode the stream and perform the same prediction process as that performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameters may include, for example, (i) a selection prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) an optional index, flag, or value based on the prediction process performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128 or indicating the prediction process.
[0602] (Decoder)
[0603] Next, a decoder 200 capable of decoding the stream output from the encoder 100 described above is described. Figure 67 2 is a block diagram showing a functional configuration of the decoder 200 according to the present embodiment. The decoder 200 is a device that decodes a stream, which is an encoded image, in units of blocks.
[0604] like Figure 67 As shown in FIG, the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformer 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a prediction parameter generator 222, and a split determiner 224. Note that the intra predictor 216 and the inter predictor 218 are configured as part of a prediction performer.
[0605] (Decoder installation example)
[0606] Figure 68 2 is a functional block diagram showing an example of an installation of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, Figure 67The multiple constituent elements of the decoder 200 shown in FIG. 2 are installed in Figure 68 The processor b1 and memory b2 are shown.
[0607] Processor b1 is a circuit that performs information processing and is coupled to memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 may be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits. In addition, for example, processor b1 may be responsible for Figure 67 The roles of two or more constituent elements other than the constituent elements for storing information among the plurality of constituent elements of the decoder 200 shown in FIG.
[0608] Memory b2 is a dedicated or general-purpose memory for storing information used by processor b1 to decode the stream. Memory b2 may be an electronic circuit and may be connected to processor b1. Alternatively, memory b2 may be included in processor b1. Alternatively, memory b2 may be a collection of multiple electronic circuits. Alternatively, memory b2 may be a magnetic disk, optical disk, or the like, or may be represented as a storage device, recording medium, or the like. Alternatively, memory b2 may be a non-volatile memory or a volatile memory.
[0609] For example, the memory b2 may store an image or a stream. Alternatively, the memory b2 may store a program for causing the processor b1 to decode the stream.
[0610] Alternatively, for example, memory b2 may assume Figure 67 The memory b2 may be used to store information in the decoder 200. More specifically, the memory b2 may be used to store information in the decoder 200. Figure 67 . More specifically, the memory b2 can store reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.).
[0611] Note that in decoder 200, it may not be implemented Figure 67 All elements among the multiple constituent elements indicated in the figure, etc., and all processes described in this document may not be performed. Figure 67 A portion of the constituent elements indicated in the , etc. may be included in another device, or a portion of the process described herein may be performed by another device.
[0612] Hereinafter, the overall flow of the process performed by the decoder 200 will be described, and then each of the constituent elements included in the decoder 200 will be described. Note that some of the constituent elements included in the decoder 200 perform the same processes as those performed by some of the encoder 100, and therefore the same processes will not be described again in detail. For example, the inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, and loop filter 212 included in the decoder 200 respectively perform similar processes to those performed by the inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and loop filter 120 included in the encoder 100.
[0613] (Overall flow of the decoding process)
[0614] Figure 69 is a flowchart illustrating one example of the overall decoding process performed by the decoder 200 .
[0615] First, the split determiner 224 in the decoder 200 determines a split mode for each of a plurality of fixed-size blocks (e.g., 128×128 pixels) included in a picture based on the parameters input from the entropy decoder 202 (step Sp_1). This split mode is the split mode selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks in the split mode.
[0616] The entropy decoder 202 decodes (specifically, performs entropy decoding) the encoded quantized coefficients and prediction parameters of the current block (step Sp_2).
[0617] Next, the inverse quantizer 204 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 206 performs inverse transform on the result to restore the prediction residual (ie, difference block) (step Sp_3).
[0618] Next, the prediction performer including all or part of the intra predictor 216, the inter predictor 218, and the prediction controller 220 generates a prediction signal for the current block (step Sp_4).
[0619] Next, the adder 208 adds the predicted image and the prediction residual to generate a reconstructed image of the current block (also referred to as a decoded image block) (step Sp_5).
[0620] When the reconstructed image is generated, the loop filter 212 performs filtering on the reconstructed image (step Sp_6).
[0621] The decoder 200 then determines whether decoding of the entire picture has ended (step Sp_7). When it is determined that decoding has not ended ("No" in step Sp_7), the decoder 200 repeats the process starting from step Sp_1.
[0622] Note that the processes of these steps Sp_1 to Sp_7 may be sequentially performed by the decoder 200, or two or more processes in the process may be performed in parallel. The processing order of two or more processes in the process may be modified.
[0623] (Split Determiner)
[0624] Figure 70 2 is a conceptual diagram for illustrating the relationship between the split determiner 224 and other constituent elements in the embodiment. As an example, the split determiner 224 may perform the following process.
[0625] For example, the split determiner 224 collects block information from the block memory 210 or the frame memory 214 and further obtains parameters from the entropy decoder 202. The split determiner 224 may then determine a splitting mode for the fixed-size block based on the block information and the parameters. The split determiner 224 may then output information indicating the determined splitting mode to the inverse transformer 206, the intra-frame predictor 216, and the inter-frame predictor 218. The inverse transformer 206 may perform an inverse transform on the transform coefficients based on the splitting mode indicated by the information from the split determiner 224. The intra-frame predictor 216 and the inter-frame predictor 218 may generate a predicted image based on the splitting mode indicated by the information from the split determiner 224.
[0626] (Entropy Decoder)
[0627] Figure 71 is a block diagram showing one example of the functional configuration of the entropy decoder 202.
[0628] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the split mode by entropy decoding the stream. For example, CABAC is used for entropy decoding. More specifically, the entropy decoding 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a de-binarizer 202c. The binary arithmetic decoder 202a uses the context value derived by the context controller 202b to arithmetically decode the stream into a binary signal. The context controller 202b derives the context value based on the characteristics of the syntactic element or the surrounding state (i.e., the probability of occurrence of the binary signal) in the same manner as the context controller 110b of the encoder 100. The de-binarizer 202c performs de-binarization to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. The binarization can be performed according to the binarization method described above.
[0629] Thus, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 may output the prediction parameters (see Figure 1 ) is output to the intra predictor 216, the inter predictor 218, and the prediction controller 220. The intra predictor 216, the inter predictor 218, and the prediction controller 220 can perform the same prediction processes as those performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.
[0630] Figure 72 is a conceptual diagram for illustrating the flow of an example CABAC process in the entropy decoder 202.
[0631] First, initialization is performed in CABAC within the entropy decoder 202. During initialization, the binary arithmetic decoder 202a is initialized and an initial context value is set. The binary arithmetic decoder 202a and the debinarizer 202c then perform arithmetic decoding and debinarization of the coded data, such as a CTU. At this point, the context controller 202b updates the context value each time row arithmetic decoding is performed. The context controller 202b then saves the context value for post-processing. For example, the saved context value is used to initialize the context value for the next CTU.
[0632] (Inverse Quantizer)
[0633] The inverse quantizer 204 inversely quantizes the quantized coefficients of the current block input from the entropy decoder 202. More specifically, the inverse quantizer 204 inversely quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inversely quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.
[0634] Figure 73 is a block diagram showing one example of the functional configuration of the inverse quantizer 204 .
[0635] The inverse quantizer 204 includes, for example, a quantization parameter generator 204 a , a predicted quantization parameter generator 204 b , a quantization parameter storage 204 d , and an inverse quantization performer 204 e .
[0636] Figure 74 is a flowchart illustrating one example of a process of inverse quantization performed by the inverse quantizer 204 .
[0637] As an example, the inverse quantizer 204 may be based on Figure 74 The inverse quantization process is performed on each CU in the process shown in FIG. More specifically, the quantization parameter generator 204 a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined that inverse quantization is to be performed (“Yes” in step Sv_11), the quantization parameter generator 204 a obtains a difference quantization parameter for the current block from the entropy decoder 202 (step Sv_12).
[0638] Next, the predicted quantization parameter generator 204b then obtains the quantization parameter for the processing unit different from the current block from the quantization parameter storage device 204d (step Sv_13). The predicted quantization parameter generator 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameter (step Sv_14).
[0639] The quantization parameter generator 204a then generates a quantization parameter for the current block based on the difference quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the difference quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b may be added together to generate the quantization parameter for the current block. Additionally, the quantization parameter generator 204a stores the quantization parameter for the current block in the quantization parameter storage device 204d (step Sv_16).
[0640] Next, the inverse quantization performer 204e inversely quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameters generated in step Sv_15 (step Sv_17).
[0641] Note that the difference quantization parameter can be decoded at the bit sequence level, picture level, slice level, brick level, or CTU level. Alternatively, the initial value of the quantization parameter can be decoded at the sequence level, picture level, slice level, brick level, or CTU level. In this case, the initial value of the quantization parameter and the difference quantization parameter can be used to generate the quantization parameter.
[0642] Note that the inverse quantizer 204 may include a plurality of inverse quantizers and may inverse quantize the quantized coefficient using an inverse quantization method selected from a plurality of inverse quantization methods.
[0643] (Inverse Converter)
[0644] The inverse transformer 206 restores the prediction residual by inversely transforming the transform coefficients that are input from the inverse quantizer 204 .
[0645] For example, when information parsed from the stream indicates that EMT or AMT is to be applied (eg, when the AMT flag is true), the inverse transformer 206 inversely transforms transform coefficients of the current block based on information indicating the parsed transform type.
[0646] Furthermore, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.
[0647] Figure 75 is a flowchart illustrating one example of a process performed by the inverse converter 206 .
[0648] For example, the inverse transformer 206 determines whether there is information in the stream indicating that an orthogonal transform is not to be performed (step St_11). Here, when it is determined that there is no such information ("No" in step St_11) (for example: there is no indication as to whether an orthogonal transform is to be performed; there is an indication that an orthogonal transform is to be performed), the inverse transformer 206 obtains information indicating the transform type decoded by the entropy decoder 202 (step St_12). Next, based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step St_13). The inverse transformer 206 then performs an inverse orthogonal transform using the determined transform type (step St_14). Figure 75 As shown in , when it is determined that there is information indicating that the orthogonal transform is not to be performed ("Yes" in step St_11) (for example, an explicit instruction not to perform the orthogonal transform; there is no instruction to perform the orthogonal transform), the orthogonal transform is not performed.
[0649] Figure 76 is a flowchart illustrating one example of a process performed by the inverse converter 206 .
[0650] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a determined value (step Su_11). The determined value may be predetermined. Here, when it is determined that the transform size is less than or equal to the determined value ("Yes" in step Su_11), the inverse transformer 206 obtains information indicating which transform type is used by the encoder 100 from the entropy decoder 202 among the at least one transform type included in the first transform type group (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transformer 206.
[0651] Based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). If it is determined that the transform size is not less than or equal to the determined value ("No" in step Su_11), the inverse transformer 206 inversely transforms the transform coefficients of the current block using the second transform type group (step Su_15).
[0652] Note that, as an example, the inverse orthogonal transform performed by inverse transformer 206 may be based on Figure 75 or Figure 76 The process shown in FIG is performed for each TU. Alternatively, an inverse orthogonal transform may be performed by using a defined transform type without decoding information indicating the transform type used for the orthogonal transform. The defined transform type may be a predefined transform type or a default transform type. Alternatively, the transform type may specifically be DST7, DCT8, or the like. In the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.
[0653] (Adder)
[0654] The adder 208 reconstructs the current block by adding the prediction residual as input from the inverse transformer 206 and the predicted image as input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.
[0655] (Block Storage)
[0656] The block memory 210 is a storage device for storing blocks included in the current picture and that can be referenced in intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.
[0657] (Loop Filter)
[0658] The loop filter 212 applies the loop filter to the reconstructed image generated by the adder 208 and outputs the filtered reconstructed image to the frame memory 214 and provides an output of the decoder 200, for example, to a display device or the like.
[0659] When the information indicating on or off of ALF parsed from the stream indicates that ALF is on, one filter may be selected from a plurality of filters based on, for example, the direction and activity of a local gradient, and the selected filter is applied to the reconstructed image.
[0660] Figure 77 2 is a block diagram showing one example of the functional configuration of the loop filter 212. Note that the loop filter 212 has a configuration similar to that of the loop filter 120 of the encoder 100.
[0661] For example, Figure 77 As shown in FIG, the loop filter 212 includes a deblocking filter executor 212a, an SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filtering process on the reconstructed image. The SAO executor 212b performs an SAO process on the reconstructed image after the deblocking filtering process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include Figure 77 All the constituent elements disclosed in the embodiment and may include only a part of the constituent elements. Additionally, the loop filter 212 may be configured to Figure 77 The above process may be performed in a different processing order than the processing order disclosed in Figure 77 All the processes shown in , etc.
[0662] (Frame Memory)
[0663] The frame memory 214 is a storage device for storing reference pictures used for inter-frame prediction, for example, and may also be referred to as a frame buffer. More specifically, the frame memory 214 stores the reconstructed image filtered by the loop filter 212.
[0664] (Predictor (Intra Predictor, Inter Predictor, Prediction Controller))
[0665] Figure 78 2 is a flowchart showing an example of a process performed by the predictor of the decoder 200. Note that the prediction performer may include all or part of the following constituent elements: the intra predictor 216; the inter predictor 218; and the prediction controller 220. The prediction performer includes, for example, the intra predictor 216 and the inter predictor 218.
[0666] The predictor generates a predicted image for the current block (step Sq_1). This predicted image may also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra-frame predicted image or an inter-frame predicted image. More specifically, the predictor generates a predicted image for the current block by generating a predicted image, restoring a prediction residual, and adding the predicted image using a reconstructed image already obtained for another block. The predictor of the decoder 200 generates the same predicted image as the predicted image generated by the predictor of the encoder 100. In other words, the predicted image is generated according to a method common to the predictors or a method corresponding to each other.
[0667] The reconstructed image may be, for example, an image in a reference picture, or an image of a decoded block (i.e., other blocks described above) in a current picture that is a picture including the current block. The decoded block in the current picture may be, for example, a neighboring block of the current block.
[0668] Figure 79 is a flowchart illustrating another example of a process performed by the predictor of the decoder 200 .
[0669] The predictor determines a method or mode for generating a predicted image (step Sr_1). For example, the method or mode may be determined based on prediction parameters or the like.
[0670] When the first method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating a predicted image, the predictor generates a predicted image according to the third method (step Sr_2c).
[0671] The first method, the second method, and the third method may be different methods for generating a predicted image. Each of the first method to the third method may be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above may be used in these prediction methods.
[0672] Figure 80 is a flowchart illustrating another example of a process performed by the predictor of the decoder 200 .
[0673] As an example, the predictor can be based on Figure 80 The prediction process is performed according to the process shown in . Note that Figure 80 The intra-block copy shown in is a mode belonging to inter-frame prediction, and the block included in the current picture is called a reference image or a reference block. In other words, a picture different from the current picture is not referenced in the intra-block copy. In addition, Figure 80The PCM mode shown in is a mode belonging to intra prediction, and in which transform and quantization are not performed.
[0674] (Intra-frame predictor)
[0675] The intra-frame predictor 216 performs intra-frame prediction based on the intra-frame prediction mode parsed from the stream by referring to the blocks in the current picture stored in the block memory 210 to generate a predicted image of the current block (i.e., the intra-frame prediction block). More specifically, the intra-frame predictor 216 performs intra-frame prediction by referring to the pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-frame prediction image, and then outputs the intra-frame prediction image to the prediction controller 220.
[0676] Note that when an intra prediction mode in which a luma block is referenced in intra prediction of a chroma block is selected, the intra predictor 216 may predict the chroma components of the current block based on the luma components of the current block.
[0677] Furthermore, when information parsed from the stream indicates that PDPC is to be applied, the intra predictor 216 corrects the intra-predicted pixel value based on horizontal / vertical reference pixel gradients.
[0678] Figure 81 is a diagram showing one example of a process performed by the intra predictor 216 of the decoder 200 .
[0679] The intra predictor 216 first determines whether to use MPM. Figure 81 As shown in FIG, the intra predictor 216 determines whether an MPM flag indicating 1 is present in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 is present ("Yes" in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 from the entropy decoder 202. Note that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode included in the plurality of intra prediction modes included in the MPM and indicated by the information obtained in step Sw_12 (step Sw_14).
[0680] When it is determined that there is no MPM flag indicating 1 ("No" in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 from the entropy decoder 202 from at least one intra prediction mode not included in the MPM. Note that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. The intra predictor 216 then determines an intra prediction mode that is not included in the plurality of intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15 (step Sw_17).
[0681] The intra predictor 216 generates a predicted image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18 ).
[0682] (Inter-frame predictor)
[0683] The inter-frame predictor 218 predicts the current block by referencing the reference picture stored in the frame memory 214. Prediction is performed in units of the current block or the current subblock in the current block. Note that a subblock is included in a block and is a unit smaller than a block. The size of the subblock can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the subblock can be switched for units such as slices, tiles, pictures, etc.
[0684] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from the stream (e.g., prediction parameters output from the entropy decoder 202), and outputs the inter-frame prediction image to the prediction controller 220.
[0685] When the information parsed from the stream indicates that the OBMC mode is to be applied, the inter predictor 218 generates an inter predicted image using motion information of neighboring blocks in addition to motion information of the current block obtained through motion estimation.
[0686] Furthermore, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives motion information by performing motion estimation according to a pattern matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter-frame predictor 218 then performs motion compensation (prediction) using the derived motion information.
[0687] Furthermore, when the BIO mode is to be applied, the inter-frame predictor 218 derives an MV based on a model assuming uniform linear motion. Alternatively, when information parsed from the stream indicates that the affine mode is to be applied, the inter-frame predictor 218 derives an MV for each subblock based on the MVs of multiple neighboring blocks.
[0688] (MV derivation process)
[0689] Figure 82 is a flowchart illustrating one example of an MV derivation process in the decoder 200 .
[0690] For example, the inter-frame predictor 218 determines whether to decode motion information (e.g., MV). For example, the inter-frame predictor 218 may make this determination based on a prediction mode included in the stream, or may make this determination based on other information included in the stream. Here, when it is determined that motion information is to be decoded, the inter-frame predictor 218 derives an MV for the current block in a mode in which motion information is decoded. When it is determined that motion information is not to be decoded, the inter-frame predictor 218 derives an MV in a mode in which motion information is not decoded.
[0691] Here, MV derivation modes include normal inter mode, normal merge mode, FRUC mode, affine mode, and the like, which will be described later. Among the modes, modes in which motion information is decoded include normal inter mode, normal merge mode, affine mode (specifically, affine inter mode and affine merge mode), and the like. Note that motion information may include not only MVs but also MV predictor selection information, which will be described later. Modes in which motion information is not decoded include FRUC mode, and the like. The inter predictor 218 selects a mode for deriving the MV for the current block from a plurality of modes, and uses the selected mode to derive the MV for the current block.
[0692] Figure 83 is a flowchart illustrating one example of a process of MV derivation in the decoder 200 .
[0693] For example, the inter-frame predictor 218 may determine whether to decode the MV difference, i.e., it may be determined based on the prediction mode included in the stream, or it may be determined based on other information included in the stream. Here, when it is determined that the MV difference is to be decoded, the inter-frame predictor 218 may derive the MV for the current block in a mode in which the MV difference is decoded. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.
[0694] When it is determined that no MV difference is to be decoded, the inter-frame predictor 218 derives the MV in a mode in which the MV difference is not decoded. In this case, the encoded MV difference is not included in the stream.
[0695] Here, as described above, the MV derivation mode includes the normal inter mode, normal merge mode, FRUC mode, affine mode, etc., which will be described later. Among the modes, the modes in which the MV difference is encoded include the normal inter mode and the affine mode (specifically, the affine inter mode), etc. The modes in which the MV difference is not encoded include the FRUC mode, the normal merge mode, the affine mode (specifically, the affine merge mode), etc. The inter predictor 218 selects a mode for deriving the MV for the current block from the plurality of modes, and derives the MV for the current block using the selected mode.
[0696] (MV derivation > normal inter-frame mode)
[0697] For example, when the information parsed from the stream indicates that the normal inter mode is to be applied, the inter predictor 218 derives an MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.
[0698] Figure 84 is a flowchart illustrating an example of a process of inter prediction by normal inter mode in the decoder 200 .
[0699] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (e.g., MVs of multiple decoded blocks temporally or spatially surrounding the current block) (step Sg_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0700] Next, the inter-frame predictor 218 extracts N (an integer of 2 or greater) MV candidates from the plurality of MV candidates obtained in step Sg_11 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the ranking in the determined priority order (step Sg_12). Note that the ranking in the priority order may be predetermined for the respective N MV predictor candidates, and the ranking may be predetermined.
[0701] Next, the inter predictor 218 decodes the MV predictor selection information from the input stream and selects one MV predictor candidate from the N MV predictor candidates as the MV predictor for the current block using the decoded MV predictor selection information (step Sg_13).
[0702] Next, the inter-frame predictor 218 decodes the MV difference from the input stream and derives the MV for the current block by adding the difference value, which is the decoded MV difference, to the selected MV predictor (step Sg_14).
[0703] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The process in steps Sg_11 to Sg_15 is performed for each block. For example, when the process in steps Sg_11 to Sg_15 is performed for each block in all blocks in the slice, inter-frame prediction for the slice using normal inter-frame mode is completed. For example, when the process in steps Sg_11 to Sg_15 is performed for each block in all blocks in the picture, inter-frame prediction for the picture using normal inter-frame mode is completed. Note that not all blocks included in the slice may undergo the process in steps Sg_11 to Sg_15, and inter-frame prediction for the slice using normal inter-frame mode may end when a portion of the blocks undergo the process. This also applies to the picture in steps Sg_11 to Sg_15. When the process is performed for a portion of the blocks in the picture, inter-frame prediction for the picture using normal inter-frame mode may end.
[0704] (MV derivation > normal merge mode)
[0705] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter predictor 218 derives an MV and performs motion compensation (prediction) using the MV.
[0706] Figure 85 is a flowchart illustrating an example of a process of performing inter-frame prediction by normal merge mode in the decoder 200 .
[0707] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (eg, MVs of multiple decoded blocks temporally or spatially surrounding the current block) (step Sh_11). In other words, the inter-frame predictor 218 generates an MV candidate list.
[0708] Next, the inter-frame predictor 218 selects one MV candidate from the multiple MV candidates obtained in step Sh_11, thereby deriving the MV for the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains MV selection information included in the stream as a prediction parameter, and selects the MV candidate identified by the MV selection information as the MV for the current block.
[0709] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). For example, the process in steps Sh_11 to Sh_13 is performed for each block. For example, when the process in steps Sh_11 to Sh_13 is performed for each of all blocks in the slice, inter-frame prediction using normal merge mode for the slice is completed. Alternatively, when the process in steps Sh_11 to Sh_13 is performed for each of all blocks in the picture, inter-frame prediction using normal merge mode for the picture is completed. Note that not all blocks included in the slice undergo the process in steps Sh_11 to Sh_13, and inter-frame prediction using normal merge mode for the slice may end when a portion of the blocks undergo the process. This also applies to the picture in steps Sh_11 to Sh_13. When the process is performed on a portion of the blocks in the picture, inter-frame prediction using normal merge mode for the picture may end.
[0710] (MV derivation > radial merger mode)
[0711] For example, when information parsed from the stream indicates that the FRUC mode is to be applied, the inter-frame predictor 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, motion information is derived on the decoder 200 side without being signaled from the encoder 100 side. For example, the decoder 200 can derive motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values in the current block.
[0712] Figure 86 is a flowchart illustrating an example of a process of inter-frame prediction by the FRUC mode in the decoder 200 .
[0713] First, the inter-frame predictor 218 generates a list of MVs indicating decoded blocks that are spatially or temporally adjacent to the current block by referring to the MVs that serve as MV candidates (the list is an MV candidate list and can also be used as an MV candidate list for normal merge mode, for example) (step Si_11). Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates an evaluation value for each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV for the current block (step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block can be derived using pattern matching in a surrounding area included in the reference picture and corresponding to the position of the selected best MV candidate. In other words, estimation using pattern matching and evaluation values in a reference picture may be performed in the surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate may be updated to the MV that produces the better evaluation value, and the updated MV may be determined as the final MV for the current block. In an embodiment, updating to the MV that produces the better evaluation value may not be performed.
[0714] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). For example, the process in steps Si_11 to Si_15 is performed for each block. For example, when the process in steps Si_11 to Si_15 is performed for each of all blocks in the slice, inter-frame prediction for the slice using the FRUC mode is completed. For example, when the process in steps Si_11 to Si_15 is performed for each of all blocks in the picture, inter-frame prediction for the picture using the FRUC mode is completed. Each sub-block can be processed similarly to the case of each block.
[0715] (MV derivation > FRUC mode)
[0716] For example, when the information parsed from the stream indicates that the affine merge mode is to be applied, the inter predictor 218 derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0717] Figure 87 is a flowchart illustrating an example of a process of inter-frame prediction by affine merge mode in the decoder 200 .
[0718] In the affine merge mode, first, the inter-frame predictor 218 derives the MV at the corresponding control point for the current block (step Sk_11). Figure 46A As shown in , the control points are the upper left corner point of the current block and the upper right corner point of the current block, or as Figure 46B As shown in , the control points are the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block.
[0719] For example, when using Figures 47A to 47C When the MV derivation method shown in Figure 47A As shown in FIG, the inter-frame predictor 218 examines the decoded blocks A (left), B (above), C (above right), D (below left), and E (above left) in order and identifies the first valid block decoded according to the affine mode. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine mode to derive the MV at the control point. For example, when block A is identified and block A has two control points, as shown in FIG. Figure 47B , the inter-frame predictor 218 calculates a motion vector v0 at the upper left control point of the current block and a motion vector v1 at the upper right control point of the current block based on the motion vectors v3 and v4 at the upper left and upper right corners of the decoded block including block A. In this way, the MV at each control point is derived.
[0720] Note that Figure 49A As shown in , when block A is identified and block A has two control points, the MV at three control points can be calculated, and as Figure 49B As shown in , when block A is identified and when block A has three control points, MVs at two control points can be calculated.
[0721] Additionally, when MV selection information is included in the stream as a prediction parameter, the inter predictor 218 may use the MV selection information to derive the MV at each control point for the current block.
[0722] Next, the inter-frame predictor 218 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or uses three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV for each of the multiple sub-blocks as an affine MV (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. In other words, motion compensation is performed on the current block to generate a predicted image for the current block.
[0723] Note that the MV candidate list described above may be generated in step Sk_11. The MV candidate list may be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods may be, for example, Figures 47A to 47C The MV derivation method shown in Figure 48A and Figure 48B The MV derivation method shown in Figure 49A and Figure 49B Any combination of the MV derivation method shown in and other MV derivation methods.
[0724] Note that the MV candidate list may include MV candidates in a mode in which prediction is performed in units of subblocks other than the affine mode.
[0725] Note that, for example, an MV candidate list including MV candidates in the affine merge mode using two control points and in the affine merge mode using three control points may be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points may be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode using two control points and the affine merge mode using three control points may be generated.
[0726] (MV derivation > affine inter-frame mode)
[0727] For example, when the information parsed from the stream indicates that the affine inter mode is to be applied, the inter predictor 218 derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.
[0728] Figure 88 is a flowchart illustrating an example of a process of inter prediction by affine inter mode in the decoder 200 .
[0729] In the affine inter mode, first, the inter predictor 218 derives the MV predictors (v0, v1) or (v0, v1, v2) of the corresponding two or three control points of the current block (step Sj_11). The control points are the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block, as shown in FIG. Figure 46A or Figure 46B As shown in .
[0730] The inter-frame predictor 218 obtains the MV predictor selection information included in the stream as a prediction parameter, and uses the MV identified by the MV predictor selection information to derive the MV predictor at each control point of the current block. Figure 48A and Figure 48BWhen the MV derivation method shown in FIG. 1 is used, the inter-frame predictor 218 selects Figure 48A or Figure 48B The MV predictor in the decoded block near the corresponding control point of the current block shown in is selected using the MV of the block identified by the information to derive the motion vector predictor (v0, v1) or (v0, v1, v2) at the control point of the current block.
[0731] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block to the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point for the current block is derived.
[0732] Next, the inter-frame predictor 218 performs motion compensation on each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or uses three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV for each of the multiple sub-blocks as an affine MV (step Sk_13). The inter-frame predictor 218 then uses these affine MVs and the encoded reference picture to perform motion compensation on the sub-blocks (step Sk_14). When the processes in steps Sk_13 and Sk_14 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merge mode for the current block ends. In other words, motion compensation is performed on the current block to generate a predicted image for the current block.
[0733] Note that the MV candidate list described above may be generated in step Sj_11 as in step Sk_11.
[0734] (MV derivation > triangle pattern)
[0735] For example, when information parsed from the stream indicates that the triangle mode is to be applied, the inter predictor 218 derives an MV in the triangle mode and performs motion compensation (prediction) using the MV.
[0736] Figure 89 is a flowchart illustrating an example of a process of inter-frame prediction by triangular mode in the decoder 200 .
[0737] In triangular mode, the inter-frame predictor 218 first splits the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as a prediction parameter, which is information related to the partition. The inter-frame predictor 218 can then split the current block into the first partition and the second partition based on the partition information.
[0738] Next, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (eg, MVs of multiple decoded blocks temporally or spatially surrounding the current block) (step Sx_12). In other words, the inter-frame predictor 218 generates an MV candidate list.
[0739] The inter-frame predictor 218 then selects an MV candidate for the first partition and an MV candidate for the second partition from the multiple MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, the inter-frame predictor 218 can obtain MV selection information for identifying each selected MV candidate from the stream as a prediction parameter. The inter-frame predictor 218 can then select the first MV and the second MV based on the MV selection information.
[0740] Next, the inter-frame predictor 218 generates a first predicted image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Similarly, the inter-frame predictor 218 generates a second predicted image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).
[0741] Finally, the inter predictor 218 generates a predicted image for the current block by performing weighted addition of the first predicted image and the second predicted image (step Sx_16).
[0742] (MV estimate > DMVR)
[0743] For example, information parsed from the stream indicates that DMVR is to be applied, and the inter predictor 218 performs motion estimation using DMVR.
[0744] Figure 90 is a flowchart illustrating an example of a process of motion estimation by DMVR in the decoder 200 .
[0745] The inter-frame predictor 218 derives the MV for the current block according to the merge mode (step S1_11). Next, the inter-frame predictor 218 derives the final MV for the current block by searching the area around the reference picture indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined according to DMVR.
[0746] Figure 91 is a flowchart showing an example of a process of motion estimation by DMVR in the decoder 200, and is similar to Figure 58B same.
[0747] First, in Figure 58AIn step 1 shown in FIG, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and the eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each of the search positions other than the starting point is minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is minimum, the inter-frame predictor 218 changes the target to the search position at which the minimum cost is obtained, and performs Figure 58A The process in step 2 is shown in FIG. When the cost at the starting point is minimum, the inter-frame predictor 218 skips Figure 58A The process in step 2 shown in FIG and the process in step 3 are performed.
[0748] exist Figure 58A In step 2 shown in FIG, the inter-frame predictor 218 performs a search similar to the process in step 1, and based on the result of the process in step 1, the search position after the target change is regarded as a new starting point. The inter-frame predictor 218 then determines whether the cost at each of the search positions other than the starting point is minimized. Here, if it is determined that the cost at one of the search positions other than the starting point is minimized, the inter-frame predictor 218 performs the process in step 4. If the cost at the starting point is minimized, the inter-frame predictor 218 performs the process in step 3.
[0749] In step 4, the inter predictor 218 regards the search position at the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0750] exist Figure 58A In step 3 shown in FIG, the inter-frame predictor 218 determines a pixel position with sub-pixel accuracy that obtains the minimum cost based on the costs at four points located above, below, left, and right relative to the starting point in step 1 or step 2, and regards the pixel position as the final search position.
[0751] The pixel position with sub-pixel accuracy is determined by performing weighted addition on each of the four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) above, below, left, and right using the cost at the corresponding one of the four search positions as a weight. The inter-frame predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0752] (Motion Compensation > BIO / OBMC / LIC)
[0753] For example, when the information parsed from the stream indicates that the predicted image is to be corrected, the inter-frame predictor 218 corrects the predicted image based on the correction mode when generating the predicted image. The correction mode is, for example, one of BIO, OBMC, and LIC described above.
[0754] Figure 92 is a flowchart showing one example of a process of generating a predicted image in the decoder 200 .
[0755] The inter predictor 218 generates a predicted image (step Sm_11 ), and corrects the predicted image according to any of the modes described above (step Sm_12 ).
[0756] Figure 93 is a flowchart illustrating another example of a process of generating a predicted image in the decoder 200 .
[0757] The inter-frame predictor 218 derives the MV for the current block (step Sn_11). Next, the inter-frame predictor 218 generates a predicted image using this MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, the inter-frame predictor 218 obtains prediction parameters included in the stream and determines whether to perform a correction process based on these prediction parameters. For example, these prediction parameters are flags indicating whether to apply one or more of the modes described above. Here, if it is determined that the correction process is to be performed ("Yes" in step Sn_13), the inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). Note that in LIC, luminance and chroma can be corrected in step Sn_14. If it is determined that the correction process is not to be performed ("No" in step Sn_13), the inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).
[0758] (Motion Compensation > OBMC)
[0759] For example, when the information parsed from the stream indicates that OBMC is to be performed, the inter predictor 218 corrects the predicted image according to OBMC when generating the predicted image.
[0760] Figure 94 : is a flowchart showing an example of a process of correcting a predicted image by OBMC in the decoder 200. Note that Figure 94 The flowchart in the instructions uses Figure 62 The correction process of the current picture and the reference picture to the predicted image is shown in .
[0761] First, if Figure 62 As shown in , the inter predictor 218 obtains a predicted image (Pred) by performing normal motion compensation using the MV assigned to the current block.
[0762] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector for the current block). The inter-frame predictor 218 then performs a first correction on the predicted image by overlapping the two predicted images Pred and Pred_L. This provides an effect of blending the boundaries between adjacent blocks.
[0763] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) derived for the coded block adjacent to the current block to the current block (reusing the MV for the current block). The inter-frame predictor 218 then performs a second correction on the predicted image by overlapping the predicted image Pred_U with the predicted image (e.g., Pred and Pred_L) on which the first correction has been performed. This provides an effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.
[0764] (Motion Compensation > BIO)
[0765] For example, when the information parsed from the stream indicates that BIO is to be performed, the inter predictor 218 corrects the predicted image according to the BIO when generating the predicted image.
[0766] Figure 95 is a flowchart illustrating an example of a process of correcting a predicted image by BIO in the decoder 200 .
[0767] like Figure 63 As shown in FIG, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. Then, the inter-frame predictor 218 derives a predicted image for the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ), and the motion vector M1 is the motion vector (MV x1 , MV y1 ).
[0768] Next, the inter-frame predictor 218 uses the motion vector M0 and the reference picture L0 to derive an interpolated image I for the current block. 0Additionally, the inter-frame predictor 218 uses the motion vector M1 and the reference picture L1 to derive an interpolated image I for the current block. 1 (Step Sy_12). Here, the interpolated image I 0 is an image included in the reference picture Ref0 and derived for the current block, and the interpolated image I 1 Is an image included in the reference picture Ref1 and derived for the current block. Interpolated image I 0 and interpolated image I 1 Each of the interpolated images I may be the same size as the current block. 0 and interpolated image I 1 Each of the images can be larger than the current block. 0 and interpolated image I 1 A predicted image obtained by using a motion vector (M0, M1) and a reference picture (L0, L1) and applying a motion compensation filter may be included.
[0769] Additionally, the inter-frame predictor 218 generates an interpolated image I 0 and interpolated image I 1 Derive the gradient image of the current block (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) (Step Sy_13). Note that the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), and the gradient image in the vertical direction is (Iy 0 , Iy 1 ). The inter-frame predictor 218 may derive a gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image may be an image indicating each of a spatial change amount of a pixel value in a horizontal direction or a spatial change amount of a pixel value in a vertical direction.
[0770] Next, the inter-frame predictor 218 uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , 1x 1 , Iy 0 , Iy 1 ) derives an optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_14). As an example, the sub-block may be a sub-CU of 4×4 pixels.
[0771] Next, the inter-frame predictor 218 uses the optical flow (vx, vy) to correct the predicted image for the current block. For example, the inter-frame predictor 218 uses the optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_15). The inter-frame predictor 218 can then use the correction values to correct the predicted image for the current block (step Sy_16). Note that the correction values can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.
[0772] Note that BIO process flows are not limited to Figure 95 The process disclosed in . You can only execute Figure 95 A part of the process disclosed in the specification may be used, or a different process may be added or used instead, or the processes may be performed in a different processing order, etc.
[0773] (Motion Compensation > LIC)
[0774] For example, when the information parsed from the stream indicates that LIC is to be performed, the inter predictor 218 corrects the predicted image according to LIC when generating the predicted image.
[0775] Figure 96 is a flowchart illustrating an example of a process of correction of a predicted image by the LIC in the decoder 200 .
[0776] First, the inter predictor 218 obtains a reference image corresponding to the current block from a decoded reference picture using MV (step Sz_11 ).
[0777] Next, the inter-frame predictor 218 extracts information indicating how the luminance value of the current block has changed between the current picture and the reference picture (step Sz_12). This extraction can be performed based on the luminance pixel values of the encoded left-neighboring reference region (surrounding reference region) and the encoded upper-neighboring reference region (surrounding reference region), as well as the luminance pixel values at corresponding positions in the reference picture specified by the derived MV. The inter-frame predictor 218 uses this information indicating how the luminance value has changed to calculate the luminance correction parameter (step Sz_13).
[0778] The inter-frame predictor 218 generates a predicted image for the current block by performing an illumination correction process in which illumination correction parameters are applied to the reference image in the reference picture specified by the MV (step Sz_14). In other words, the predicted image is corrected based on the illumination correction parameters, and the predicted image is used as the reference image in the reference picture specified by the MV. This correction can be performed for illumination or color.
[0779] (Predictive Controller)
[0780] The prediction controller 220 selects an intra-frame predicted image or an inter-frame predicted image and outputs the selected image to the adder 208. In general, the configuration, function, and process of the prediction controller 220, the intra-frame predictor 216, and the inter-frame predictor 218 on the decoder 200 side may correspond to the configuration, function, and process of the prediction controller 128, the intra-frame predictor 124, and the inter-frame predictor 126 on the encoder 100 side.
[0781] (decoding using predicted chroma samples)
[0782] In a first aspect, a determination is made as to whether a luma sample can be used to predict a block of chroma samples of a current block, wherein the block is decoded using the predicted chroma samples. For example, an embodiment may employ a process of using a decoding result of a luminance signal in a decoding method or encoding method to determine whether to enable a tool such as CCLM to predict a color difference signal.
[0783] Figure 97 is a flow chart illustrating one example of a process 1000 for decoding a block using predicted chroma samples, which may be performed, for example, by Figure 7 Encoder 100 or Figure 67 For convenience, reference will be made to the decoder 200. Figure 67 The decoder 200 is described Figure 97 .
[0784] At S1001 , the decoder 200 determines whether the current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples. Figure 99 and Figure 100 is a conceptual diagram illustrating an example of determining whether the current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples. In some formats (e.g., YUV420 format), a 16×16 pixel region of chroma corresponds to a 32×32 pixel region of luminance. Figure 99 and Figure 100 As shown in , chroma blocks within a 32×32 luma region aligned with a 16×16 chroma grid are determined to be within an M×N non-overlapping region aligned with an M×N grid of chroma samples. Chroma blocks that are not within a 32×32 luma region are not determined to be within an M×N non-overlapping region aligned with an M×N grid of chroma samples. Even if the current chroma block straddles the boundary of the corresponding luma block, luma samples can be used to obtain chroma samples if it is included in the same VPDU. For example, Figure 99 Chroma block A in uses corresponding samples in luma block B to predict samples in chroma block A-1, and uses corresponding samples in luma block C to predict samples in chroma block A-2.
[0785] like Figure 100As shown in , the chroma samples of the shown chroma blocks can be predicted using the luma samples for the chroma blocks because the chroma blocks are included in a grid (a 16×16 grid as shown) and the co-located luma blocks are also inside the co-located 32×32 region.
[0786] In some embodiments, for example, by default, when other conditions such as those discussed below with reference to S1002 are met, etc., luma samples may not be used to predict blocks of chroma samples that are not determined to be within an M×N non-overlapping region aligned with an M×N grid of chroma samples, but luma samples may be used to predict blocks of chroma samples that are determined to be within an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0787] like Figure 97 As shown in FIG, when it is not determined at S1001 that the current chroma block is within the M×N non-overlapping region aligned with the M×N grid of chroma samples, process 1000 proceeds from S1001 to S1004, in which the decoder 200 predicts the chroma sample block without using luma samples. Process 1000 proceeds from S1004 to S1005, in which the decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1001 that the current chroma block is within the M×N non-overlapping region, process 1000 proceeds from S1001 to S1002.
[0788] At S1002, the decoder 200 determines whether to split the current luma VPDU into smaller blocks. A VPDU is a unit processed in parallel during encoding or decoding, for example, a size of 64×64. The size of a VPDU may be determined by a standard or may be encoded in the stream.
[0789] Whether to split the current luma VPDU into smaller blocks can be determined in various ways, and the following reference is made to Figure 102 and Figure 103 Let's discuss some examples in more detail.
[0790] When it is not determined at S1002 that the current luma VPDU is to be split into smaller blocks, the process 1000 proceeds from S1002 to S1004, where the decoder 200 predicts the chroma sample block without using the luma samples. The process 1000 proceeds from S1004 to S1005, where the decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1002 that the current luma VPDU is to be split into smaller blocks, the process 1000 proceeds from S1002 to S1003, where the decoder 200 uses the luma samples to predict the chroma sample block. The process 1000 proceeds from S1003 to S1005, where the decoder 200 decodes the block using the predicted chroma samples. In some embodiments, additional considerations may be taken into account to determine whether to use luma samples to decode the chroma sample block, for example, as described below with reference to Figures 104 to 110 discussed.
[0791] Figure 98 is a flow chart illustrating another example of a process 2000 for decoding a block using predicted chroma samples, which may be performed, for example, by Figure 7 Encoder 100 or Figure 67 For convenience, reference will be made to the decoder 200. Figure 67 The decoder 200 is described Figure 98 .
[0792] At S2001, the decoder 200 determines whether to split the first VPDU and the second VPDU into smaller blocks. Whether to split the current luma VPDU into smaller blocks can be determined in various ways and is described below with reference to Figure 102 and Figure 103 Let's discuss some examples in more detail.
[0793] When it is determined at S2001 that the first VPDU is not to be split into smaller blocks and the second VPDU is to be split into smaller blocks, process 2000 proceeds from S2001 to S2002, where decoder 200 predicts the chroma sample block without using luma samples. Process 2000 proceeds from S2002 to S2004, where decoder 200 decodes the block using the predicted chroma samples.
[0794] When it is not determined at S2001 that the first luma VPDU is not to be split into smaller blocks and the second VPDU is to be split into smaller blocks, process 2000 proceeds from S2001 to S2003, in which decoder 200 uses luma samples to predict a block of chroma samples. Process 2000 proceeds from S2003 to S2004, in which decoder 200 decodes the block using the predicted chroma samples. In some embodiments, additional considerations may be taken into account to determine whether to use luma samples to decode a block of chroma samples, for example, as described below with reference to Figures 104 to 110 discussed.
[0795] Figure 101 This is a conceptual diagram illustrating VPDUs. VPDUs are non-overlapping and represent the buffer size of the pipeline stages. Figure 101 The left side of FIG (labeled a) shows an example of a 128×128 CTU with four 64×64 VPDUs. Figure 101 The right side of (marked as b) shows an example of a 128×128 CTU with 16 32×32 VPDUs. For example, if the VPDU is 64×64, both M and N are set to 16. When the VPDU is to be further divided, the CU size after division becomes 2M×2N (32×32) or smaller. In the YUV420 format, the 16×16 chroma area corresponds to the 32×32 luminance area, so the pixels in the 16×16 grid of chroma can be predicted based on the pixels of the corresponding 32×32 grid in luminance. Therefore, when the decoding of the 32×32 area of luminance is completed, the prediction process of luminance chroma can be started on the 16×16 area of chroma. In the case of the YUV444 format, the luminance M×N area corresponds to the chroma M×N area. If 2M×2N is half the size of the VPDU in both horizontal and vertical directions, then in Figure 97 In step S1002 or Figure 98 In step S2001, it may be determined whether the VPDU is further divided into one or more layers, but this is performed in the case of 1 / 4 of the VPDU. The embodiment may determine whether the divided CU becomes 2M×2N or smaller, for example, whether the CU is further divided into two or more layers.
[0796] Figure 102 This is a conceptual diagram illustrating an example of determining whether the current VPDU can use luma samples to predict chroma sample blocks based on whether the luma VPDU is split into blocks, wherein the left side shows the luma CTU and the right side shows the corresponding chroma CTU. As shown, luma VPDU0 will be split into blocks, and luma VPDU1 will not be split into blocks. Therefore, referring to Figure 97According to process 1000 , luma samples may be used to predict chroma samples of VPDU0 , and luma samples may not be used to predict chroma samples of VPDU1 .
[0797] Figure 103 is a conceptual diagram illustrating two example ways of determining whether a luma VPDU is to be split into smaller blocks. Figure 103 In the first example (labeled a) shown on the left, whether the luma VPDU is to be split can be determined based on the split flag associated with the luma VPDU. As shown, when the split flag has a value of 1, the VPDU is to be split (and, with reference to Figure 97 1000, using luma samples to predict chroma samples of a block). When the split flag has a value of 0, the VPDU is not split (and, see Figure 97 (The process 1000 does not use luma samples to predict chroma samples of the block.) Other split flag values may be used to determine whether the luma VPDU is split.
[0798] exist Figure 103 In the second example (labeled b) shown on the right, whether the luma VPDU is to be split can be determined based on the quadtree split depth of the luma block of the VPDU. As shown, the quadtree split depth of the luma block of VPDU0 is greater than 1, so the reference Figure 97 In the process 1000, when decoding the block of VPDU0, the luma samples can be used to predict the chroma samples. In contrast, the quadtree split depth of the block of VPDU1 is less than or equal to 1, so the reference Figure 97 In the process 1000, when decoding the VPDU0 block, the luma samples may not be used to predict the chroma samples. Other split depth values may be used to determine whether the luma VPDU is split.
[0799] Figure 104 is a conceptual diagram illustrating additional considerations that may be taken into account to determine whether to use luma samples to predict chroma samples of a block. As shown, whether the current block size is equal to or less than a threshold block size may be employed as an additional consideration in determining whether to use luma samples to predict chroma samples of a block.
[0800] The threshold block size can be a default block size, a signaled block size, or a determined block size, and can be a luma or chroma block size. For example, if the threshold block size is a 16×16 luma block size, the luma block size of VPDU0 is larger than 16×16, so it can be determined that luma samples are not used to determine the chroma samples of the block. Figure 97 S1002 or Figure 98 At S2001 , a threshold block size is used to determine whether to split the luma VPDU into smaller blocks.
[0801] Figure 97 Aspects of the Process 1000 and Figure 98 Aspects of process 2000 can be modified in various ways. For example, process 1000 or 2000 can be modified to perform more actions than shown, can be modified to perform fewer actions than shown, can be modified to perform actions in a different order, can be modified to combine or split actions, etc. For example, before S1001 or S1002, process 1000 can be modified to be more specific based on other considerations (e.g., reference Figure 103 In another example, the process 1000 may be modified to omit S1001. In another example, Figure 98 The embodiment of process 2000 may be modified to perform step S1001 before performing step S2001. In another example, S2001 may determine whether the first VPDU and the second VPDU are split into smaller blocks.
[0802] Figure 105 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. Figure 105 , an example combined condition is whether both the luma VPDU and the corresponding chroma VPDU have a quadtree split depth greater than or equal to 2. Luma VPDU0 has a quadtree split depth greater than or equal to 2, and chroma VPDU0 has a quadtree split depth greater than or equal to 2, so the chroma samples for chroma VPDU0 can be predicted using luma samples. However, luma VPDU1 has a quadtree split depth not greater than or equal to 2, so one of the conditions is not met, and the chroma samples for chroma VPDU1 will be predicted without using luma samples.
[0803] Figure 106 is a conceptual diagram for illustrating another example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. Figure 106As shown in , the example condition combination is: (i) whether the luma VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDU quadtree split depth is equal to 1; and (iii) whether the chroma split threshold condition of 32×32 is met (for example, when the chroma size is 32×32, the block is not split). Luma VPDU0 has a quadtree split depth greater than or equal to 2, which satisfies condition (i); chroma VPDU0 has a quadtree split depth equal to 1, which satisfies condition (ii), and the chroma VPDU is not split into blocks of size less than 32×32, so all three conditions are met and the chroma samples of VPDU0 can be predicted using luma samples. However, the size of the block of chroma VPDU1 is less than the 32×32 threshold, so condition (iii) is not met, and the chroma samples for VPDU1 will be predicted without using luma samples.
[0804] Figure 107 is a conceptual diagram for illustrating another example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. Figure 107 As shown in , the combination of example conditions is: (i) whether the luma VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDU quadtree split depth is equal to 1; and (iii) whether the chroma split threshold condition of 32×32 is met (for example, when the chroma size is 32×32, the block is not split). Figure 107 qtDepthC in indicates the chroma quadtree split depth, and Figure 107 mtDepthC in _ indicates the chroma multitree split depth. A quadtree split can be followed by another quadtree split or a multitree type split (binary or ternary split). To specify that the chroma quadtree split terminates at depth 1, add the condition chromaSplit32×32 == CU_DONT_SPLIT (refer to Figure 106 Condition iii) discussed, this means that there is no further splitting at the chroma 32×32 level. Assuming that the luma quadtree split depth qtDepthl is greater than or equal to 2, only chroma VPDU0 meets all three conditions, and the chroma samples for VPDU0 can be predicted using luma samples. The chroma quadtree split depth of chroma VPDU1 is 2, and the blocks are split into blocks smaller than 32×32, so the chroma samples for VPDU1 will be predicted without using luma samples. The chroma quadtree split depth of chroma VPDU2 is 1, but the blocks are split into blocks smaller than 32×32, so the chroma samples for VPDU2 will be predicted without using luma samples. The chroma quadtree split depth of chroma VPDU3 is 1, but the blocks are split into blocks smaller than 32×32, so the chroma samples for VPDU3 will be predicted without using luma samples.
[0805] Figure 108 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luma samples to predict chroma samples for a block. In this example, the combination of conditions is: (i) whether the luma VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDU quadtree split depth is equal to 1; and (iii) whether a horizontal chroma split of size 32×32 is followed by no vertical or horizontal ternary split. For VPDU0, the conditions are met; the VPDU is split horizontally into two 16×32 blocks, and these blocks are not further split using horizontal or vertical ternary splits. Therefore, luma samples can be used to predict chroma samples in all blocks of VPDU0. For VPDU1, the lower 16×32 block meets the conditions, no further ternary split is performed, and luma samples can be used to predict chroma samples for the lower 16×32 block. The upper 16×32 block of VPDU1 does not meet the conditions because a further vertical ternary split is performed, and chroma samples for the upper 16×32 block of VPDU1 are predicted without using luma samples.
[0806] Figure 109 is a conceptual diagram for illustrating another example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. Figure 109 As shown in , the example condition combinations are: (i) whether the luma VPDU quadtree split depth is equal to 1; (ii) whether the luma split threshold condition of 64×64 is met (e.g., when the luma size is 64×64, the block is not split); (iii) whether the corresponding chroma VPDU quadtree split depth is equal to 1; and (iv) whether the chroma split threshold condition of 32×32 is met (e.g., when the chroma size is 32×32, the block is not split). VPDU0 meets all four conditions, and the chroma samples for VPDU0 can be predicted using luma samples. Chroma VPDU1 has a chroma quadtree split depth of 2, and the block is split into blocks smaller than 32×32, so the chroma samples for VPDU1 will be predicted without using luma samples.
[0807] Figure 110 is a conceptual diagram for illustrating another example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. Figure 110As shown in , if any of the conditions is true, the luma samples can be used to predict the chroma samples in the block. Example condition combinations are: (i) whether the luma VPDU quadtree split depth is greater than or equal to 2, and whether the chroma VPDU quadtree split depth is greater than or equal to 2; (ii) whether the luma VPDU quadtree split depth is equal to 1, and the luma split threshold condition of 64×64 is met (for example, when the luma size is 64×64, the block is not split), and the corresponding chroma VPDU quadtree split depth is equal to 1, and the chroma split threshold condition of 32×32 is met (for example, when the chroma size is 32×32, the block is not split); (iii) Whether the luma VPDU quadtree split depth is greater than or equal to 2, the corresponding chroma VPDU quadtree split depth is equal to 1, and the 32×32 chroma split threshold condition is met (for example, when the chroma size is 32×32, the block is not split); and (iv) Whether the luma VPDU quadtree split depth is greater than or equal to 2, the corresponding chroma VPDU quadtree split depth is equal to 1, the chroma split of the 32×32 block is horizontal, and chroma blocks smaller than 32×32 are either not split or split vertically. The block of VPDU1 violates all four conditions, and therefore the chroma samples of VPDU1 will be predicted without using luma samples. Considering the scanning order, it can be used Figure 110 The example condition of is used to limit the chroma prediction latency (based on luma samples) within 32×32 samples. For example, in VPDU1, chroma block 0 must wait for luma block 0 reconstruction to be predicted. Chroma block 1 must wait for luma blocks 0 and 1 reconstruction to be predicted. To avoid this latency, chroma prediction can be performed without using luma samples.
[0808] The blocks described in each of the aspects may be replaced with rectangular or non-rectangular shaped partitions. Figure 111 Examples of non-rectangular partitions are shown, for example, triangular partitions, L-shaped partitions, pentagonal partitions, hexagonal partitions, and polygonal partitions. Other non-rectangular partitions can be used, and combinations of various shapes can be used. The term "partition" described in each of the aspects can be replaced with the term "prediction unit." The term "partition" described in each of the aspects can also be replaced with the term "sub-prediction unit." The term "partition" described in each of the aspects can also be replaced with the term "decoding unit."
[0809] Other conditions may be used. For example, in an embodiment, when a coding mode for predicting color difference based on illumination (e.g., CCLM) is enabled, the first partition of the VPDU may always be a quaternary partition. In another embodiment, when a certain number of quaternary partitions are not applied to a VPDU in at least one VPDU in a CTU (e.g., the head VPDU of a CTU in scan order), CCLM may be disabled in all VPDUs in the CTU.
[0810] CCLM can be defined as an intra-frame prediction mode using mode information such as intra_chroma_pred_mode. The index number indicating the intra-frame prediction mode and each mode can be associated one-to-one in a table, but when CCLM is disabled, the entry of the table for CCLM is unnecessary, so the index number is encoded, thereby facilitating a reduction in the number of bits for encoding the signal. In an embodiment, the table indicating the intra-frame prediction mode can be switched depending on whether CCLM is valid or invalid. For example, with reference to the division flag information indicating the quad division of illuminance, if the illuminance is not divided into a certain size or smaller in the VPDU, it can be determined that CCLM is invalid, and a table corresponding to the case where CCLM is invalid is adopted. Otherwise, the corresponding table used when CCLM can be used can be adopted. In an embodiment, a table including entries for CCLM can be used without switching the table, but when CCLM is invalid, the entry for CCLM can be not referenced.
[0811] refer to Figure 98For example, in some embodiments, the current block may be in a first VPDU. In some embodiments, the current block may be in a second VPDU. In some embodiments, the fir...
Claims
1. An encoder, comprising: Circuit; as well as a memory coupled to the circuit; The circuit performs the following operations during operation: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting the block of chroma samples using luma samples; and The block is encoded using the predicted chroma samples.
2. An encoder comprising: a block splitter operable to split the first image into a plurality of blocks; an intra predictor that, in operation, predicts a block comprised in said first image using a reference block comprised in said first image; an inter-frame predictor that, in operation, predicts a block included in the first image using a reference block included in a second image different from the first image; a loop filter operable to filter blocks comprised in said first image; a transformer operative to transform a prediction error between an original signal and a prediction signal generated by the intra predictor or the inter predictor to generate transform coefficients; a quantizer operative to quantize the transform coefficients to generate quantized coefficients; as well as an entropy encoder operative to variably encode the quantized coefficients to generate an encoded bitstream comprising the encoded quantized coefficients and control information; Among them, the prediction block includes: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; and In response to determining that the first virtual pipelined decoding unit is not split into smaller blocks and determining that the second virtual pipelined decoding unit is not split into smaller blocks, the block of chroma samples is predicted using luma samples.
3. A decoder comprising: Circuit; a memory coupled to the circuit; The circuit performs the following operations during operation: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting the block of chroma samples using luma samples; and The block is decoded using the predicted chroma samples.
4. A decoding device comprising: a decoder that, in operation, decodes the encoded bitstream to output quantized coefficients; an inverse quantizer operable to inverse quantize the quantized coefficients to output transform coefficients; an inverse transformer operable to inversely transform the transform coefficients to output a prediction error; an intra predictor that, in operation, uses a reference block included in a first image to predict a block included in said first image; an inter-frame predictor that, in operation, predicts a block included in the first image using a reference block included in a second image different from the first image; a loop filter operable to filter blocks comprised in said first image; as well as output, which in operation outputs a picture including the first image, Among them, the prediction block includes: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; and In response to determining that the first virtual pipelined decoding unit is not split into smaller blocks and determining that the second virtual pipelined decoding unit is not split into smaller blocks, the block of chroma samples is predicted using luma samples.
5. A coding method comprising: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting the block of chroma samples using luma samples; as well as The block is encoded using the predicted chroma samples.
6. A decoding method comprising: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; In response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting the block of chroma samples using luma samples; as well as The block is decoded using the predicted chroma samples.