Systems and methods for video coding
By using predicted chroma sample blocks in video decoding, the method determines whether to use luminance samples for prediction based on the splitting of the virtual pipeline decoding unit. This solves the problems of low coding efficiency, poor image quality, and large circuit scale in existing technologies, and achieves more efficient encoding and decoding.
Patent Information
- Application Number
- CN202511243163.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-18
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video decoding technologies struggle to effectively improve coding efficiency, enhance image quality, and reduce circuit size when dealing with increasing volumes of digital video data.
By using predicted chroma sample blocks in video decoding, the method determines whether to use luminance samples for prediction based on whether the virtual pipeline decoding unit is split into smaller blocks, thereby achieving the encoding and decoding of blocks.
It improves encoding efficiency, enhances image quality, reduces the utilization of encoding/decoding processing resources, and reduces circuit size.
Smart Images

Figure CN120956887A_ABST
Abstract
Description
[0001] This application is a divisional application of the same patent application, filed on June 18, 2020, with application number 202080044292.2. Technical Field
[0002] This disclosure relates to video coding, and more particularly to video coding and decoding systems, components, and methods, for example, for performing encoding of blocks using predicted chroma samples. Background Technology
[0003] With advancements in video decoding technology from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Codec), MPEG-LA, H.265 / HEVC (High-Efficiency Video Codec), and H.266 / VVC (Multi-Functional Video Codec), there remains a continuous need for improvements and optimizations to video decoding technologies to handle the ever-increasing volumes of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video decoding, particularly in the use of predicted chroma samples to perform block encoding. Summary of the Invention
[0004] In one aspect, an encoder includes: circuitry; and a memory coupled to the circuitry. The circuitry determines whether a first Virtual Pipeline Decoding Unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is split into smaller blocks, a chroma sample block is predicted without using luma samples. In response to determining that the first VPDU is split into smaller blocks and that the second VPDU is split into smaller blocks, a chroma sample block is predicted using luma samples. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is not split into smaller blocks, a chroma sample block is predicted using luma samples. The predicted chroma samples are used to encode the block.
[0005] In one aspect, an encoder includes: a block splitter that splits a first image into multiple blocks in operation; an intra predictor that predicts blocks included in the first image using reference blocks included in the first image in operation; an inter predictor that predicts blocks included in the first image using reference blocks included in a second image different from the first image in operation; a loop filter that filters the blocks included in the first image in operation; a transformer that transforms a prediction error between the original signal and a prediction signal generated by the intra predictor or the inter predictor in operation to generate transform coefficients; a quantizer that quantizes the transform coefficients in operation to generate quantized coefficients; and an entropy encoder that variablely encodes the quantized coefficients in operation to generate an encoded bitstream, the encoded bitstream including the encoded quantized coefficients and control information. The prediction block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU has not been split into smaller blocks and that the second VPDU has been split into smaller blocks, a chroma sample block is predicted without using luma samples. In response to determining that the first VPDU has been split into smaller blocks and that the second VPDU has been split into smaller blocks, a chroma sample block is predicted using luma samples. In response to determining that the first VPDU has not been split into smaller blocks and that the second VPDU has not been split into smaller blocks, a chroma sample block is predicted using luma samples. The predicted chroma samples are then used to encode the block.
[0006] In one aspect, a decoder includes: circuitry; and memory coupled to the circuitry. The circuitry determines whether a first Virtual Pipeline Decoding Unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is split into smaller blocks, a chroma sample block is predicted without using luminance samples. In response to determining that the first VPDU is split into smaller blocks and that the second VPDU is split into smaller blocks, a chroma sample block is predicted using luminance samples. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is not split into smaller blocks, a chroma sample block is predicted using luminance samples. The predicted chroma samples are used to decode the block.
[0007] In one aspect, a decoding apparatus includes: a decoder that decodes an encoded bitstream in operation to output quantized coefficients; an inverse quantizer that in operation inverse-quantizes the quantized coefficients to output transform coefficients; an inverse transformer that in operation inversely transforms the transform coefficients to output a prediction error; an intra-frame predictor that in operation predicts a block included in a first image using a reference block included in a first image; an inter-frame predictor that in operation predicts a block included in the first image using a reference block included in a second image different from the first image; a loop filter that in operation filters the block included in the first image; and an output that in operation outputs a picture including the first image. The predicted block includes determining whether a first Virtual Pipeline Decoding Unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, a chroma sample block is predicted without using luminance samples. In response to determining that the first VPDU has been split into smaller blocks and that the second VPDU has been split into smaller blocks, luminance samples are used to predict chroma sample blocks. In response to determining that the first VPDU has not been split into smaller blocks and that the second VPDU has not been split into smaller blocks, luminance samples are used to predict chroma sample blocks. The predicted chroma samples are then used to decode the blocks.
[0008] In one aspect, an encoding method includes: determining whether a first Virtual Pipeline Decoding Unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is split into smaller blocks, predicting chroma sample blocks without using luma samples. In response to determining that the first VPDU is split into smaller blocks and that the second VPDU is split into smaller blocks, predicting chroma sample blocks using luma samples. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is not split into smaller blocks, predicting chroma sample blocks using luma samples. Encoding the blocks using the predicted chroma samples.
[0009] In one aspect, a decoding method includes: determining whether a first Virtual Pipeline Decoding Unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is split into smaller blocks, predicting chroma sample blocks without using luma samples. In response to determining that the first VPDU is split into smaller blocks and that the second VPDU is split into smaller blocks, predicting chroma sample blocks using luma samples. In response to determining that the first VPDU is not split into smaller blocks and that the second VPDU is not split into smaller blocks, predicting chroma sample blocks using luma samples. Decoding the blocks using the predicted chroma samples.
[0010] In video decoding technology, there is a desire to propose new methods to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of embodiments of this disclosure, including constituent elements of embodiments of this disclosure considered individually or in various combinations, can facilitate one or more of the following: improved coding efficiency; enhanced image quality; reduced utilization of processing resources associated with encoding / decoding; reduced circuit size; improved encoding / decoding processing speed, etc.
[0011] Additionally, some implementations of embodiments of this disclosure, including constituent elements of embodiments of this disclosure considered individually or in various combinations, can facilitate appropriate selection of one or more elements (e.g., filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations) during encoding and decoding. Note that this disclosure includes information regarding configurations and methods that can provide advantages beyond those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing the increase in processing resource usage.
[0012] Additional benefits and advantages of the disclosed embodiments will become apparent from the specification and accompanying drawings. Benefits and / or advantages may be obtained individually from the various embodiments and features in the specification and drawings, and it is not necessary to provide all embodiments and features to obtain one or more of such benefits and / or advantages.
[0013] It should be noted that general or specific embodiments may be implemented as systems, methods, integrated circuits, computer programs, storage media, or any alternative combination thereof. Attached Figure Description
[0014] Figure 1 This is a schematic diagram illustrating an example of the functional configuration of a transmission system according to an embodiment.
[0015] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream.
[0016] Figure 3 This is a conceptual diagram used to illustrate an example of a slice configuration.
[0017] Figure 4 This is a conceptual diagram used to illustrate an example of a tile configuration.
[0018] Figure 5 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.
[0019] Figure 6 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.
[0020] Figure 7 This is a block diagram illustrating the functional configuration of an encoder according to an embodiment.
[0021] Figure 8 This is a functional block diagram illustrating an example of encoder installation.
[0022] Figure 9 It is a flowchart that indicates an example of the overall encoding process performed by the encoder.
[0023] Figure 10 This is a conceptual diagram used to illustrate an example of block splitting.
[0024] Figure 11 This is a block diagram illustrating an example of the functional configuration of a splitter according to an embodiment.
[0025] Figure 12 This is a conceptual diagram used to illustrate an example of a splitting pattern.
[0026] Figure 13A This is a conceptual diagram used to illustrate an example of a syntactic tree for splitting patterns.
[0027] Figure 13B This is a conceptual diagram used to illustrate another example of a syntactic tree for splitting patterns.
[0028] Figure 14 It is a graph indicating example transformation basis functions used for various transformation types.
[0029] Figure 15 This is a conceptual diagram used to illustrate an example spatial transformation (SVT).
[0030] Figure 16 This is a flowchart illustrating an example of a process performed by a converter.
[0031] Figure 17 This is a flowchart illustrating another example of a process performed by a converter.
[0032] Figure 18 This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment.
[0033] Figure 19 This is a flowchart illustrating an example of the quantization process performed by a quantizer.
[0034] Figure 20 This is a block diagram illustrating an example of the functional configuration of an entropy encoder according to an embodiment.
[0035] Figure 21 This is a conceptual diagram illustrating an example flow of the context-based adaptive binary arithmetic decoding (CABAC) process in an entropy encoder.
[0036] Figure 22 This is a block diagram illustrating an example of the functional configuration of a loop filter according to an embodiment.
[0037] Figure 23A This is a conceptual diagram used to illustrate an example of the filter shape used in an adaptive loop filter (ALF).
[0038] Figure 23B This is a conceptual diagram used to illustrate another example of the filter shape used in ALF.
[0039] Figure 23C This is a conceptual diagram used to illustrate another example of the filter shape used in ALF.
[0040] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF).
[0041] Figure 23E This is a conceptual diagram used to illustrate an example of the filter shape used in CC-ALF.
[0042] Figure 23F This is a conceptual diagram used to illustrate an example process of Joint Chromaticity CCALF (JC-CCALF).
[0043] Figure 23G This is a table that shows example weighted index candidates that can be used in JC-CCALF.
[0044] Figure 24 This is a block diagram illustrating an example of the specific configuration of a loop filter used as a deblocking filter (DBF).
[0045] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics about block boundaries.
[0046] Figure 26It is a conceptual diagram used to illustrate the block boundaries on which the deblocking filtering process is performed.
[0047] Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.
[0048] Figure 28 This is a flowchart illustrating an example of the process performed by the encoder's predictor.
[0049] Figure 29 This is a flowchart illustrating another example of the process performed by the encoder's predictor.
[0050] Figure 30 This is a flowchart illustrating another example of the process performed by the encoder's predictor.
[0051] Figure 31 This is a conceptual diagram used to illustrate the sixty-seven intra-prediction modes used in the intra-prediction in the embodiments.
[0052] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor.
[0053] Figure 33 This is a concept diagram used to illustrate a reference image.
[0054] Figure 34 This is a concept diagram used to illustrate a list of reference images.
[0055] Figure 35 This is a flowchart illustrating the basic processing flow of an example inter-frame prediction.
[0056] Figure 36 This is a flowchart illustrating an example of the derivation process of the motion vector.
[0057] Figure 37 This is a flowchart illustrating another example of the derivation process of the motion vector.
[0058] Figure 38A This is a conceptual diagram used to illustrate example representations of patterns used for MV derivation.
[0059] Figure 38B This is a conceptual diagram used to illustrate example representations of patterns used for MV derivation.
[0060] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in normal inter-frame mode.
[0061] Figure 40 This is a flowchart illustrating an example of the inter-frame prediction process in normal merging mode.
[0062] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process in the merging mode.
[0063] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process for the current image using the HMVP merging pattern.
[0064] Figure 43 This is a flowchart illustrating an example of the frame rate upconversion (FRUC) process.
[0065] Figure 44 This is a conceptual diagram used to illustrate an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.
[0066] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image.
[0067] Figure 46A This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block based on the motion vectors of multiple adjacent blocks.
[0068] Figure 46B This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block in an affine pattern using three control points.
[0069] Figure 47A This is a conceptual diagram used to illustrate an example MV derivation at the control point in affine mode.
[0070] Figure 47B This is a conceptual diagram used to illustrate an example MV derivation at the control point in affine mode.
[0071] Figure 47C This is a conceptual diagram used to illustrate an example MV derivation at the control point in affine mode.
[0072] Figure 48A This is a conceptual diagram used to illustrate an affine pattern that uses two control points.
[0073] Figure 48B This is a conceptual diagram used to illustrate an affine pattern that uses three control points.
[0074] Figure 49A This is a conceptual diagram illustrating an example of a method for MV derivation at control points when the number of control points used for the encoded block and the number of control points used for the current block are different from each other.
[0075] Figure 49BThis is a conceptual diagram illustrating another example of the method used for MV derivation at control points when the number of control points used for the encoded block and the number of control points used for the current block are different from each other.
[0076] Figure 50 This is a flowchart illustrating an example of a process in affine merging mode.
[0077] Figure 51 This is a flowchart illustrating an example of the process in affine inter-frame mode.
[0078] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangle prediction images.
[0079] Figure 52B This is a conceptual diagram used to illustrate an example of a first portion of a first partition that overlaps with a second partition, as well as a first set and a second set of samples that can be weighted as part of a correction process.
[0080] Figure 52C This is a conceptual diagram used to illustrate the first part of the first partition, which is a portion of the first partition that overlaps with a portion of an adjacent partition.
[0081] Figure 53 This is a flowchart illustrating an example of a process in triangle mode.
[0082] Figure 54 This is a conceptual diagram used to illustrate an example of an advanced temporal motion vector prediction (ATMVP) pattern in which MV is derived on a sub-block basis.
[0083] Figure 55 This is a flowchart illustrating the relationship between merge mode and Dynamic Motion Vector Refresh (DMVR).
[0084] Figure 56 This is a conceptual diagram used to illustrate an example of DMVR.
[0085] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV.
[0086] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in DMVR.
[0087] Figure 58B This is a flowchart illustrating an example of the motion estimation process in a DMVR.
[0088] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.
[0089] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.
[0090] Figure 61 This is a flowchart illustrating an example of the process of correcting a predicted image using Overlapping Block Motion Compensation (OBMC).
[0091] Figure 62 This is a conceptual diagram used to illustrate an example of the predictive image correction process performed via OBMC.
[0092] Figure 63 It is a conceptual diagram used to illustrate a model that assumes uniform linear motion.
[0093] Figure 64 This is a flowchart illustrating an example of the process of inter-frame prediction based on BIO.
[0094] Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor that can perform inter-frame prediction based on BIO.
[0095] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses an luminance correction process performed by a LIC.
[0096] Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.
[0097] Figure 67 This is a block diagram illustrating the functional configuration of the decoder according to an embodiment.
[0098] Figure 68 This is a functional block diagram illustrating an example of decoder installation.
[0099] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by the decoder.
[0100] Figure 70 It is a conceptual diagram used to illustrate the relationship between the split determinant and other constituent elements.
[0101] Figure 71 This is a block diagram illustrating an example of the functional configuration of an entropy decoder.
[0102] Figure 72 This is a conceptual diagram used to illustrate an example flow of the CABAC process in the entropy decoder.
[0103] Figure 73 This is a block diagram illustrating an example of the functional configuration of an inverse quantizer.
[0104] Figure 74 This is a flowchart illustrating an example of the dequantization process performed by the dequantizer.
[0105] Figure 75 This is a flowchart illustrating an example of a process performed by an inverse transformer.
[0106] Figure 76 This is a flowchart illustrating another example of the process performed by the inverse converter.
[0107] Figure 77 This is a block diagram illustrating an example of the functional configuration of a loop filter.
[0108] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of the decoder.
[0109] Figure 79 This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0110] Figure 80 This is a flowchart illustrating another example of the process performed by the predictor of the decoder.
[0111] Figure 81 This is a diagram illustrating an example of the process performed by the decoder's intra-frame predictor.
[0112] Figure 82 This is a flowchart illustrating an example of the MV derivation process in the decoder.
[0113] Figure 83 This is a flowchart illustrating another example of the MV derivation process in the decoder.
[0114] Figure 84 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using normal inter-frame mode.
[0115] Figure 85 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using normal merging mode.
[0116] Figure 86 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using FRUC mode.
[0117] Figure 87 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using an affine merging mode.
[0118] Figure 88 This is a flowchart illustrating an example of the process of inter-frame prediction in the decoder using affine inter-frame patterns.
[0119] Figure 89 This is a flowchart illustrating an example of the process of inter-frame prediction using triangle patterns in the decoder.
[0120] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed via DMVR in the decoder.
[0121] Figure 91 This is a flowchart illustrating an example process of motion estimation via DMVR in the decoder.
[0122] Figure 92 This is a flowchart illustrating an example of the process of generating a predicted image in the decoder.
[0123] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in the decoder.
[0124] Figure 94 This is a flowchart illustrating an example of the process of correcting the predicted image using OBMC in the decoder.
[0125] Figure 95 This is a flowchart illustrating an example of the process of correcting the predicted image using BIO in the decoder.
[0126] Figure 96 This is a flowchart illustrating an example of the process of correcting the predicted image via LIC in the decoder.
[0127] Figure 97 This is a flowchart illustrating an example of the process of decoding a block using predicted chroma samples.
[0128] Figure 98 This is a flowchart illustrating an example of the process of decoding a block using predicted chroma samples.
[0129] Figure 99 This is a conceptual diagram used to illustrate an example of determining whether the current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0130] Figure 100 This is a conceptual diagram used to illustrate an example of determining whether the current chroma block is inside an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0131] Figure 101 This is a conceptual diagram used to illustrate a Virtual Pipeline Decoding Unit (VPDU).
[0132] Figure 102 This is a conceptual diagram used to illustrate an example of determining whether a current VPDU can use luminance samples to predict chrominance sample blocks.
[0133] Figure 103 This is a conceptual diagram used to illustrate an example of how to determine whether a luminance VPDU should be split into smaller blocks.
[0134] Figure 104 This is a conceptual diagram used to illustrate additional considerations that can be taken into account to determine whether to use luminance samples to predict the chromaticity samples of a block.
[0135] Figure 105 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0136] Figure 106 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0137] Figure 107 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0138] Figure 108 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0139] Figure 109 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0140] Figure 110 This is a conceptual diagram used to illustrate an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples of a block.
[0141] Figure 111 This is a conceptual diagram used to illustrate an example of a non-rectangular partition.
[0142] Figure 112 This is a diagram illustrating an example overall configuration of a content delivery system used to implement a content distribution service.
[0143] Figure 113 This is a conceptual diagram of an example display screen used to show a webpage.
[0144] Figure 114 This is a conceptual diagram of an example display screen used to show a webpage.
[0145] Figure 115 This is a block diagram showing an example of a smartphone.
[0146] Figure 116 This is a block diagram illustrating an example of the functional configuration of a smartphone. Detailed Implementation
[0147] In the accompanying drawings, unless the context otherwise indicates, the same reference numerals denote the same elements. The size and relative position of the elements in the drawings are not necessarily drawn to scale.
[0148] In the following description, several embodiments will be illustrated with reference to the accompanying drawings. Note that each of the embodiments described below illustrates a general or specific example. The numerical values, shapes, materials, components, arrangements and connections of components, steps, relationships and sequences of steps indicated in the following embodiments are merely examples and are not intended to limit the scope of the claims.
[0149] Embodiments of the encoder and decoder will now be described. The embodiments are examples of encoders and decoders to which the processes and / or configurations presented in the description of aspects of this disclosure are applicable. The processes and / or configurations may also be implemented in encoders and decoders different from those according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented:
[0150] (1) Any of the components of the encoder or decoder of the embodiments presented in the description of aspects of this disclosure may be replaced by or combined with another component presented anywhere in the description of aspects of this disclosure.
[0151] (2) In the encoder or decoder according to the embodiment, any changes can be made to the functions or processes performed by one or more components of the encoder or decoder, such as adding, replacing, or removing functions or processes. For example, any function or process can be replaced by or combined with another function or process presented anywhere in the description of this aspect of the disclosure.
[0152] (3) In the method implemented by the encoder or decoder according to the embodiment, any changes may be made, such as the addition, substitution, and removal of one or more processes included in the method. For example, any process in the method may be replaced by or combined with another process presented anywhere in the description of this aspect of the disclosure.
[0153] (4) One or more components included in the encoder or decoder according to the embodiment may be combined with components presented anywhere in the description of this disclosure, may be combined with components including one or more functions presented anywhere in the description of this disclosure, and may be combined with components that implement one or more processes implemented by components presented in the description of this disclosure.
[0154] (5) A component that includes one or more functions of an encoder or decoder according to an embodiment, or a component that implements one or more processes of an encoder or decoder according to an embodiment, may be combined with or replaced by: a component presented anywhere in the description of this disclosure, a component that includes one or more functions presented anywhere in the description of this disclosure, or a component that implements one or more processes presented anywhere in the description of this disclosure.
[0155] (6) In a method implemented by an encoder or decoder according to an embodiment, any process included in the method may be replaced by or combined with a process presented anywhere in the description of aspects of this disclosure or by any corresponding or equivalent process.
[0156] (7) One or more processes included in the method implemented by the encoder or decoder according to the embodiment may be combined with processes presented anywhere in the description of aspects of this disclosure.
[0157] (8) The implementation of the processes and / or configurations presented in the description of aspects of this disclosure is not limited to the encoder or decoder according to the embodiments. For example, the processes and / or configurations may be implemented in a device for a different purpose than the motion picture encoder or motion picture decoder disclosed in the embodiments.
[0158] (Definition of the term)
[0159] The corresponding terms can be defined as examples as indicated below.
[0160] An image is a data unit configured with a set of pixels; it is a picture, or a block of pixels. In addition to video, images also include still images.
[0161] An image is an image processing unit configured with a set of pixels, and can also be referred to as a frame or field. For example, an image can be in the form of a luminance sample array in monochrome format or in the form of a luminance sample array and two corresponding chrominance sample arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0162] A block is a processing unit that is a collection of a defined number of pixels. Blocks can have any number of different shapes. For example, a block can be a rectangle with M×N pixels (M columns × N rows), a square with M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, pieces, bricks, CTUs, superblocks, basic splitting units, VPDUs, processing splitting units for hardware, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and subblocks. Blocks can take the form of an M×N array of samples or an M×N array of transform coefficients. For example, a block can be a square or rectangular pixel region comprising a luminance matrix and two chrominance matrices.
[0163] A pixel or sample is the smallest point in an image. A pixel or sample includes pixels at integer positions and pixels at sub-pixel positions, such as pixels generated based on pixels at integer positions.
[0164] Pixel values or sample values are the feature values of a pixel. Pixel values or sample values can include one or more of the following: luminance value, chrominance value, RGB gradient level, depth value, binary value of zero or 1, etc.
[0165] Chromaticity, or color intensity, is the intensity of a color, typically represented by the symbols Cb and Cr. It specifies the value of an array of samples or a single sample value representing the value of one of two color difference signals associated with the primary color.
[0166] Brightness or illuminance is the luminance of an image, typically represented by the symbol or subscript Y or L, which specifies the value of a sample array or a single sample value representing the value of a monochromatic signal associated with the primary color.
[0167] Flags include one or more bits that indicate the value of, for example, a parameter or index. Flags can be binary flags, indicating the binary value of the flag, or they can indicate the non-binary value of a parameter.
[0168] Signals convey information; they are symbolized or encoded into signals. Signals include discrete digital signals and continuous analog signals.
[0169] A stream or bit stream is a digital data string of a digital data stream. A stream or bit stream can be a single stream, or it can be configured with multiple streams having multiple levels. A stream or bit stream can be sent in serial communication using a single transmission path, or it can be sent in packet communication using multiple transmission paths.
[0170] The term "difference" refers to various mathematical differences, such as the simple difference (xy), the absolute value of the difference (|xy|), the difference of squares (x^2-y^2), the square root of the difference (√(xy)), the weighted difference (ax-by: a and b are constants), the offset difference (x-y+a: a is the offset), etc. In the case of scalars, the simple difference may be sufficient and includes difference calculations.
[0171] The term "sum" refers to various mathematical sums, such as the simple sum (x+y), the absolute value of the sum (|x+y|), the sum of squares (x^2+y^2), the square root of the sum (√(x+y)), the weighted difference (ax+by: a and b are constants), the offset sum (x+y+a: a is the offset), etc. In the case of scalars, the simple sum may be sufficient, and summations are also included.
[0172] A frame is a combination of a top field and a bottom field, where sample rows 0, 2, 4, ... originate from the top field, and sample rows 1, 3, 5, ... originate from the bottom field.
[0173] A slice is an integer number of decode tree units contained in an independent slice fragment and all subsequent dependent slice fragments (if any) preceding the next independent slice fragment (if any) within the same access unit.
[0174] A slice is a rectangular region of the decoder tree within a specific slice column and row in an image. A slice can be a rectangular region of a frame designed to be independently decoded and encoded, but loop filtering across slice edges can still be applied.
[0175] A Code Tree Unit (CTU) can be a block of decoder tree representing luminance samples from an image with three sample arrays, or two corresponding blocks of decoder tree representing chrominance samples. Alternatively, a CTU can be a block of decoder tree representing samples from a monochrome image and an image decoded using three separate color planes and a syntax structure for decoding the samples. A superblock can be a 64×64 pixel square block consisting of one or two pattern information blocks, or recursively divided into four 32×32 blocks, which themselves can be further subdivided.
[0176] (System Configuration)
[0177] First, the transmission system according to an embodiment will be described. Figure 1 This is a schematic diagram illustrating an example configuration of a transmission system 400 according to an embodiment.
[0178] The transmission system 400 is a system for transmitting a stream generated by encoding an image and for decoding the transmitted stream. As shown, the transmission system 400 includes, for example, […]. Figure 1 The encoder 100, network 300, and decoder 200 are shown in the figure.
[0179] An image is input to encoder 100. Encoder 100 generates a stream by encoding the input image and outputs the stream to network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed through encoding.
[0180] It should be noted that the image before encoding by encoder 100 is also referred to as the raw image, raw signal, or raw sample. An image can be video or a still image. Image is a general concept encompassing sequence, picture, and block; therefore, unless otherwise stated, an image is not limited to a spatial region of a specific size and a temporal region of a specific size. An image is an array of pixels or pixel values, and the signal representing the image or pixel values is also referred to as a sample. A stream can be referred to as a bit stream, an encoded bit stream, a compressed bit stream, or an encoded signal. Furthermore, encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.
[0181] Network 300 sends the stream generated by encoder 100 to decoder 200. Network 200 can be any combination of the Internet, wide area network (WAN), local area network (LAN), or network. Network 300 is not limited to a two-way communication network and can also be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasts and satellite broadcasts. Alternatively, network 300 can be replaced by a recording medium on which the stream is recorded (e.g., digital versatile disc (DVD) and Blu-ray disc (BD)).
[0182] Decoder 200 generates a decoded image as an uncompressed image, for example, by decoding the stream sent by network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method used by encoder 100.
[0183] It should be noted that decoder 200 can also be referred to as image decoder or video decoder, and the decoding method performed by decoder 200 can also be referred to as decoding method, image decoding method or video decoding method.
[0184] (Data Structures)
[0185] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream. For convenience, reference will be made to... Figure 1 The transmission system 400 is used to describe Figure 2 Streams include, for example, video sequences. (e.g.) Figure 2As shown in (a), the video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplementary enhancement information (SEI), and multiple pictures.
[0186] In a video with multiple layers, a VPS may include decoding parameters that are common between some of the layers, as well as decoding parameters that are associated with some of the layers included in the video or with a single layer.
[0187] The SPS includes parameters for the sequence, that is, decoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the decoding parameters may indicate the width or height of the image. It should be noted that multiple SPSs may exist.
[0188] PPS includes parameters for the images, that is, decoding parameters that the decoder 200 refers to in order to decode each of the images in the sequence. For example, the decoding parameters may include reference values for the quantization width used to decode the images and flags indicating the application of weighted predictions. It should be noted that multiple PPSs may exist. Each of the SPS and PPS can be simply referred to as a parameter set.
[0189] like Figure 2 As shown in (b), the image may include an image header and one or more slices. The image header includes decoding parameters, which the decoder 200 refers to to decode the one or more slices.
[0190] like Figure 2 As shown in (c), the slice includes a slice header and one or more bricks. The slice header includes decoding parameters that the decoder 200 refers to in order to decode the one or more bricks.
[0191] like Figure 2 As shown in (d), the bricks comprise one or more decoding tree units (CTUs).
[0192] It should be noted that an image may not include any slices and may include groups of slices instead of slices. In this case, a group of slices includes at least one slice. Alternatively, bricks may include slices.
[0193] CTU is also known as a superblock or basic split unit. For example... Figure 2 As shown in (e), the CTU includes a CTU header and at least one decoding unit (CU). As shown, the CTU includes four decoding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes decoding parameters that the decoder 200 refers to in order to decode at least one CU.
[0194] A CU can be divided into multiple smaller CUs. As shown, CU(10) is not divided into smaller decoding units; CU(11) is divided into four smaller decoding units CU(110), CU(111), CU(112), and CU(113); CU(12) is not divided into smaller decoding units; and CU(13) is divided into seven smaller decoding units CU(1310), CU(1311), CU(1312), CU(1313), CU(132), CU(133), and CU(134). Figure 2 As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is used to predict the CU, and the residual coefficient information is information indicating the prediction residuals, as described later. While the CU is substantially the same as the prediction unit (PU) and transform unit (TU), it should be noted that, for example, the subblock transform (SBT) described later may include multiple TUs smaller than the CU. Additionally, the CU can be processed for each virtual pipelined decoding unit (VPDU) included in the CU. A VPDU is, for example, a fixed unit that can be processed at a stage when pipelined processing is performed in hardware.
[0195] It should be noted that a flow may not include... Figure 2 All the hierarchical layers are shown in the diagram. The order of the hierarchical layers can be interchanged, or any one of the hierarchical layers can be replaced by another hierarchical layer. Here, the image that is the target of a process to be performed by the device (e.g., encoder 100 or decoder 200) is called the current image. The current image represents the current image to be encoded when the process is an encoding process, and the current image represents the current image to be decoded when the process is a decoding process. Similarly, for example, a CU or CU block that is the target of a process to be performed by the device (e.g., encoder 100 or decoder 200) is called the current block. The current block represents the current block to be encoded when the process is an encoding process, and the current block represents the current block to be decoded when the process is a decoding process.
[0196] (Image structure: slice / partition)
[0197] An image can be configured with one or more slice units or one or more fragment units to facilitate parallel decoding / decoding of the image.
[0198] A slice is a basic decoding unit included in an image. An image may include, for example, one or more slices. Additionally, a slice may include one or more decoding tree units (CTUs).
[0199] Figure 3 This is a conceptual diagram used to illustrate an example of slice configuration. For example, in Figure 3The image contains 11×8 CTUs and is divided into four slices (slices 1 to 4). Slice 1 contains sixteen CTUs, slice 2 contains twenty-one CTUs, slice 3 contains twenty-nine CTUs, and slice 4 contains twenty-two CTUs. Each CTU in the image belongs to one slice. The shape of each slice is obtained by horizontally splitting the image. The boundaries of each slice do not need to coincide with image endpoints and can coincide with any of the boundaries between CTUs in the image. The processing order (encoding or decoding order) of the CTUs in a slice is, for example, the raster scan order. A slice includes a slice header and encoded data. Slice features can be written into the slice header. These features can include the CTU address of the top CTU in the slice, the slice type, etc.
[0200] A tile is a unit rectangular area included in an image. Image tiles can be assigned a number called TileId according to the raster scan order.
[0201] Figure 4 This is a conceptual diagram used to illustrate an example of a sharding configuration. For example, in Figure 4 The image contains 11×8 CTUs and is divided into four rectangular regions (partitions 1 to 4). When using partitioning, the processing order of the CTUs may differ from that without partitioning. Without partitioning, multiple CTUs in the image are typically processed in raster scan order. When using multiple partitions, at least one CTU in each of the multiple partitions is processed in raster scan order. For example, as... Figure 4 As shown, the processing order of the CTUs included in shard 1 is from the left end of the first column of shard 1 toward the right end of the first column of shard 1, and then continues from the left end of the second column of shard 1 toward the right end of the second column of shard 1.
[0202] It should be noted that a slice may include one or more slices, and a slice may include one or more slices.
[0203] It should be noted that an image can be configured with one or more fragment sets. A fragment set can include one or more fragment groups, or one or more fragments. An image can be configured with one of fragment sets, fragment groups, and fragments. For example, suppose the order in which multiple fragments are scanned for each fragment set in raster scan order is the basic encoding order of the fragments. Suppose the set of one or more fragments consecutive in basic encoding order in each fragment set is a fragment group. Such an image can be split by the splitter 102 described later (see [link to splitter description]). Figure 7 Configure it using ).
[0204] (Scalable encoding)
[0205] Figure 5 and Figure 6 This is a conceptual diagram illustrating an example of a scalable flow structure, and for convenience, references will be made to... Figure 1 Describe it.
[0206] like Figure 5 As shown, encoder 100 can generate a temporally / spatially scalable stream by dividing each of a plurality of images into any of a plurality of layers and encoding the images within the layers. For example, encoder 100 encodes the images of each layer, thereby achieving scalability even when enhancement layers exist above base layers. This encoding of each image is also called scalable encoding. In this way, decoder 200 is able to switch the image quality of the images displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as the processing power of decoder 200 and external factors such as the state of communication bandwidth. As a result, decoder 200 is able to decode content while freely switching between low and high resolution. For example, a user of the stream watches half of the streaming video on a smartphone on their way home and continues watching the video at home on a device such as a TV connected to the internet. It should be noted that each of the smartphones and devices described above includes decoder 200 with the same or different performance. In this case, when the device decodes layers up to higher layers in the stream, the user can watch high-quality video at home. In this way, encoder 100 does not need to generate multiple streams with different image qualities having the same content, and thus the processing load can be reduced.
[0207] Furthermore, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 can generate a video with enhanced image quality by performing super-resolution imaging on the images in the base layer based on the metadata. Super-resolution imaging may include, for example, improvements in the SN ratio at the same resolution, increases in resolution, etc. Metadata may include, for example, information for identifying linear or nonlinear filter coefficients (as used in the super-resolution process), or information for identifying parameter values in filtering processes, machine learning, or least squares methods used in super-resolution processing.
[0208] In an embodiment, a configuration can be provided in which an image is divided into segments, for example, based on the meaning of objects in the image. In this case, the decoder 200 can decode only a portion of the image by selecting the segments to be decoded. Alternatively, the attributes of the objects (people, cars, balls, etc.) and the objects' positions in the image (coordinates within the same image) can be stored as metadata. In this case, the decoder 200 can identify the location of the desired object based on the metadata and determine the segment that includes that object. For example, as... Figure 6As shown, metadata can be stored using a data storage structure different from that used for image data (e.g., SEI (Supplemental Enhancement Information) messages in HEVC). This metadata indicates, for example, the location, size, or color of the primary object.
[0209] Metadata can be stored in units of multiple images (e.g., streams, sequences, random access units). In this way, decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information to the image unit information, it can identify images containing objects (people) and determine the object's location in the image.
[0210] (encoder)
[0211] An encoder according to an embodiment will be described. Figure 7 This is a block diagram illustrating the functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes video in blocks.
[0212] like Figure 7 As shown, encoder 100 is a device for encoding images in blocks, including splitter 102, subtractor 104, transformer 106, quantizer 108, entropy encoder 110, inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, loop filter 120, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and prediction parameter generator 130. As shown, intra-frame predictor 124 and inter-frame predictor 126 are part of the prediction controller.
[0213] The encoder 100 is implemented, for example, as a general-purpose processor and memory. In this case, when the software program stored in the memory is executed by the processor, the processor acts as a splitter 102, a subtractor 104, a converter 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse converter 114, an adder 116, a loop filter 120, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, subtractor 104, converter 106, quantizer 108, entropy encoder 110, inverse quantizer 112, inverse converter 114, adder 116, loop filter 120, intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128.
[0214] (Encoder installation example)
[0215] Figure 8 This is a functional block diagram illustrating an example installation of encoder 100. Encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The encoder 100 shown in the figure has multiple constituent elements mounted on it. Figure 8 The processor a1 and memory a2 are shown in the figure.
[0216] Processor a1 is a circuit that performs information processing and is coupled to memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes images. Processor a1 can be a processor such as a CPU. Alternatively, processor a1 can be an assembly of multiple electronic circuits. Additionally, for example, processor a1 can perform... Figure 7 The roles of two or more of the constituent elements of the encoder 100 shown in the figure.
[0217] Memory a2 is a dedicated or general-purpose memory used by processor a1 to encode images. Memory a2 can be an electronic circuit and can be connected to processor a1. Alternatively, memory a2 can be included within processor a1. Alternatively, memory a2 can be an assembly of multiple electronic circuits. Alternatively, memory a2 can be a disk, optical disk, etc., or can be represented as a storage device, recording medium, etc. Additionally, memory a2 can be non-volatile memory or volatile memory.
[0218] For example, memory a2 can store the image to be encoded or a bitstream corresponding to the encoded image. Additionally, memory a2 can store a program for instructing processor a1 to encode the image.
[0219] Alternatively, for example, memory a2 can undertake Figure 7 The encoder 100 shown herein has multiple constituent elements, including the roles of two or more constituent elements used for storing information. For example, memory a2 can serve as... Figure 7 The roles of block memory 118 and frame memory 122 are shown in the diagram. More specifically, memory a2 can store reconstructed blocks, reconstructed images, etc.
[0220] It should be noted that encoder 100 may not implement... Figure 7 All of the constituent elements indicated herein may not perform all the processes described herein. Figure 7 A portion of the constituent elements indicated herein may be included in another device, or a portion of the process described herein may be performed by another device.
[0221] The following describes the overall flow of the process performed by encoder 100, and then describes each of the constituent elements included in encoder 100.
[0222] (The overall flow of the coding process)
[0223] Figure 9 This is a flowchart illustrating an example of the overall encoding process performed by encoder 100, and for convenience, reference will be made to... Figure 7 Describe it.
[0224] First, the splitter 102 of the encoder 100 splits each of the images included in the input image into multiple blocks of a fixed size (e.g., 128 × 128 pixels) (step Sa_1). The splitter 102 then selects a splitting pattern for the fixed-size blocks (also called block shapes) (step Sa_2). In other words, the splitter 102 further splits the fixed-size blocks into multiple blocks forming the selected splitting pattern. For each of the multiple blocks, the encoder 100 performs steps Sa_3 to Sa_9 for the block (i.e., the current block to be encoded).
[0225] The prediction controller 128 and the prediction actuators (including the intra-frame predictor 124 and the inter-frame predictor 126) generate a prediction image of the current block (step Sa-3). The prediction image may also be referred to as the prediction signal, the prediction block, or the prediction sample.
[0226] Next, subtractor 104 generates the difference between the current block and the predicted image as the prediction residual (step Sa_4). The prediction residual can also be referred to as the prediction error.
[0227] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate multiple quantized coefficients (step Sa_5). The multiple quantized coefficients can sometimes be referred to as a coefficient block.
[0228] Next, the entropy encoder 110 encodes (specifically, performs entropy coding) multiple quantized coefficients and prediction parameters related to the generation of the predicted image to generate a stream (step Sa_6). The stream can sometimes be referred to as an encoded bitstream or a compressed bitstream.
[0229] Next, the inverse quantizer 112 performs inverse quantization on the multiple quantized coefficients, and the inverse transformer 114 performs inverse transformation on the result to recover the prediction residual (step Sa_7).
[0230] Next, adder 116 adds the predicted image to the recovered prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image can also be referred to as a reconstructed block or a decoded image block.
[0231] When the reconstructed image is generated, the loop filter 120 performs filtering on the reconstructed image as needed (step Sa_9).
[0232] Then, encoder 100 determines whether the encoding of the entire image has ended (step Sa_10). If it is determined that the encoding has not ended ("No" in step Sa_10), the process starting from step Sa_2 is repeated for the next block of the image.
[0233] Although in the example described above, encoder 100 selects a split pattern for fixed-size blocks and encodes each block according to that split pattern, it should be noted that each block can be encoded according to a corresponding split pattern from a plurality of split patterns. In this case, encoder 100 can evaluate the cost of each of the plurality of split patterns and, for example, select the stream obtained by encoding according to the split pattern that produces the minimum cost as the output stream.
[0234] As shown, the processes in steps Sa_1 to Sa_10 are executed sequentially by encoder 100. Alternatively, two or more processes can be executed in parallel, and the processes can be reordered, etc.
[0235] The encoding process employed by encoder 100 is a hybrid encoding using predictive coding and transform coding. Furthermore, predictive coding is performed by a coding loop configured with subtractor 104, transformer 106, quantizer 108, inverse quantizer 112, inverse transformer 114, adder 116, loop filter 120, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128. In other words, the prediction executor configured with intra-frame predictor 124 and inter-frame predictor 126 is part of the coding loop.
[0236] (Splitter)
[0237] The splitter 102 splits each image included in the original image into multiple blocks and outputs each block to the subtractor 104. For example, the splitter 102 first splits the image into fixed-size blocks (e.g., 128 × 128 pixels). Other fixed block sizes may be used. Fixed-size blocks are also referred to as decoding tree units (CTUs). Then, the splitter 102 splits each fixed-size block into variable-size blocks (e.g., 64 × 64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a splitting mode. Variable-size blocks may also be referred to as decoding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or some blocks in the image can be processed in units of CUs, PUs, or TUs.
[0238] Figure 10 This is a conceptual diagram used to illustrate an example of block splitting according to an embodiment. Figure 10 In the diagram, solid lines represent the block boundaries of blocks split using quadtree block splitting, while dashed lines represent the block boundaries of blocks split using binary tree block splitting.
[0239] Here, block 10 is a square block with 128×128 pixels (128×128 block). This 128×128 block 10 is first divided into four square blocks of 64×64 pixels (quadtree block splitting).
[0240] The 64×64 pixel block in the upper left corner is further vertically divided into two rectangular 32×64 pixel blocks, and the 32×64 pixel block on the left is further vertically divided into two rectangular 16×64 pixel blocks (binary block splitting). As a result, the 64×64 pixel block in the upper left corner is split into two 16×64 pixel blocks 11 and 12 and a 32×64 pixel block 13.
[0241] The 64×64 pixel block in the upper right corner is horizontally split into two rectangular 64×32 pixel blocks, 14 and 15 (binary tree block split).
[0242] The 64×64 pixel square in the lower left corner is first divided into four 32×32 pixel squares (quadtree block splitting). The upper left and lower right blocks within these four 32×32 pixel squares are further split. The upper left 32×32 pixel square is vertically split into two 16×32 pixel rectangular blocks, and the rightmost 16×32 pixel block is further horizontally split into two 16×16 pixel blocks (binary tree block splitting). The lower right 32×32 pixel block is horizontally split into two 32×16 pixel blocks (binary tree block splitting). The upper right 32×32 pixel square is horizontally split into two 32×16 pixel rectangular blocks (binary tree block splitting). As a result, the 64×64 pixel block of the square in the lower left corner was split into a 16×32 pixel block 16, two 16×16 pixel blocks 17 and 18 of squares, two 32×32 pixel blocks 19 and 20 of squares, and two 32×16 pixel blocks 21 and 22 of rectangles.
[0243] The 64×64 pixel block 23 in the lower right corner was not split.
[0244] As described above, in Figure 10 In this example, based on recursive quadtree and binary tree block splitting, block 10 is split into thirteen variable-size blocks 11 to 23. This type of splitting is also known as quadtree plus binary tree (QTBT) splitting.
[0245] It should be noted that, Figure 10In this context, a block is split into four or two blocks (quadtree or binary tree block split), but the split is not limited to these examples. For instance, a block can be split into three blocks (ternary block split). Splits that include this type of ternary block split are also known as multi-type tree (MBT) splits.
[0246] Figure 11 This is a block diagram illustrating an example of the functional configuration of a splitter 102 according to one embodiment. Figure 11 As shown, splitter 102 may include block split determiner 102a. As an example, block split determiner 102a may perform the following process.
[0247] For example, block splitting determiner 102a can obtain or retrieve block information from block memory 118 and / or frame memory 122, and determine a splitting mode (e.g., the splitting mode described above) based on the block information. Splitter 102 splits the original image according to the splitting mode and outputs at least one block obtained by splitting to subtractor 104.
[0248] Alternatively, for example, the block splitting determiner 102a outputs one or more parameters indicating the determined splitting pattern (e.g., the splitting pattern described above) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 can transform the prediction residual based on one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 can generate the prediction image based on one or more parameters. Additionally, the entropy encoder 110 can entropy encode one or more parameters.
[0249] As indicated below, parameters related to the splitting mode can be written into a stream, as an example.
[0250] Figure 12 This is a conceptual diagram used to illustrate examples of splitting patterns. Examples of splitting patterns include: splitting into four regions (QT), where the block is split horizontally into two regions and vertically into two regions; splitting into three regions (HT or VT), where the block is split in the same direction at a ratio of 1:2:1; splitting into two regions (HB or VB), where the block is split in the same direction at a ratio of 1:1; and no splitting (NS).
[0251] It should be noted that the splitting mode does not have a block splitting direction when splitting into four regions or not splitting, while the splitting mode has splitting direction information when splitting into two or three regions.
[0252] Figure 13A This is a conceptual diagram used to illustrate an example of a syntactic tree for splitting patterns.
[0253] Figure 13B This is a conceptual diagram used to illustrate another example of a syntactic tree for splitting patterns.
[0254] Figure 13A and Figure 13B This is a conceptual diagram used to illustrate an example of a syntactic tree for splitting patterns. Figure 13A In the example, first, there is information indicating whether to perform a split (S: split flag), and next, there is information indicating whether to split into four regions (QT: QT flag). Next, there is information indicating which of the three or two regions should be split into (TT: TT flag, or BT: BT flag), and then, there is information indicating the splitting direction (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of the at least one block obtained by splitting according to such a splitting pattern can be further split repeatedly in a similar process. In other words, as an example, it is possible to recursively determine whether to perform a split, whether to split into four regions, which of the horizontal and vertical directions is the direction of the splitting method to be performed, and which of the three or two regions should be split into, and can be determined according to... Figure 13A The encoding order publicly displayed in the syntax tree shown will determine the encoding of the results in the stream.
[0255] Additionally, although the information items indicating S, QT, TT, and Ver are respectively in Figure 13A The syntax tree shown is arranged in the listed order, but the information items indicating S, QT, Ver, and BT can be arranged in the listed order respectively. In other words, in Figure 13B In the example, first, there is information indicating whether to perform a split (S: split flag), and next, there is information indicating whether to perform a split into four regions (QT: QT flag). Next, there is information indicating the split direction (Ver: vertical flag, or Hor: horizontal flag), and then, there is information indicating which of the two or three regions should be split into (BT: BT flag, or TT: TT flag).
[0256] It should be noted that the splitting patterns described above are examples, and splitting patterns other than those described can be used, or a portion of the described splitting patterns can be used.
[0257] (Subtractor)
[0258] Subtractor 104 subtracts the prediction image (prediction samples input from prediction controller 128, as indicated below) from the original image, which is input from and split by splitter 102 in blocks. In other words, subtractor 104 calculates the prediction residual (also known as error) for the current block. Subtractor 104 then outputs the calculated prediction residual to transformer 106.
[0259] The original image can be an image that has been input into encoder 100 as a signal (e.g., luminance signal and two chrominance signals) representing each picture included in the video. The signal representing the image can also be referred to as a sample.
[0260] (Transformer)
[0261] Transformer 106 transforms the prediction residuals in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to quantizer 108. More specifically, transformer 106 applies, for example, a defined Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) to the prediction residuals in the spatial domain. The defined DCT or DST can be predefined.
[0262] It should be noted that the transformer 106 can adaptively select a transformation type from multiple transformation types and transform the prediction residuals into transformation coefficients by using transformation basis functions corresponding to the selected transformation type. This type of transformation is also known as explicit multi-kernel transformation (EMT) or adaptive multiple transformation (AMT). The transformation basis functions can also be referred to as bases.
[0263] Transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. It should be noted that these transformation types can also be represented as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 This is a graph indicating the example transformation basis functions used for the example transformation type. Figure 14 In this context, N indicates the number of input pixels. For example, the selection of a transform type from multiple transform types can depend on the prediction type (one of intra-frame prediction and inter-frame prediction) and can also depend on the intra-frame prediction mode.
[0264] Information indicating whether to apply EMT or AMT (referred to as, for example, EMT flag or AMT flag) and information indicating the selected transformation type are typically signaled at the CU level. It should be noted that signaling this information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, or CTU level).
[0265] Additionally, transformer 106 can re-transform the transform coefficients (which are the transform results). This re-transformation is also known as adaptive quadratic transform (AST) or non-separable quadratic transform (NSST). For example, transformer 106 performs the re-transformation on a sub-block (e.g., a 4×4 pixel sub-block) basis, which is included in the transform coefficient block corresponding to the intra-frame prediction residual. Information indicating whether NSST is applied and information related to the transform matrix used for NSST are typically signaled at the CU level. It should be noted that signaling this information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, or CTU level).
[0266] Transformer 106 can employ both separable and non-separable transformations. A separable transformation is a method in which a transformation is performed multiple times by individually performing the transformation for each of multiple directions according to the dimension of the input. A non-separable transformation is a method of performing a collective transformation in which two or more dimensions of the multidimensional input are collectively treated as a single dimension.
[0267] In one example of an inseparable transformation, when the input is a 4×4 pixel block, the 4×4 pixel block is considered as a single array consisting of sixteen elements, and the transformation applies a 16×16 transformation matrix to the array.
[0268] In another example of an inseparable transformation, a 4×4 pixel input block is treated as a single array comprising sixteen elements, and a transformation (hypercube given transformation) can then be performed on the array with multiple given rotations.
[0269] In the transformation within transformer 106, the type of transformation to be applied to the transform basis functions in the frequency domain based on the region transformation in the CU can be switched. Examples include spatial transformation (SVT).
[0270] Figure 15 This is a conceptual diagram used to illustrate an example of SVT.
[0271] In SVT, such as Figure 15 As shown, the CU is split horizontally or vertically into two equal regions, and only one of the regions is transformed into the frequency domain. The basic transformation type can be set for each region. For example, DST7 and DST8 can be used. For instance, in the two regions obtained by vertically splitting the CU into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, DST7 can be used for the region at position 1 in both regions. Similarly, in the two regions obtained by horizontally splitting the CU into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, DST7 can be used for the region at position 1 in both regions. Although in Figure 15 In the example shown, only one region in the two regions of the CU is transformed while the other is not, but each region in both regions can be transformed. Furthermore, the splitting method can include not only splitting into two regions but also splitting into four regions. Additionally, the splitting method can be more flexible. For example, information indicating the splitting method can be encoded and signaled in the same way as CU splitting. It should be noted that SVT can also be called Subblock Transform (SBT).
[0272] The AMT and EMT described above can be referred to as MTS (Multiple Transform Selection). When applying MTS, transform types such as DST7 and DCT8 can be selected, and information indicating the selected transform type can be encoded as index information for each CU. There is another process called IMTS (Implicit MTS) for selecting the transform type to be used for orthogonal transforms performed without encoded index information. When applying IMTS, for example, when the CU has a rectangular shape, DST7 can be used for the short side and DST2 for the long side to perform an orthogonal transform for the rectangular shape. Alternatively, for example, when the CU has a square shape, DCT2 can be used when MTS is valid in the sequence and DST7 can be used when MTS is invalid in the sequence to perform an orthogonal transform for the rectangular shape. DCT2 and DST7 are just examples. Other transform types can be used, and the combination of transform types can be changed for different combinations of transform types. IMTS can be used only for intra-prediction blocks, or it can be used for both intra-prediction blocks and inter-prediction blocks.
[0273] The three processes MTS, SBT, and IMTS have been described above as selection processes for selectively switching transform types for orthogonal transforms. However, all three selection processes can be used, or only a portion of the selection processes can be selectively used. For example, one or more selection processes can be identified based on flag information in a header such as SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching transform types can be a different selection process from the three selection processes described above, or each of the three selection processes can be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing the orthogonal transform of the entire CU and determining the transform type based on determined rules without encoding information indicating the transform type. Function [3] is a function for performing the orthogonal transform of a portion of the CU and encoding information indicating the transform type used in the transform. The function [4] is used to perform orthogonal transformations on a portion of the CU and to determine the transformation type based on established rules without encoding information indicating the type of transformation used in the transformation. The established rules can be predetermined.
[0274] It should be noted that the application of MTS, IMTS, and / or SBT can be determined for each processing unit. For example, the application of MTS, IMTS, and / or SBT can be determined for each sequence, image, brick, slice, CTU, or CU.
[0275] It should be noted that the tool for selectively switching transformation types in this disclosure can be described as a method, selection process, or procedure for selectively selecting the basis used in the transformation process. Additionally, the tool for selectively switching transformation types can be described as a mode for adaptively selecting the transformation type.
[0276] Figure 16 This is a flowchart illustrating an example of the process performed by converter 106, and references will be made for convenience. Figure 7 Describe it.
[0277] For example, transformer 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined that an orthogonal transformation should be performed ("Yes" in step St_1), transformer 106 selects a transformation type for orthogonal transformation from a plurality of transformation types (step St_2). Next, transformer 106 performs orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Transformer 106 then outputs information indicating the selected transformation type to entropy encoder 110 so that entropy encoder 110 can encode the information (step St_4). On the other hand, when it is determined that an orthogonal transformation should not be performed ("No" in step St_1), transformer 106 outputs information indicating that an orthogonal transformation should not be performed so that entropy encoder 110 can encode the information (step St_5). It should be noted that whether to perform an orthogonal transformation in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transformation can be performed using a defined transformation type without encoding the information indicating the transformation type used in the orthogonal transformation. The defined transformation type can be predefined.
[0278] Figure 17 This is a flowchart illustrating an example of the process performed by converter 106, and references will be made for convenience. Figure 7 Please describe it. It should be noted that... Figure 17 The example shown is in the case where the type of transformation used in orthogonal transformations is selectively switched (such as in...). Figure 16 Examples of orthogonal transformations (as shown in the example).
[0279] As an example, the first transform type group may include DCT2, DST7, and DCT8. As another example, the second transform type group may include DCT2. Transform types included in the first transform type group and transform types included in the second transform type group may partially overlap or may be completely different from each other.
[0280] Transformer 106 determines whether the transform size is less than or equal to a predetermined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the predetermined value ("Yes" in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using a transform type included in the first transform type group (step Su_2). Next, transformer 106 outputs information indicating the transform type to be used from at least one transform type included in the first transform type group to the entropy encoder 110 so that the entropy encoder 110 can encode the information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to a predetermined value ("No" in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using a second transform type group (step Su_4). The predetermined value can be a threshold and can be a predetermined value.
[0281] In step Su_3, the information indicating the transformation type used in the orthogonal transformation can be a combination of information indicating the transformation type to be applied vertically in the current block and the transformation type to be applied horizontally in the current block. The first type group may include only one transformation type and may not encode the information indicating the transformation type used in the orthogonal transformation. The second transformation type group may include multiple transformation types and may encode the information indicating the transformation type used in the orthogonal transformation among one or more transformation types included in the second transformation type group.
[0282] Alternatively, the transformation type can be indicated based on the transformation size, without encoding the information indicating the transformation type. It should be noted that such determination is not limited to determining whether the transformation size is less than or equal to a given value, and other procedures can also be used to determine the transformation type used in orthogonal transformations based on the transformation size.
[0283] (Quantizer)
[0284] Quantizer 108 quantizes the transform coefficients output from transformer 106. More specifically, quantizer 108 scans the transform coefficients of the current block in a determined scan order and quantizes the scanned transform coefficients based on quantization parameters (QP) corresponding to the transform coefficients. Quantizer 108 then outputs the quantized transform coefficients of the current block (also referred to quantized coefficients hereinafter) to entropy encoder 110 and inverse quantizer 112. The determined scan order can be predetermined.
[0285] The determined scan order is the order in which the transform coefficients are quantized / inverse quantized. For example, the determined scan order can be defined as ascending frequency (from low frequency to high frequency) or descending frequency (from high frequency to low frequency).
[0286] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, when the value of the quantization parameter increases, the quantization step size also increases. In other words, when the value of the quantization parameter increases, the error of the quantized coefficients (quantization error) increases.
[0287] Additionally, a quantization matrix can be used for quantization. For example, several types of quantization matrices can be used corresponding to frequency transform sizes such as 4×4 or 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luma and chroma pixel components. It should be noted that quantization refers to digitizing values sampled at determined intervals corresponding to a determined level. In this art, quantization can be referred to using other terms (e.g., rounding and scaling), and rounding and scaling can be employed. The determined intervals and determined levels can be predetermined.
[0288] Methods for using a quantization matrix can include using a quantization matrix that has already been set directly on the encoder 100 side, and using a quantization matrix that has been set as the default (default matrix). On the encoder 100 side, a quantization matrix suitable for the image features can be set by directly setting the quantization matrix. However, this may have the disadvantage of increasing the amount of decoding required to encode the quantization matrix. It should be noted that a quantization matrix for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.
[0289] There exists a method for quantizing high-frequency and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be considered equivalent to using a quantization matrix (a flat matrix) whose coefficients have the same values.
[0290] The quantization matrix can be encoded at, for example, the sequence level, image level, slice level, brick level, or CTU level. The quantization matrix can be specified using, for example, a Sequence Parameter Set (SPS) or a Picture Parameter Set (PPS). The SPS includes parameters for the sequence, and the PPS includes parameters for the image. Each of the SPS and PPS can be simply referred to as a parameter set.
[0291] When using a quantization matrix, quantizer 108 uses the values of the quantization matrix to scale the quantization width for each transform coefficient, for example, based on quantization parameters. A quantization process performed without using a quantization matrix can be a process for quantizing the transform coefficients according to a quantization width calculated based on quantization parameters. It should be noted that in a quantization process performed without using any quantization matrix, the quantization width can be multiplied by a predetermined value common to all transform coefficients in the block. This predetermined value can be pre-determined.
[0292] Figure 18 This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment. For example, quantizer 108 includes a differential quantization parameter generator 108a, a predicted quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage device 108d, and a quantization actuator 108e.
[0293] Figure 19 This is a flowchart illustrating an example of the quantization process performed by quantizer 108, and for convenience, reference will be made to... Figure 7 and Figure 18 Describe it.
[0294] As an example, the quantizer 108 can be based on Figure 19 The flowchart shown in the figure shows the quantization process performed for each CU. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined that quantization should be performed ("yes" in step Sv_1), the quantization parameter generator 108c generates quantization parameters for the current block (step Sv_2) and stores the quantization parameters in the quantization parameter storage device 108d (step Sv_3).
[0295] Next, the quantization executor 108e uses the quantization parameters generated in step Sv_2 to quantize the transform coefficients of the current block (step Sv_4). The predicted quantization parameter generator 108b then obtains the quantization parameters for a different processing unit than the current block from the quantization parameter storage device 108d (step Sv_5). The predicted quantization parameter generator 108b generates the predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generator 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generator 108c and the predicted quantization parameters of the current block generated by the predicted quantization parameter generator 108b (step Sv_7). The differential quantization parameters can be generated by calculating the difference. The differential quantization parameter generator 108a outputs the differential quantization parameters to the entropy encoder 110 so that the entropy encoder 110 can encode the differential quantization parameters (step Sv_8).
[0296] It should be noted that differential quantization parameters can be encoded at, for example, the sequence level, image level, slice level, brick level, or CTU level. Additionally, the initial values of the quantization parameters can be encoded at the sequence level, image level, slice level, brick level, or CTU level. During initialization, the initial values of the quantization parameters and the differential quantization parameters can be used to generate the quantization parameters.
[0297] It should be noted that quantizer 108 may include multiple quantizers and may apply dependent quantization, wherein the transform coefficients are quantized using a quantization method selected from multiple quantization methods.
[0298] (Entropy encoder)
[0299] Figure 20 This is a block diagram illustrating an example of the functional configuration of the entropy encoder 110 according to an embodiment, and reference will be made for convenience. Figure 7 The entropy encoder 110 generates a stream by entropy encoding quantized coefficients input from quantizer 108 and prediction parameters input from prediction parameter generator 130. For example, context-based adaptive binary arithmetic decoding (CABAC) is used as the entropy encoding. More specifically, the entropy encoder 110 shown includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, where a multi-level signal, such as quantized coefficients and prediction parameters, is transformed into a binary signal. Examples of binarization methods include truncated Rice binarization, exponential Golomb coding, and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntax elements or the surrounding state (i.e., the probability of occurrence of the binary signal). Examples of methods for deriving the context value include bypassing, referencing syntax elements, referencing the upper and left adjacent blocks, referencing hierarchy information, etc. The binary arithmetic encoder 110c uses the derived context to arithmetically encode the binary signal.
[0300] Figure 21 This is a conceptual diagram illustrating an example flow of the CABAC process in the entropy encoder 110. First, initialization is performed in CABAC within the entropy encoder 110. During initialization, initialization and setting of the initial context value are performed in the binary arithmetic encoder 110c. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic encoding of multiple quantization coefficients within the CTU. Each time arithmetic encoding is performed, the context controller 110b can update the context value. The context controller 110b can then save the context value for post-processing. For example, the saved context value can be used to initialize the context value for the next CTU.
[0301] (Inverse quantizer)
[0302] Inverse quantizer 112 inverse quantizes the quantized coefficients input from quantizer 108. More specifically, inverse quantizer 112 inverse quantizes the quantized coefficients of the current block in a determined scan order. Inverse quantizer 112 then outputs the inverse-quantized transform coefficients of the current block to inverse transform 114. The determined scan order can be predetermined.
[0303] (Inverse Transformer)
[0304] Inverse transformer 114 recovers the prediction residual by performing an inverse transform on the transform coefficients that have been input from inverse quantizer 112. More specifically, inverse transformer 114 recovers the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by transformer 106. Inverse transformer 114 then outputs the recovered prediction residual to adder 116.
[0305] It should be noted that, because information is typically lost during quantization, the recovered prediction residual does not match the prediction residual calculated by subtractor 104. In other words, the recovered prediction residual usually includes quantization error.
[0306] (Adder)
[0307] Adder 116 reconstructs the current block by adding the prediction residual that has been input from inverse transformer 114 and the prediction image that has been input from prediction controller 128. Thus, a reconstructed image is generated. Adder 116 then outputs the reconstructed image to block memory 118 and loop filter 120. The reconstructed block can also be referred to as a locally decoded block.
[0308] (Block memory)
[0309] Block memory 118 is a storage device for storing blocks in the current image (e.g., for intra-frame prediction). More specifically, block memory 118 stores the reconstructed image output from adder 116.
[0310] (Frame Memory)
[0311] Frame memory 122 is, for example, a storage device for storing reference images used in inter-frame prediction, and is also referred to as a frame buffer. More specifically, frame memory 122 stores the reconstructed image filtered by loop filter 120.
[0312] (Loop filter)
[0313] Loop filter 120 applies a loop filter to the reconstructed image output by adder 116 and outputs the filtered reconstructed image to frame memory 122. A loop filter is a filter used in the coding loop. Examples of loop filters include, for example, adaptive loop filter (ALF), deblocking filter (DB or DBF), sample adaptive offset (SAO) filter, etc.
[0314] Figure 22 This is a block diagram illustrating an example of the functional configuration of a loop filter 120 according to an embodiment. For example, as Figure 22As shown, the loop filter 120 includes a deblocking filter actuator 120a, a SAO actuator 120b, and an ALF actuator 120c. The deblocking filter actuator 120a performs deblocking filter processing on the reconstructed image. The SAO actuator 120b performs the SAO process on the reconstructed image after the deblocking filter process. The ALF actuator 120c performs the ALF process on the reconstructed image after the SAO process. The ALF and deblocking filters will be described in detail later. The SAO process is used to enhance image quality by reducing ringing (a phenomenon where pixel values are distorted like waves around edges) and correcting deviations in pixel values. Examples of SAO processes include edge offsetting and band offsetting processes. It should be noted that in some embodiments, the loop filter 120 may not include... Figure 22 The loop filter 120 may include all the constituent elements disclosed herein, and may include some of the constituent elements, as well as additional elements. Additionally, the loop filter 120 may be configured to interact with... Figure 22 The processing order disclosed in the document may vary, and different processing orders may be used to execute the above process, or not all processes may be executed, etc.
[0315] (Loop filter > Adaptive loop filter)
[0316] In ALF, a least-squares error filter is applied to remove compression artifacts. For example, a filter selected from multiple filters is applied for each 2×2 pixel sub-block in the current block, based on the direction and activity of the local gradient.
[0317] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of several categories (e.g., fifteen or twenty-five categories). The classification of sub-blocks can be based on, for example, gradient directionality and activity. In the example, a category index C (e.g., C = 5D + A) is calculated or determined based on gradient directionality D (e.g., 0 to 2 or 0 to 4) and gradient activity A (e.g., 0 to 4). Then, based on the category index C, each sub-block is classified into one of several categories.
[0318] For example, gradient directionality D is calculated by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, gradient activity A is calculated, for example, by summing the gradients in multiple directions and quantizing the sum.
[0319] Based on the results of this classification, the filter to be used for each sub-block can be determined from multiple filters.
[0320] The filter shape to be used in ALF is, for example, a circularly symmetrical filter shape. Figures 23A to 23C This is a conceptual diagram used to illustrate an example of the filter shape used in ALF. Figure 23AA 5×5 diamond filter is shown. Figure 23B A 7×7 rhombus filter is shown, and Figure 23C A 9×9 diamond-shaped filter is shown. Information indicating the filter shape is typically communicated at the picture level using signals. It should be noted that this signaling of filter shape does not necessarily need to be performed at the picture level; it can also be performed at another level (e.g., sequence level, slice level, partition level, CTU level, or CU level).
[0321] For example, the on / off state of ALF can be determined at the picture level or the CU level. For instance, it can be determined at the CU level whether to apply ALF to luminance, and at the picture level whether to apply ALF to chrominance. Information indicating whether ALF is on or off is typically signaled at the picture or CU level. It should be noted that signaling information indicating whether ALF is on or off does not necessarily need to be performed at the picture or CU level; it can be performed at another level (e.g., sequence level, slice level, partition level, or CTU level).
[0322] Additionally, as described above, a filter is selected from multiple filters, and the ALF procedure for the sub-block is performed. The set of coefficients to be used for each of the multiple filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that signaling the coefficient set does not necessarily need to be performed at the picture level and can be performed at another level (e.g., sequence level, slice level, fragment level, CTU level, CU level, or sub-block level).
[0323] (Loop filter > Cross-component adaptive loop filter)
[0324] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF). Figure 23E It is used to illustrate in CC-ALF (e.g., Figure 23D A conceptual diagram illustrating an example of the filter shape used in CC-ALF. Figure 23D and Figure 23E The example form of CC-ALF operates by applying a linear diamond filter to the luminance channel of each chromaticity component. For example, the filter coefficients can be sent in the APS, scaled by a factor of 2^10, and rounded for a fixed-point representation. For example, in... Figure 23D In this context, the Y sample (first component) is used for CCALF for Cb and CCALF for Cr (a component different from the first component).
[0325] The application of filters can be controlled on a variable block size and is signaled by a flag received for context decoding for each sample block. The block size, along with the CC-ALF enable flag, can be received at the slice level for each chroma component. CC-ALF can support various block sizes, such as (in chroma samples) 16×16 pixels, 32×32 pixels, 64×64 pixels, and 128×128 pixels.
[0326] (Loop filter > Joint chroma cross-component adaptive loop filter)
[0327] An example of Union Chromaticity-CCALF is in Figure 23F and Figure 23G As shown in the image. Figure 23F This is a conceptual diagram used to illustrate an example process for joint chromaticity CCALF. Figure 23G This is a table showing example weighted index candidates. As shown, a CCALF filter is used to generate a CCALF-filtered output as a chroma refinement signal for one color component, while a weighted version of the same chroma refinement signal is applied to another color component. In this way, the complexity of the existing CCALF is reduced by approximately half. The weights can be decoded into a symbol flag and a weighted index. The weighted index (denoted as weight_index) can be decoded into 3 bits, specifying the magnitude of the JC-CCALF weight JcCcWeight, which is a non-zero magnitude. For example, the magnitude of JcCcWeight can be determined as follows:
[0328] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index >> 2;
[0329] Otherwise, JcCcWeight equals 4 / (weight_index-4).
[0330] The block-level on / off control for ALF filtering of Cb and Cr can be separate. This is the same as in CCALF, and two separate sets of block-level on / off control flags can be decoded. Unlike CCALF, the block sizes for Cb and Cr on / off control are the same in this paper, so only one block size variable can be decoded.
[0331] (Loop filter > Deblocking filter)
[0332] During the deblocking filtering process, the loop filter 120 performs a filtering process on the block boundaries in the reconstructed image in order to reduce the distortion that occurs at the block boundaries.
[0333] Figure 24 This shows a loop filter 120 used as a deblocking filter (see [link]). Figure 7 and Figure 22 A block diagram of an example configuration of the deblocking filter actuator 120a.
[0334] The deblocking filter actuator 120a includes: a boundary determiner 1201; a filter determiner 1203; a filter actuator 1205; a process determiner 1208; a filter characteristic determiner 1207; and switches 1202, 1204, and 1206.
[0335] Boundary determiner 1201 determines whether the pixel to be deblocked (i.e., the current pixel) exists around the block boundary. Boundary determiner 1201 then outputs the determination result to switch 1202 and processing determiner 1208.
[0336] If boundary determiner 1201 has determined that the current pixel exists around the block boundary, switch 1202 outputs an unfiltered image to switch 1204. Conversely, if boundary determiner 1201 has determined that the current pixel does not exist around the block boundary, switch 1202 outputs an unfiltered image to switch 1206. It should be noted that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.
[0337] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0338] If the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, switch 1204 outputs the unfiltered image obtained through switch 1202 to filter executor 1205. If the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel, switch 1204 outputs the unfiltered image obtained through switch 1202 to switch 1206.
[0339] When an unfiltered image is obtained via switches 1202 and 1204, filter actuator 1205 performs deblocking filtering on the current pixel using the filter characteristics determined by filter characteristic determiner 1207. Filter actuator 1205 then outputs the filtered pixel to switch 1206.
[0340] Under the control of the processing determinant 1208, the switch 1206 selectively outputs one of the pixels that have not yet been deblocked and the pixels that have been deblocked by the filter actuator 1205.
[0341] The processing determiner 1208 controls the switch 1206 based on the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around a block boundary, and when the filter determiner 1203 has determined that deblocking filtering should be performed on the current pixel, the processing determiner 1208 causes the switch 1206 to output the pixel that has already been deblocked. Additionally, in other cases, the processing determiner 1208 causes the switch 1206 to output pixels that have not yet been deblocked. By repeatedly outputting pixels in this manner, a filtered image is output from the switch 1206. It should be noted that... Figure 24 The configuration shown is an example of a configuration in the deblocking filter executor 120a. The deblocking filter executor 120a can have various configurations.
[0342] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics about block boundaries.
[0343] During deblocking filtering, pixel values and quantization parameters can be used to select one of two deblocking filters with different characteristics (i.e., a strong filter and a weak filter). In the case of a strong filter, when pixels p0 to p2 and pixels q0 to q2 cross block boundaries, such as... Figure 25 As shown, by performing calculations, for example, according to the following expression, the pixel values of the corresponding pixels q0 to q2 are changed to pixel values q'0 to q'2.
[0344] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8
[0345] q'1=(p0+q0+q1+q2+2) / 4
[0346] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0347] It should be noted that in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the corresponding pixels p0 to p2 and q0 to q2. Additionally, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of the block boundary relative to pixel q2. Furthermore, on the right-hand side of each expression, the coefficient multiplied by the corresponding pixel value of the pixel to be used for deblocking filtering is the filter coefficient.
[0348] Furthermore, in deblocking filtering, clipping can be performed to ensure that the calculated pixel value change does not exceed a threshold. For example, during clipping, a threshold determined based on the quantization parameters can be used to clip the pixel value calculated according to the above expression to a value obtained by "calculating pixel value ± 2 × threshold". In this way, over-smoothing can be prevented.
[0349] Figure 26 It is a conceptual diagram used to illustrate the block boundaries on which the deblocking filtering process is performed. Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.
[0350] The block boundaries on which the deblocking filtering process is performed are, for example, the boundaries between CUs, Pus, or TUs that have 8×8 pixel blocks, such as... Figure 26 As shown in the diagram. The deblocking filtering process can be performed, for example, in units of four rows or four columns. First, for... Figure 26 Blocks P and Q are shown in the diagram to determine the boundary strength (Bs) value, as follows: Figure 27 As instructed in the document.
[0351] according to Figure 27 The Bs value determines whether to perform deblocking filtering on block boundaries belonging to the same image using different intensities. When the Bs value is 2, deblocking filtering is performed on the chroma signal. When the Bs value is 1 or greater and certain conditions are met, deblocking filtering is performed on the luminance signal. These conditions can be predetermined. It should be noted that the conditions used to determine the Bs value are not limited to... Figure 27 The values indicated in the text can be used to determine the Bs value based on another parameter.
[0352] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))
[0353] Figure 28 This is a flowchart illustrating an example of the process performed by the predictor of encoder 100. It should be noted that the predictor includes all or some of the following constituent elements: intra-frame predictor 124; inter-frame predictor 126; and prediction controller 128. The prediction actuator includes, for example, intra-frame predictor 124 and inter-frame predictor 126.
[0354] The predictor generates a predicted image for the current block (step Sb_1). This predicted image can also be referred to as a predicted signal or a predicted block. It should be noted that the predicted signal is, for example, an intra-frame predicted image (image prediction signal) or an inter-frame predicted image (inter-frame prediction signal). The predictor generates the predicted image for the current block using the reconstructed image obtained from another block through the generation of the predicted image, the generation of the prediction residual, the generation of the quantized coefficients, the recovery of the prediction residual, and the addition of the predicted image.
[0355] The reconstructed image can be, for example, an image in a reference image, or an image of an encoded block in the current image that includes the current block (i.e., the other blocks described above). An encoded block in the current image is, for example, a neighboring block of the current block.
[0356] Figure 29This is a flowchart illustrating another example of the process performed by the predictor of encoder 100.
[0357] The predictor generates a predicted image using the first method (step Sc_1a), the second method (step Sc_1b), and the third method (step Sc_1c). The first, second, and third methods can be different from each other in generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0358] Next, the prediction processor evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates a cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the predicted images by comparing the cost C of the predicted images. Note that the cost C can be calculated, for example, according to an expression of the RD optimization model (e.g., C = D + λ × R). In this expression, D represents the compression artifacts of the predicted image and is expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. Additionally, R represents the bit rate of the stream. Additionally, λ represents a multiplier, for example, according to the Lagrange multiplier method.
[0359] Then, the predictor selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects the method or mode used to obtain the final predicted image. For example, the predictor selects the predicted image with the lowest cost C based on the cost C calculated for the predicted image. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 can be based on parameters used during the encoding process. The encoder 100 can transform information used to identify the selected predicted image, method, or mode into a stream. This information can be, for example, flags. In this way, the decoder 200 is able to generate a predicted image based on this information according to the method or mode selected by the encoder 100. Note that in Figure 29 In the example shown, after generating the predicted image using the appropriate method, the predictor selects any one of the predicted images. However, the predictor can select a method or mode based on the parameters used in the encoding process described above before generating the predicted image, and can generate the predicted image according to the selected method or mode.
[0360] For example, the first and second methods can be intra-frame prediction and inter-frame prediction, respectively, and the predictor can select the final predicted image for the current block from the predicted images generated according to the prediction method.
[0361] Figure 30 This is a flowchart illustrating another example of the process performed by the predictor of encoder 100.
[0362] First, the predictor uses intra-frame prediction to generate a predicted image (step Sd_1a) and inter-frame prediction to generate a predicted image (step Sd_1b). Note that the predicted image generated by intra-frame prediction is also called the intra-frame predicted image, and the predicted image generated by inter-frame prediction is also called the inter-frame predicted image.
[0363] Next, the predictor evaluates each of the intra-frame and inter-frame predicted images (step Sd_2). The cost C described above can be used in the evaluation. The predictor can then select the predicted image for which the minimum cost C has been computed from the intra-frame and inter-frame predicted images as the final predicted image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the predicted image for the current block is selected.
[0364] (Intra-frame predictor)
[0365] Intra-predictor 124 generates a prediction signal (i.e., an intra-predicted image) by performing intra-prediction (also referred to as intra-prediction) of the current block, referencing one or more blocks in the current image and stored in block memory 118. More specifically, intra-predictor 124 generates an intra-predicted image by performing intra-prediction, referencing pixel values (e.g., luminance and / or chrominance values) of one or more adjacent blocks, and then outputs the intra-predicted image to prediction controller 128.
[0366] For example, the intra predictor 124 performs intra prediction by using one of a plurality of predefined intra prediction modes. Intra prediction modes typically include one or more non-directional prediction modes and multiple directional prediction modes. The defined modes can be predefined.
[0367] One or more non-directional prediction modes include, for example, the planar prediction mode and the DC prediction mode defined in the H.265 / High Efficiency Video Decoding (HEVC) standard.
[0368] Multiple directional prediction modes include, for example, the thirty-three directional prediction modes defined in the H.265 / HEVC standard. Note that in addition to the thirty-three directional prediction modes, multiple directional prediction modes may also include thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31This is a conceptual diagram illustrating a total of sixty-seven intra-prediction modes (two non-directional prediction modes and sixty-five directional prediction modes) that can be used in intra-prediction. Solid arrows represent the thirty-three directions defined in the H.265 / HEVC standard, and dashed arrows represent the additional thirty-two directions. Figure 31 (Two non-directional prediction modes are not shown in the image).
[0369] In various processing examples, the luma block can be referenced in the intra-prediction of the chroma block. In other words, the chroma component of the current block can be predicted based on the luma component of the current block. This intra-prediction is also known as cross-component linear model (CCLM) prediction. An intra-prediction mode for the chroma block that references such a luma block (also known as, for example, CCLM mode) can be added as one of the intra-prediction modes for the chroma block.
[0370] Intra-predictor 124 can correct the intra-predicted pixel values based on horizontal / vertical reference pixel gradients. Intra-prediction accompanied by this correction is also known as position-dependent intra-prediction combination (PDPC). Information indicating whether PDPC is applied (e.g., referred to as a PDPC flag) is typically signaled at the CU level. Note that signaling this information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, or CTU level).
[0371] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor 124.
[0372] Intra-predictor 124 selects one intra-prediction mode from multiple intra-prediction modes (step Sw_1). Intra-predictor 124 then generates a predicted image based on the selected intra-prediction mode (step Sw_2). Next, intra-predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra-prediction modes. For example, two of the six intra-prediction modes may be planar mode and DC prediction mode, and the other four modes may be directional prediction modes. Intra-predictor 124 determines whether the intra-prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0373] Here, when it is determined that the intra-prediction mode selected in step Sw_1 is included in the MPM ("Yes" in step Sw_4), the intra-predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the selected intra-prediction mode in the MPM (step Sw_6). Note that the MPM flag set to 1 and the information indicating the intra-prediction mode can be encoded into prediction parameters by the entropy encoder 110.
[0374] When it is determined that the selected intra-prediction mode is not included in the MPM (No in step Sw_4), the intra-predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra-predictor 124 does not set any MPM flag. The intra-predictor 124 then generates information indicating the selected intra-prediction mode that is not included in the MPM in at least one intra-prediction mode (step Sw_8). Note that the MPM flag set to 0 and the information indicating the intra-prediction mode can be encoded as prediction parameters by the entropy encoder 110. The information indicating the intra-prediction mode indicates, for example, any one of 0 to 60.
[0375] (Inter-frame predictor)
[0376] Inter-frame predictor 126 generates a predicted image (inter-frame predicted image) by performing inter-frame prediction (also called inter-frame prediction) on the current block, referencing one or more blocks in a reference image. This reference image is different from the current image and is stored in frame memory 122. Inter-frame prediction is performed on a unit of the current block or the current sub-block within the current block (e.g., a 4×4 block). Sub-blocks are included within blocks and are smaller than blocks. The size of sub-blocks can be in the form of slices, bricks, pictures, etc.
[0377] For example, inter-frame predictor 126 performs motion estimation in a reference image for the current block or sub-block and finds a reference block or reference sub-block that best matches the current block or sub-block. Inter-frame predictor 126 then obtains motion information (e.g., motion vectors) that compensates for motion or changes from the reference block or reference sub-block to the current block or sub-block. Inter-frame predictor 126 generates an inter-frame predicted image of the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. Inter-frame predictor 126 outputs the generated inter-frame predicted image to prediction controller 128.
[0378] Motion information used in motion compensation can be signaled in various forms as inter-frame prediction signals. For example, motion vectors can be signaled. As another example, the difference between motion vectors and motion vector predictors can be signaled.
[0379] (List of reference images)
[0380] Figure 33 This is a concept diagram used to illustrate a reference image. Figure 34 This is a conceptual diagram used to illustrate an example of a list of reference images. The list of reference images is a list indicating at least one reference image stored in frame memory 122. Note that in... Figure 33In the diagram, each image within a rectangle indicates its reference relationship, each arrow indicates its reference relationship, the horizontal axis indicates time, and the I, P, and B symbols within the rectangles indicate intra-frame predicted images, single-predicted images, and double-predicted images, respectively. The numbers within the rectangles also indicate the decoding order. For example... Figure 33 As shown, the decoding order of the images is I0, P1, B2, B3, and B4, and the display order of the images is I0, B3, B2, B4, and P1. Figure 34 As shown, the reference image list is a list representing candidate reference images. For example, an image (or slice) may include at least one reference image list. For instance, one reference image list is used when the current image is a single-prediction image, and two reference image lists are used when the current image is a double-prediction image. Figure 33 and Figure 34 In the example, image B3, which is the current image currPic, has two lists of reference images: the L0 list and the L1 list. When the current image currPic is image B3, the candidate reference images for the current image currPic are I0, P1, and B2, and the list of reference images (i.e., the L0 list and the L1 list) indicates these images. The inter-frame predictor 126 or the prediction controller 128 specifies which image in each list to actually reference in the form of a reference image index refidxLx. Figure 34 In the text, reference images P1 and B2 are specified by reference image indices refIdxL0 and refIdxL1.
[0381] Such a list of reference images can be generated for each unit, such as a sequence, picture, slice, brick, CTU, or CU. Furthermore, the reference image index indicating the reference image to be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, brick level, CTU level, or CU level. Additionally, a common list of reference images can be used in multiple inter-frame prediction modes.
[0382] (Basic process of inter-frame prediction)
[0383] Figure 35 This is a flowchart illustrating the basic processing flow of an example of inter-frame prediction.
[0384] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates the difference between the current block and the prediction image as the prediction residual (step Se_4).
[0385] Here, when generating the predicted image, the inter-frame predictor 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Furthermore, when determining the MV, the inter-frame predictor 126 determines the MV by selecting motion vector candidates (MV candidates) (step Se_1) and deriving the MV (step Se_2). The selection of MV candidates is performed, for example, by the inter-frame predictor 126 generating a list of MV candidates and selecting at least one MV candidate from the list. Note that previously derived MVs can be added to the MV candidate list. Alternatively, when deriving the MV, the inter-frame predictor 126 can also select at least one MV candidate from at least one MV candidate and determine the selected at least one MV candidate as the MV for the current block. Alternatively, the inter-frame predictor 126 can determine the MV for the current block by performing estimation in a reference image region specified by each of the selected at least one MV candidate. Note that the estimation in the reference image region can be referred to as motion estimation.
[0386] Additionally, although steps Se_1 to Se_3 are performed by the inter-frame predictor 126 in the example described above, the processes of steps Se_1, Se_2, etc., can be performed by another constituent element included in the encoder 100.
[0387] Note that an MV candidate list can be generated for each procedure in inter-frame prediction mode, or a common MV candidate list can be used in multiple inter-frame prediction modes. The procedures in steps Se_3 and Se_4 correspond to... Figure 9 Steps Sa_3 and Sa_4 are shown in the diagram. The process in step Sa_3 corresponds to... Figure 30 The process in step Sd_1b.
[0388] (Derivation process of motion vectors)
[0389] Figure 36 This is a flowchart illustrating an example of the derivation process of the motion vector.
[0390] The inter-frame predictor 126 can derive the MV of the current block in a mode used for encoding motion information (e.g., MV). In this case, for example, the motion information can be encoded as prediction parameters and can be signaled. In other words, the encoded motion information is included in the stream.
[0391] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where motion information is not encoded. In this case, motion information is not included in the stream.
[0392] Here, the MV derivation mode can include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., which will be described later. Modes in which motion information is encoded include the normal inter-frame mode, normal merging mode, and affine mode (specifically, affine inter-frame mode and affine merging mode). Note that motion information can include not only MV but also motion vector predictor selection information, which will be described later. Modes in which motion information is not encoded include FRUC mode, etc. The inter-frame predictor 126 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.
[0393] Figure 37 This is a flowchart illustrating another example of the derivation of motion vectors.
[0394] The inter-frame predictor 126 can derive the MV for the current block in a mode where the MV difference is encoded. In this case, for example, the MV difference can be encoded as a prediction parameter and can be signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. Note that the MV predictor is a motion vector predictor.
[0395] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where the MV difference is not encoded. In this case, the encoded MV difference is not included in the stream.
[0396] Here, as described above, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., which will be described later. Among the modes, those in which the MV difference is encoded include the normal inter-frame mode, affine mode (specifically, affine inter-frame mode), etc. Among the modes in which the MV difference is not encoded, those include the FRUC mode, normal merging mode, affine mode (specifically, affine merging mode), etc. The inter-frame predictor 126 selects a mode from the multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.
[0397] (Motion vector derivation mode)
[0398] Figure 38A and Figure 38B This is a conceptual diagram used to illustrate example classifications of patterns used for MV derivation. For example, such as... Figure 38A As shown, based on whether motion information and MV difference are encoded, MV derivation modes are broadly classified into three modes: inter-frame mode, merge mode, and frame rate up-conversion (FRUC) mode. Inter-frame mode is a mode in which motion estimation is performed, and where motion information and MV difference are encoded. For example, as... Figure 38BAs shown, inter-frame modes include affine inter-frame mode and normal inter-frame mode. Merging mode is a mode in which no motion estimation is performed, and where a motion difference (MV) is selected from the encoded surrounding blocks and used to derive the MV for the current block. Merging mode is a mode in which motion information is essentially encoded and the MV difference is not encoded. For example, as... Figure 38B The merging modes shown include normal merging mode (also known as normal merging mode or regular merging mode), motion vector difference merging (MMVD) mode, combined inter-frame merging / intra-frame prediction (CIIP) mode, triangle mode, ATMVP mode, and affine merging mode. Here, among the merging modes, the MV difference is anomalously encoded in the MMVD mode. Note that affine merging mode and affine inter-frame mode are modes included within affine modes. An affine mode is used to derive the MV of each of the multiple sub-blocks included in the current block as the MV of the current block, assuming an affine transformation. FRUC mode is a mode used to derive the MV of the current block by performing estimation between encoded regions, where neither motion information nor any MV difference is encoded. Note that the corresponding modes will be described in more detail later.
[0399] Notice, Figure 38A and Figure 38B The classification of modes shown is an example, and the classification is not limited to this. For example, when MV difference is encoded in CIIP mode, CIIP mode is classified as an inter-frame mode.
[0400] (MV derivation > Normal inter-frame mode)
[0401] The normal inter-frame mode is an inter-frame prediction mode used to derive the MV of the current block from a reference image region specified by the MV candidate based on blocks similar to the current block in the image. In this normal inter-frame mode, the MV difference is encoded.
[0402] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in normal inter-frame mode.
[0403] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple encoded blocks surrounding the current block in time or space) (step Sg_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0404] Next, the inter-frame predictor 126 extracts N (2 or larger integers) MV candidates from the plurality of MV candidates obtained in step Sg_1 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the determined priority order (step Sg_2). Note that the priority order can be predetermined for each of the N MV candidates.
[0405] Next, the inter-frame predictor 126 selects a motion vector predictor candidate from the N predicted motion vector candidates as the motion vector predictor (also referred to as the MV predictor) for the current block (step Sg_3). At this point, the inter-frame predictor 126 encodes motion vector predictor selection information in the stream to identify the selected motion vector predictor. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.
[0406] Next, the inter-frame predictor 126 derives the MV of the current block by referring to the encoded reference image (step Sg_4). At this point, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as the MV difference within the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130. Note that the encoded reference image includes multiple blocks that have been reconstructed after encoding.
[0407] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference image (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in a slice, the inter-frame prediction of the slice using normal inter-frame mode ends. For example, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in an image, the inter-frame prediction of the image using normal inter-frame mode ends. Note that not all blocks included in a slice may undergo the processes in steps Sg_1 to Sg_5, and the inter-frame prediction of the slice using normal inter-frame mode may end when only a portion of a block undergoes the process. This also applies to the processes in steps Sg_1 to Sg_5. When the process is performed on only a portion of a block in an image, the inter-frame prediction of the image using normal inter-frame mode may end.
[0408] Note that the predicted image is the inter-frame prediction signal as described above. Additionally, information indicating the inter-frame prediction mode used to generate the predicted image (normal inter-frame mode in the example above) is encoded, for example, as prediction parameters in the encoded signal.
[0409] Note that the MV candidate list can also be used as a list in other modes. Additionally, procedures associated with the MV candidate list can be applied to procedures associated with lists used in another mode. Procedures associated with the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering MV candidates, or deleting MV candidates.
[0410] (MV derivation > Normal merge mode)
[0411] Normal merging mode is used to select MV candidates from the MV candidate list as the MV of the current block, thereby deriving the inter-frame prediction mode of the MV. Note that normal merging mode is a type of merging mode and can be simply referred to as merging mode. In this embodiment, normal merging mode and merging mode are distinguished, and merging mode is used in a broader sense.
[0412] Figure 40 This is a flowchart illustrating an example of inter-frame prediction in normal merging mode.
[0413] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple encoded blocks surrounding the current block in time or space) (step Sh_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0414] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter-frame predictor 126 encodes MV selection information in the stream to identify the selected MV candidate. In other words, the inter-frame predictor 126 outputs the MV selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0415] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference image (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed on all blocks in a slice, the inter-frame prediction of the slice using the normal merging mode ends. Additionally, when the processes in steps Sh_1 to Sh_3 are performed on all blocks in an image, the inter-frame prediction of the image using the normal merging mode ends. Note that not all blocks included in a slice can undergo the processes in steps Sh_1 to Sh_3, and the inter-frame prediction of the slice using the normal merging mode may end when only a portion of a block undergoes the process. This also applies to the processes in steps Sh_1 to Sh_3. When the process is performed on only a portion of a block in an image, the inter-frame prediction of the image using the normal merging mode may end.
[0416] Alternatively, for example, information indicating the inter-frame prediction mode (normal merging mode in the example above) used to generate the predicted image and included in the encoded signal is encoded as prediction parameters in the stream.
[0417] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process performed on the current image through normal merging mode.
[0418] First, the inter-frame predictor 126 generates a list of MV candidates, in which MV candidates are registered. Examples of MV candidates include: spatially adjacent MV candidates, which are the MVs of multiple encoded blocks spatially surrounding the current block; temporally adjacent candidate MVs, which are the MVs of blocks surrounding the position of the current block in the encoded reference picture on which the current block is projected; combined MV candidates, which are MVs generated by combining the MV values of spatially adjacent MV predictors and temporally adjacent MV predictors; and zero MV candidates, which are MVs with a value of zero.
[0419] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates registered in the MV candidate list and determines that MV candidate as the MV of the current block.
[0420] In addition, the entropy encoder 110 writes and encodes merge_idx in the stream, which is a signal indicating which MV candidate has been selected.
[0421] Note that in Figure 41 The MV candidates registered in the MV candidate list described are examples. The number of MV candidates may differ from the number of MV candidates in the diagram, and the MV candidate list may be configured in such a way that some of the types of MV candidates in the diagram may not be included, or one or more MV candidates other than those in the diagram may be included.
[0422] The final MV can be determined by performing Dynamic Motion Vector Refresh (DMVR), described later, using the MV of the current block derived from the normal merge mode. Note that in normal merge mode, motion information is encoded and the MV difference is not. In MMVD mode, an MV candidate is selected from the MV candidate list, and the MV difference is encoded, as in normal merge mode. Figure 38B As shown, MMVD can be classified as a merge mode along with normal merge modes. Note that the MV difference in MMVD mode does not always need to be the same as the MV difference used in inter-frame mode. For example, MV difference derivation in MMVD mode can be a process that requires less processing power than MV difference derivation in inter-frame mode.
[0423] Alternatively, a combined inter-frame merge / intra-frame prediction (CIIP) mode can be performed. This mode is used to overlap the prediction images generated in inter-frame prediction and the prediction images generated in intra-frame prediction to generate a prediction image for the current block.
[0424] Note that the MV candidate list can also be called a candidate list. Additionally, merge_idx contains MV selection information.
[0425] (MV Derivation > HMVP Pattern)
[0426] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process for the current image using the HMVP merging pattern.
[0427] In normal merge mode, the MV for, for example, the current block (CU) is determined by selecting an MV candidate from a list of MV candidates generated from a reference encoded block (e.g., CU). Here, another MV candidate can be registered in the MV candidate list. The mode in which this other MV candidate is registered is called the HMVP mode.
[0428] In HMVP mode, HMVP’s First-In-First-Out (FIFO) server is used to manage MV candidates, separate from the MV candidate list used in normal merge mode.
[0429] In a FIFO buffer, motion information such as the MV (Motion Value) of the most recently processed blocks is stored first. When managing the FIFO buffer, each time a block is processed, the MV for the most recent block (i.e., the CU that was just processed) is stored in the FIFO buffer, and the MV for the oldest CU (i.e., the earliest processed CU) is removed from the FIFO buffer. Figure 42 In the example shown, HMVP1 is the MV for the latest block, and HMVP5 is the MV for the oldest block.
[0430] Then, for example, the inter-frame predictor 126 checks, starting with HMVP1, whether each MV managed in the FIFO buffer is a different MV from all MV candidates already registered in the MV candidate list for normal merge mode. When it is determined that an MV is different from all MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer as an MV candidate in the MV candidate list for normal merge mode. At this point, one or more of the MV candidates in the FIFO buffer can be registered (added to the MV candidate list).
[0431] By using the HMVP pattern in this way, not only can the MVs of blocks spatially or temporally adjacent to the current block be added, but also the MVs of blocks processed in the past can be added. This expands the variation of MV candidates for the normal merge pattern, increasing the potential for improved decoding efficiency.
[0432] Note that MV can be motion information. In other words, the information stored in the MV candidate list and FIFO buffer can include not only MV values, but also reference image information, reference orientation, number of images, etc. Additionally, a block can be, for example, a CU.
[0433] Notice, Figure 42 The MV candidate list and FIFO buffer shown are examples. The size of the MV candidate list and FIFO buffer can be... Figure 42 The differences in, or can be configured to, with Figure 42 The MV candidates are registered in a different order. Additionally, the process described herein may be shared between encoder 100 and decoder 200.
[0434] Note that the HMVP pattern can be applied to patterns other than the normal merge pattern. For example, motion information such as the MV of blocks previously processed in the affine pattern can be stored first and used as MV candidates, which can better improve efficiency. The pattern obtained by applying the HMVP pattern to the affine pattern can be called the historical affine pattern.
[0435] (MV Derivation > FRUC Pattern)
[0436] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In an embodiment, motion estimation is performed on the decoder side without using any pixel values in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values in the current block include frame rate up-conversion (FRUC) mode, pattern matching motion vector derivation (PMMVD) mode, etc.
[0437] Figure 43 The diagram shows an example of the FRUC process in flowchart form. First, a list is used to indicate the MVs for the encoded blocks (each of which is spatially or temporally adjacent to the current block) as MV candidates by referring to the MVs (this list can be an MV candidate list or can be used as an MV candidate list for the normal merge mode) (step Si_1).
[0438] Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation values of the corresponding MV candidates included in the MV candidate list are calculated, and an MV candidate is selected based on the evaluation values. Based on the selected motion vector candidate, the motion vector for the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector for the current block. Alternatively, for example, pattern matching in the region surrounding a position in the reference image, where the position in the reference image corresponds to the selected motion vector candidate, can be used to derive the motion vector for the current block. In other words, the estimation using pattern matching and evaluation values can be performed in the region surrounding the best MV candidate, and when an MV that produces a better evaluation value exists, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV for the current block. In some embodiments, updating the motion vector that produces a better evaluation value may not be performed.
[0439] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference image (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in an image, the inter-frame prediction of the image using the FRUC mode ends. Note that not all blocks included in a slice can undergo the processes in steps Si_1 to Si_5, and the inter-frame prediction of the slice using the FRUC mode may end when only a portion of a block undergoes the process. Similarly, the inter-frame prediction of the image using the FRUC mode may end when the processes in steps Si_1 to Si_5 are performed on a portion of blocks included in an image in a similar manner.
[0440] A similar process can be performed on a sub-block basis.
[0441] Evaluation values can be calculated using various methods. For example, a reconstructed image of a region in a reference image corresponding to a motion vector is compared with a reconstructed image of a defined region (which could be, for example, a region in another reference image or a region in a neighboring block of the current image, as indicated below). The defined region can be predetermined.
[0442] The difference between pixel values in two reconstructed images can be used as an evaluation value for the motion vector. Note that information other than the difference can be used to calculate the evaluation value.
[0443] Next, an example of pattern matching is described in detail. First, a candidate MV included in the candidate MV list (e.g., a merge list) is selected as the starting point for estimation by pattern matching. For example, a first pattern matching or a second pattern matching can be used as the pattern matching. The first pattern matching and the second pattern matching can be referred to as bilateral matching and template matching, respectively.
[0444] (MV derivation > FRUC > bilateral matching)
[0445] In the first pattern matching, pattern matching is performed between two blocks located along the motion trajectory of the current block and included in two different reference images. Therefore, in the first pattern matching, the region in the other reference image that follows the motion trajectory of the current block is used as the defined region for calculating the candidate evaluation values described above. The defined region can be predetermined.
[0446] Figure 44 This is a conceptual diagram used to illustrate an example of first-order pattern matching (bilateral matching) between two blocks along a motion trajectory in two reference images. Figure 44 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the best matching pair between pairs of two blocks located along the motion trajectory of the current block (Cur block) and included in two different reference images (Ref0, Ref1). More specifically, for the current block, the difference between the reconstructed image at a specified position in the first encoded reference image (Ref0) specified by the MV candidate and the reconstructed image at a specified position in the second encoded reference image (Ref1) specified by the symmetric MV obtained by scanning the MV candidate at display time intervals is derived, and the value of the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value and is likely to produce good results can be selected as the final MV from among the multiple MV candidates.
[0447] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) of two reference blocks are specified to be proportional to the temporal distances (TD0, TD1) between the current image (CurPic) and the two reference images (Ref0, Ref1). For example, when the current image is temporally located between two reference images and the temporal distances from the current image to the corresponding two reference images are equal, a mirror-symmetric bidirectional motion vector is derived in the first pattern matching.
[0448] (MV derivation > FRUC > Template matching)
[0449] In the second pattern matching (template matching), pattern matching is performed between a block in the reference image and a template in the current image (the template being a block in the current image adjacent to the current block (adjacent blocks being, for example, the upper and / or left adjacent blocks)). Therefore, in the second pattern matching, the blocks in the current image adjacent to the current block are used as the defined regions for calculating the evaluation values of the MV candidates described above.
[0450] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 45 As shown, in the second pattern matching, the motion vector of the current block (Cur block) is derived by estimating the block in the reference image (Ref0) that best matches the block adjacent to the current block in the current image (Cur Pic). More specifically, the difference between the reconstructed image in the encoded region (which is adjacent to the current block in the upper left, left, or upper left) and the reconstructed image in the corresponding region in the encoded reference image (Ref0) specified by the MV candidate is derived, and the value of the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.
[0451] Information indicating whether a FRUC mode is applied (e.g., referred to as the FRUC flag) can be signaled at the CU level. Additionally, when a FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the applicable mode matching method (e.g., first mode matching or second mode matching) can be signaled at the CU level. Note that signaling this information does not necessarily need to be performed at the CU level; it can be performed at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0452] (MV derivation > Affine mode)
[0453] Affine patterns are used to generate motion vectors (MVs) using affine transformations. For example, MVs can be derived on a sub-block basis based on the motion vectors of multiple neighboring blocks. This pattern is also known as an affine motion compensation prediction pattern.
[0454] Figure 46A This is a conceptual diagram used to illustrate an example of MV derivation based on motion vectors of multiple adjacent blocks, on a sub-block basis. Figure 46AIn this context, the current block comprises, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the top-left control point in the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V1 at the top-right control point in the current block is derived based on the motion vectors of adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vector (v1) for the corresponding sub-blocks in the current block can be derived. x ,v y ).
[0455] [Mathematics 1]
[0456]
[0457] Here, x and y indicate the horizontal and vertical positions of the sub-block, respectively, and w indicates the determined weighting coefficients. The determined weighting coefficients can be predetermined.
[0458] This information indicating an affine mode (e.g., referred to as an affine flag) can be signaled at the CU level. It should be noted that signaling information indicating an affine mode does not necessarily need to be performed at the CU level, and can also be performed at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0459] Additionally, affine modes can include several modes for different methods of deriving motion vectors at the top-left and top-right control points. For example, affine modes include two modes: affine inter-frame mode (also known as affine normal inter-frame mode) and affine merge mode.
[0460] (MV derivation > Affine mode)
[0461] Figure 46B This is a conceptual diagram used to illustrate an example of MV derivation in a sub-block manner within an affine pattern that uses three control points. Figure 46B In this context, the current block comprises, for example, sixteen 4×4 blocks. Here, the motion vector V0 at the top-left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector V1 at the top-right control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V2 at the bottom-left control point of the current block is derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v1, v2, v2) for the corresponding sub-blocks in the current block can be derived. x ,v y ).
[0462] [Mathematics 2]
[0463]
[0464] Here, x and y indicate the horizontal and vertical positions of the sub-block, respectively, and w and h can be weighting coefficients, which can be predetermined weighting coefficients. In an embodiment, w can indicate the width of the current block, and h can indicate the height of the current block.
[0465] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0466] Additionally, this affine mode using three control points can include different methods for deriving motion vectors at the top-left, top-right, and bottom-left control points. For example, as with the affine mode using two control points, the affine mode using three control points can include both affine inter-frame mode and affine merge mode.
[0467] Note that in affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels, and can also be another size. For example, the size of each sub-block can be 8x8 pixels.
[0468] (MV derivation > Affine pattern > Control point)
[0469] Figure 47A , Figure 47B and Figure 47C This is a conceptual diagram used to illustrate an example of MV derivation at the control point in affine mode.
[0470] like Figure 47A As shown, in affine mode, for example, a motion vector predictor at the corresponding control point of the current block is calculated based on multiple motion vectors corresponding to blocks encoded according to the affine mode among the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block. More specifically, the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in the listed order, and the first valid block encoded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on multiple motion vectors corresponding to the identified block.
[0471] For example, such as Figure 47BAs shown, when block A, which is adjacent to the current block on the left, has been encoded according to an affine pattern using two control points, motion vectors v3 and v4 projected at the top-left and top-right corners of the encoded block including block A are derived. Then, motion vector v0 at the top-left control point and motion vector v1 at the top-right control point of the current block are calculated based on the derived motion vectors v3 and v4.
[0472] For example, such as Figure 47C As shown, when block A, which is adjacent to the current block on the left, has been encoded according to an affine pattern using three control points, motion vectors v3, v4, and v5 projected onto the top-left, top-right, and bottom-left positions of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, motion vector v0 at the top-left control point of the current block, motion vector v1 at the top-right control point of the current block, and motion vector v2 at the bottom-left control point of the current block are calculated.
[0473] Figures 47A to 47C The MV derivation method shown can be used for Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in the figure, or can be used for the description later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown in the figure.
[0474] Figure 48A and Figure 48B This is a conceptual diagram used to illustrate an example of MV derivation at the control point in affine mode.
[0475] Figure 48A This is a conceptual diagram used to illustrate an example affine pattern in which two control points are used.
[0476] In affine mode, such as Figure 48A As shown, the MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block.
[0477] Figure 48B This is a conceptual diagram used to illustrate an example affine pattern that uses three control points.
[0478] In affine mode, such as Figure 48BAs shown, the MV selected from the MVs of the coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of the coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block. Furthermore, the MV selected from the MVs of the coded blocks F and G adjacent to the current block is used as the motion vector v2 at the bottom-left control point of the current block.
[0479] Notice, Figure 48A and Figure 48B The MV derivation method shown can be used in the descriptions that follow. Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in the figure, or can be used for the description later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown in the figure.
[0480] Here, when affine patterns using different numbers of control points (e.g., two and three control points) can be switched at the CU level and signaled, the number of control points in the encoded block and the number of control points in the current block can be different from each other.
[0481] Figure 49A and Figure 49B This is a conceptual diagram illustrating an example of a method for MV derivation at control points when the number of control points in the encoded block and the number of control points in the current block are different from each other.
[0482] For example, such as Figure 49A As shown, the current block has three control points at the top left, top right, and bottom left corners, and the block A adjacent to the current block on the left has been encoded according to an affine pattern using two control points. In this case, motion vectors v3 and v4 projected onto the top left and top right corners of the encoded block including block A are derived. Then, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated based on the derived motion vectors v3 and v4. Furthermore, motion vector v2 at the bottom left control point is calculated based on the derived motion vectors v0 and v1.
[0483] For example, such as Figure 49BAs shown, the current block has two control points at the top left and top right corners, and the block A adjacent to the current block on the left has been encoded according to an affine pattern using three control points. In this case, motion vectors v3, v4, and v5 projected onto the top left, top right, and bottom left corners of the encoded block, including block A, are derived. Then, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0484] Notice, Figure 49A and Figure 49B The MV derivation method shown can be used in the descriptions that follow. Figure 50 The MV derivation at each control point of the current block in step Sk_1 shown in the figure, or can be used for the description later. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown in the figure.
[0485] (MV Derivation > Affine Pattern > Affine Merging Pattern)
[0486] Figure 50 This is a flowchart illustrating an example of a process in affine merging mode.
[0487] In the affine merging mode as shown, firstly, the inter-frame predictor 126 derives the MV at the corresponding control point of the current block (step Sk_1). Figure 46A As shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 46B As shown, the control points are the top-left corner, top-right corner, and bottom-left corner of the current block. The inter-frame predictor 126 can encode MV selection information to identify two or three derived MVs in the stream.
[0488] For example, when using Figures 47A to 47C When the MV derivation method is shown in the figure, such as Figure 47A As shown, the inter-frame predictor 126 examines the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in the listed order and identifies the first valid block encoded according to the affine pattern.
[0489] Inter-frame predictor 126 uses the first valid block identified, encoded according to the identified affine pattern, to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47BAs shown, the inter-frame predictor 126 calculates motion vector v0 at the top-left control point of the current block and motion vector v1 at the top-right control point of the current block based on motion vectors v3 and v4 at the top-left and top-right control points of the encoded block, including block A. For example, the inter-frame predictor 126 calculates motion vector v0 at the top-left control point of the current block and motion vector v1 at the top-right control point of the current block by projecting motion vectors v3 and v4 at the top-left and top-right control points of the encoded block onto the current block.
[0490] Alternatively, when block A is identified and block A has three control points, such as Figure 47C As shown, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point, the motion vector v1 at the top-right control point, and the motion vector v2 at the bottom-left control point of the current block based on the motion vectors v3, v4, and v5 of the coded block A, including the top-left, top-right, and bottom-left control points. For example, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point, the motion vector v1 at the top-right control point, and the motion vector v2 at the bottom-left control point of the current block by projecting the motion vectors v3, v4, and v5 of the coded block at the top-left, top-right, and bottom-left control points onto the current block.
[0491] Note that, as described above Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, as described above. Figure 49B As shown, when block A is identified and block A has three control points, the MV at two control points can be calculated.
[0492] Next, the inter-frame predictor 126 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates an affine MV for each of the multiple sub-blocks, for example, using two motion vectors v0 and v1 and the above expression (1A) or using three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter-frame predictor 126 then uses these affine MVs and the encoded reference image to perform motion compensation for the sub-blocks (step Sk_3). The process of generating a predicted image using the affine merging pattern for the current block ends when the processes in steps Sk_2 and Sk_3 are performed for each of all sub-blocks included in the current block. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.
[0493] Note that the MV candidate list described above can be generated in step Sk_1. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown in the figure Figure 49A and Figure 49B The MV derivation method shown in the figure, and any combination of other MV derivation methods.
[0494] Note that the MV candidate list can include MV candidates in modes that perform prediction on a sub-block basis, other than affine mode.
[0495] Note that, for example, an MV candidate list can be generated that includes MV candidates in an affine merge pattern using two control points and an affine merge pattern using three control points. Alternatively, an MV candidate list can be generated separately that includes MV candidates in an affine merge pattern using two control points and an MV candidate list that includes MV candidates in an affine merge pattern using three control points. Alternatively, an MV candidate list can be generated that includes MV candidates in one of the affine merge patterns using two control points and an affine merge pattern using three control points. The MV candidates (multiple) can be, for example, MVs for encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or MVs for valid blocks within a block.
[0496] Note that the index indicating one of the MVs in the MV candidate list can be sent as MV selection information.
[0497] (MV Derivation > Affine Mode > Affine Inter-Frame Mode)
[0498] Figure 51 This is a flowchart illustrating an example of the process in affine inter-frame mode.
[0499] In affine inter-frame mode, firstly, the inter-frame predictor 126 derives the MV predictor (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_1). The control points can be, for example, the top-left corner, the top-right corner, and the bottom-left corner of the current block, such as... Figure 46A or Figure 46B As shown in the image.
[0500] For example, when using Figure 48A and Figure 48B When the MV derivation method shown in the figure is used, the inter-frame predictor 126 selects in Figure 48A or Figure 48B The MV of any one of the encoded blocks near the corresponding control point of the current block is shown in the figure to derive the MV predictor (v0, v1) or (v0, v1, v2) at the corresponding two or three control points of the current block. At this time, the inter-frame predictor 126 encodes MV predictor selection information in the stream to identify the selected two or three MV predictors.
[0501] For example, the inter-frame predictor 126 can use cost evaluation or the like to determine the block from which the MV predictor at the control point is selected from the encoded blocks adjacent to the current block, and can write a flag indicating which MV predictor has been selected into the bit stream. In other words, the inter-frame predictor 126 outputs MV predictor selection information, such as the flag, as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.
[0502] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 uses the expression (1A) or expression (1B) described above to compute the MV of each sub-block corresponding to the updated MV predictor as an affine MV (step Sj_3). The inter-frame predictor 126 then uses these affine MVs and encoded reference images to perform motion compensation for the sub-blocks (step Sj_4). When the MV predictor is updated in step Sj_2, the process in steps Sj_3 and Sj_4 is performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor as the MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0503] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation for the current block using the determined MV and the encoded reference image (step Sj_6).
[0504] Note that the MV candidate list described above can be generated in step Sj_1. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown in the figure Figure 49A and Figure 49BThe MV derivation method shown in the figure, and any combination of other MV derivation methods.
[0505] Note that the MV candidate list can include MV candidates in modes that perform prediction on a sub-block basis, other than affine mode.
[0506] Note that, for example, an MV candidate list can be generated that includes MV candidates in an affine inter-frame pattern using two control points and an affine inter-frame pattern using three control points. Alternatively, an MV candidate list can be generated separately, including MV candidates in an affine inter-frame pattern using two control points and MV candidates in an affine inter-frame pattern using three control points. Alternatively, an MV candidate list can be generated that includes MV candidates in one of the affine inter-frame patterns using two control points and an affine inter-frame pattern using three control points. The MV candidates (multiple) can be, for example, MVs for encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or MVs for valid blocks within a block.
[0507] Note that the index indicating one of the MVs in the MV candidate list can be sent as MV predictor selection information.
[0508] (MV Derivation > Triangle Pattern)
[0509] In the example above, the inter-frame predictor 126 generates a rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 can generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and can combine multiple prediction images to generate a final rectangular prediction image. The shape other than a rectangle can be, for example, a triangle.
[0510] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangular predicted images.
[0511] Inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a first partition with a triangular shape in the current block using a first MV of a first partition. Similarly, inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a second partition with a triangular shape in the current block using a second MV of a second partition. Then, inter-frame predictor 126 combines these prediction images to generate a prediction image with a rectangular shape that is the same as the rectangular shape of the current block.
[0512] Note that a first prediction image with a rectangular shape corresponding to the current block can be generated using a first prediction image (MV) as a prediction image for the first partition. Alternatively, a second prediction image with a rectangular shape corresponding to the current block can be generated using a second prediction image (MV) as a prediction image for the second partition. A prediction image for the current block can be generated by performing a weighted summation of the first and second prediction images. Note that the weighted summation can be of a portion of the boundary between the first and second partitions.
[0513] Figure 52B This is a conceptual diagram illustrating an example of a first portion of a first partition that overlaps with a second partition, and a first set and a second set of samples that can be weighted as part of a correction process. The first portion may have, for example, one-quarter of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to the edges of the first partition, where N is a positive integer, for example, N could be an integer 2. As shown, Figure 52B The example on the left shows a rectangular partition with a rectangular portion whose width is one-quarter of the width of the first partition, wherein a first set of samples includes samples outside the first part and samples inside the first part, and a second set of samples includes samples inside the first part. Figure 52B The central example shows a rectangular partition with a rectangular portion whose height is one-quarter of the height of the first partition, wherein a first set of samples includes samples outside the first part and samples inside the first part, and a second set of samples includes samples inside the first part. Figure 52B The example on the right shows a triangular partition with polygonal portions, the height of which corresponds to two samples, wherein the first set of samples includes samples outside the first part and samples inside the first part, and the second set of samples includes samples inside the first part.
[0514] The first part can be the portion of the first partition that overlaps with the adjacent partition. Figure 52C This is a conceptual diagram illustrating a first portion of a first partition, which is a portion of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition with an overlapping portion is shown. Partitions with other shapes (e.g., triangular partitions) may be used, and the overlapping portion may overlap with partitions that are spatially or temporally adjacent.
[0515] Additionally, although an example is given in which inter-frame prediction is used to generate a predicted image for each of the two partitions, intra-frame prediction can be used to generate a predicted image for at least one partition.
[0516] Figure 53This is a flowchart illustrating an example of a process in triangle mode.
[0517] In triangle mode, firstly, the inter-frame predictor 126 splits the current block into a first partition and a second partition (step Sx_1). At this point, the inter-frame predictor 126 can encode the partition information, which is related to the partitioning, as prediction parameters in the stream. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.
[0518] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple encoded blocks surrounding the current block in time or space) (step Sx_2). In other words, the inter-frame predictor 126 generates a list of MV candidates.
[0519] Inter-frame predictor 126 then selects the MV candidate for the first partition and the MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_1, respectively, as the first MV and the second MV (step Sx_3). At this time, inter-frame predictor 126 encodes the MV selection information used to identify the selected MV candidate as prediction parameters in the stream. In other words, inter-frame predictor 126 outputs the MV selection information as prediction parameters to entropy encoder 110 through prediction parameter generator 130.
[0520] Next, the inter-frame predictor 126 performs motion compensation using a selected first MV and an encoded reference image to generate a first predicted image (step Sx_4). Similarly, the inter-frame predictor 126 performs motion compensation using a selected second MV and an encoded reference image to generate a second predicted image (step Sx_5).
[0521] Finally, the inter-frame predictor 126 generates a prediction image for the current block by performing a weighted summation of the first and second prediction images (step Sx_6).
[0522] Note that, although in Figure 52A In the example shown, the first and second partitions are triangles, but the first and second partitions can be trapezoids or other shapes that are different from each other. Furthermore, although the current block is... Figure 52A and Figure 52C The example shown includes two partitions, but the current block can include three or more partitions.
[0523] Furthermore, the first and second partitions can overlap. In other words, the first and second partitions can include the same pixel region. In this case, the predicted image in the first partition and the predicted image in the second partition can be used to generate a predicted image for the current block.
[0524] Additionally, while an example has been shown in which inter-frame prediction is used to generate a predicted image for each of two partitions, intra-frame prediction can be used to generate a predicted image for at least one partition.
[0525] Note that the MV candidate list used to select the first MV and the MV candidate list used to select the second MV can be different from each other, or the MV candidate list used to select the first MV can also be used as the MV candidate list used to select the second MV.
[0526] Note that partition information may include indexes indicating the split direction in which at least the current block is split into multiple partitions. MV selection information may include indexes indicating the selected first MV and indexes indicating the selected second MV. An index may indicate multiple pieces of information. For example, an index that jointly indicates part or all of the partition information and part or all of the MV selection information may be encoded.
[0527] (MV Derivation > ATMVP Pattern)
[0528] Figure 54 This is a conceptual diagram used to illustrate an example of an advanced temporal motion vector prediction (ATMVP) pattern in which MV is derived on a sub-block basis.
[0529] The ATMVP pattern is a pattern that is classified as a merge pattern. For example, in the ATMVP pattern, the MV candidate for each sub-block is registered in the MV candidate list for use in the normal merge pattern.
[0530] More specifically, in the ATMVP model, firstly, as Figure 54 As shown, a temporal MV reference block associated with the current block is identified in an encoded reference image specified by the MV (MV0) of the neighboring block located at the lower left position relative to the current block. Next, in each sub-block of the current block, an MV is identified to encode the region corresponding to the sub-block in the temporal MV reference block. MVs identified in this way are included as MV candidates in an MV candidate list for use in the current block. When selecting an MV candidate for each sub-block from the MV candidate list, the sub-block undergoes motion compensation, where the MV candidate is used as the MV for the sub-block. In this way, a predicted image is generated for each sub-block.
[0531] Although Figure 54 In the example shown, the block located at the bottom left relative to the current block is used as a reference block for the surrounding MV, but it should be noted that another block can be used. Additionally, the size of the child block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the child block can be toggled in units such as slices, bricks, images, etc.
[0532] (Motion estimation > DMVR)
[0533] Figure 55 This is a flowchart illustrating the relationship between the merging mode and the decoder motion vector refinement DMVR.
[0534] Inter-frame predictor 126 derives the motion vector for the current block based on the merging mode (step S1_1). Next, inter-frame predictor 126 determines whether to perform motion vector estimation, i.e., motion estimation (step S1_2). Here, when it is determined that motion estimation should not be performed ("No" in step S1_2), inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). In other words, in this case, the motion vector for the current block is determined based on the merging mode.
[0535] When it is determined in step Sl_1 that motion estimation will be performed ("Yes" in step Sl_2), the inter-frame predictor 126 derives the final motion vector for the current block by estimating the region surrounding the reference image specified by the motion vector derived in step Sl_1 (step Sl_3). In other words, in this case, the motion vector of the current block is determined according to the DMVR.
[0536] Figure 56 This is a conceptual diagram illustrating an example of the DMVR process used to determine the MV.
[0537] First, for example in merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference image (L0), which is an encoded image in the L0 list, based on the MV candidate (L0). Similarly, reference pixels are identified from the second reference image (L1), which is an encoded image in the L1 list, based on the MV candidate (L1). A template is generated by calculating the average of these reference pixels.
[0538] Next, a template is used to estimate each of the MV candidates in the surrounding regions of the first reference image (L0) and the second reference image (L1), and the MV that produces the minimum cost is determined as the final MV. Note that the cost can be calculated, for example, using the difference between each pixel value in the template and the corresponding pixel value in the estimated region, the value of the MV candidate, etc.
[0539] It is not always necessary to perform the exact same procedure described here. Other procedures can be used to derive the final MV by estimating the region surrounding the MV candidate.
[0540] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV. Unlike Figure 56 The example of DMVR shown in [the image] Figure 57 In the example shown, the cost is calculated without generating a template.
[0541] First, the inter-frame predictor 126 estimates the surrounding region of the reference block in each of the reference images included in the L0 and L1 lists based on the initial MVs as MV candidates obtained from each MV candidate list. For example, as Figure 57 As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets a search position for the reference images in the L0 list. Based on the position indicated by the vector difference indicating the search position to be set, specifically the initial MV (i.e., InitMV_L0), the vector difference with the search position is MVd_L0. The inter-frame predictor 126 then determines the estimated position in the reference images in the L1 list. This search position is indicated by the vector difference from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. More specifically, the inter-frame predictor 126 determines the vector difference MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines the search position in each reference image in the L0 and L1 lists as the position symmetrical with respect to the position indicated by the initial MV. The inter-frame predictor 126 calculates the sum of the absolute differences (SADs) between the values of the pixels at the search position in the block as the cost for each search position, and finds the search position that produces the minimum cost.
[0542] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in a DMVR, and Figure 58B This is a flowchart illustrating an example of the motion estimation process.
[0543] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also known as the start point) and eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost is minimized at each of the search positions other than the start point. Here, when it is determined that the cost is minimized at a search position other than the start point, the inter-frame predictor 126 changes its target to the search position where the minimum cost is obtained and executes the process in step 2. When the cost is minimized at the start point, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.
[0544] In step 2, the inter-frame predictor 126 performs a search similar to that in step 1, taking the search position after the target change as the new starting point based on the result of the process in step 1. Then, the inter-frame predictor 126 determines whether the cost is minimized at each search position other than the starting point. Here, when it is determined that the cost is minimized at each search position other than the starting point, the inter-frame predictor 126 performs the process in step 4. When the cost is minimized at the starting point, the inter-frame predictor 126 performs the process in step 3.
[0545] In step 4, the inter-frame predictor 126 considers the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0546] In step 3, the inter-frame predictor 126 determines the pixel position with sub-pixel accuracy that yields the minimum cost based on the costs at four points located above, below, to the left, and to the right relative to the starting point in step 1 or step 2, and treats this pixel position as the final search position. The pixel position with sub-pixel accuracy is determined by performing a weighted summation of each of the four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) for the above, below, left, and right sides using the costs corresponding to one of the four search positions as weights. The inter-frame predictor 126 then determines the vector difference between the position indicated by the initial MV and the final search position.
[0547] (Motion compensation > BIO / OBMC / LIC)
[0548] Motion compensation involves modes used to generate and correct predicted images. These modes include, for example, bidirectional optical flow (BIO), overlapping block motion compensation (OBMC), local illumination compensation (LIC), etc., which will be described later.
[0549] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.
[0550] Inter-frame predictor 126 generates a predicted image (step Sm_1) and corrects the predicted image, for example, according to any of the modes described above (step Sm_2).
[0551] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.
[0552] Inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, inter-frame predictor 126 uses the motion vector to generate a predicted image (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, when it is determined that a correction process should be performed ("Yes" in step Sn_3), inter-frame predictor 126 generates a final predicted image by correcting the predicted image (step Sn_4). Note that in the LIC described later, illumination and chroma can be corrected in step Sn_4. When it is determined that no correction process should be performed ("No" in step Sn_3), inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).
[0553] (Motion compensation > OBMC)
[0554] Note that in addition to the motion information obtained through motion estimation for the current block, motion information for neighboring blocks can also be used to generate inter-frame prediction images. More specifically, inter-frame prediction images can be generated for each sub-block in the current block by performing a weighted summation of the prediction image based on the motion information obtained through motion estimation (in the reference image) and the prediction image based on the motion information of neighboring blocks (in the current image). This inter-frame prediction (motion compensation) is also known as Overlapping Block Motion Compensation (OBMC) or OBMC mode.
[0555] In OBMC mode, information indicating the size of the sub-block used for OBMC can be signaled at the sequence level (e.g., referred to as the OBMC block size). Additionally, information indicating whether OBMC mode is applied can be signaled at the CU level (e.g., referred to as the OBMC flag). Note that signaling this information does not necessarily need to be performed at the sequence and CU levels, and can be performed at another level (e.g., picture level, slice level, brick level, CTU level, or sub-block level).
[0556] The OBMC model will be described in more detail. Figure 61 and Figure 62 These are flowcharts and conceptual diagrams used to illustrate an outline of the predictive image correction process performed by OBMC.
[0557] First, such as Figure 62 As shown, the MV assigned to the current block is used to obtain the predicted image (Pred) through normal motion compensation. Figure 62 In the image, the arrow "MV" points to the reference image and indicates the content referenced by the current block of the current image in order to obtain the predicted image.
[0558] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) derived for the coded block adjacent to the current block to the left (reusing the motion vector used for the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which points to the reference image from the current block. A first correction to the predicted image is performed by overlaying the two predicted images, Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0559] Similarly, a predicted image (Pred_U) is obtained by applying the MV (MV_U) already derived for the coded block adjacent to the current block (reusing the MV for the current block) to the current block. The MV (MV_U) is indicated by the arrow "MV_U", which points to the reference image from the current block. A second correction to the predicted image is performed by overlapping the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) that have already undergone the first correction. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is an image in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.
[0560] While the example above uses a two-path correction method with adjacent blocks on the left and top, it should be noted that correction methods can also use three-path or more-path correction methods with adjacent blocks on the right and / or bottom.
[0561] Note that the region where this overlap is performed can be only a portion of the region near the block boundary, rather than the entire pixel region of the block.
[0562] Note that the prediction image correction process based on OBMC for obtaining a prediction image Pred from a reference image by overlapping and appending prediction images Pred_L and Pred_U has already been described above. However, when correcting prediction images based on multiple reference images, a similar process can be applied to each of the multiple reference images. In this case, after obtaining the corrected prediction images from the corresponding reference images by performing OBMC image correction based on multiple reference images, the obtained corrected prediction images are further overlapped to obtain the final prediction image.
[0563] Note that in OBMC, the current block cell can be a PU, or a sub-block cell obtained by further splitting the PU.
[0564] An example of a method for determining whether to apply OBMC is the use of obmc_flag as a signal indicating whether OBMC is applied. As a concrete example, encoder 100 can determine whether the current block belongs to a region with complex motion. When the block belongs to a region with complex motion, encoder 100 sets obmc_flag to the value "1" and applies OBMC during encoding; when the block does not belong to a region with complex motion, encoder 100 sets obmc_flag to the value "0" and encodes the block without applying OBMC. Decoder 200 switches between applying and not applying OBMC by decoding obmc_flag in the write stream.
[0565] (Motion compensation > BIO)
[0566] Next, the derivation method for MV is described. First, a mode for deriving MV based on a model assuming uniform linear motion is described. This mode is also known as the bidirectional optical flow (BIO) mode. Additionally, this bidirectional optical flow can be written as BDOF instead of BIO.
[0567] Figure 63 This is a conceptual diagram used to illustrate a model that assumes uniform linear motion. Figure 63 In the middle, (v x v y The vector τ0 and τ1 indicate the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1). x0 MV y0 ) indicates the MV corresponding to the reference image Ref0, and (MV x1 MV y1 ) indicates the MV corresponding to the reference image Ref1.
[0568] Here, it is assumed that the velocity vector (v) x v y It exhibits uniform linear motion, (MV) x0 MV y0 ) and (MV x1 MV y1 ) are respectively represented as (v xτ0 v yτ0 ) and (-v xτ1 -v yτ1 ), and the following optical flow equation (2) is given.
[0569] [Mathematics 3]
[0570]
[0571] Here, I(k) indicates the luminance value k (k = 0, 1) of the motion-compensated reference image after motion compensation. The optical flow equation states that the sum of the following terms equals zero: (i) the temporal derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, an MV candidate list can be corrected pixel by pixel.
[0572] Note that motion vectors can be derived on the decoder side 200 using methods other than deriving motion vectors based on a model assuming uniform linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.
[0573] Figure 64 This is a flowchart illustrating an example of the process of inter-frame prediction based on BIO. Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor 126 that can perform inter-frame prediction based on BIO.
[0574] like Figure 65 As shown, the inter-frame predictor 126 includes, for example, a memory 126a, an interpolated image deducer 126b, a gradient image deducer 126c, an optical flow deducer 126d, a correction value deducer 126e, and a predicted image corrector 126f. Note that the memory 126a may be a frame memory 122.
[0575] Inter-frame predictor 126 uses two reference images (Ref0, Ref1) different from the image including the current block (Cur Pic) to derive two motion vectors (M0, M1). Then, inter-frame predictor 126 uses the two motion vectors (M0, M1) to derive the predicted image for the current block (step Sy_1). Note that motion vector M0 is the motion vector (MV) corresponding to the reference image Ref0. x0 MV y0 Furthermore, motion vector M1 is the motion vector (MV) corresponding to the reference image Ref1. x1 MV y1 ).
[0576] Next, the interpolated image derivator 126b uses the motion vector M0 and the reference image L0 via the reference memory 126a to derive the interpolated image I for the current block. 0 Next, the interpolation image derivator 126b uses the motion vector M1 and the reference image L1 via the reference memory 126a to derive the interpolation image I for the current block. 1 (Step Sy_2). Here, the interpolated image I 0It is the image included in the reference image Ref0 and derived for the current block, and the interpolated image I 1 It is the image included in reference image Ref1 and derived for the current block. Interpolated image I 0 and interpolated image I 1 Each element in the interpolated image can be the same size as the current block. Alternatively, the interpolated image I... 0 and interpolated image I 1 Each of the elements can be an image larger than the current block. Furthermore, the interpolated image I... 0 and interpolated image I 1 This can include a predicted image obtained by using motion vectors (M0, M1) and a reference image (L0, L1) and applying a motion compensation filter.
[0577] Additionally, the gradient image derivator 126c is based on the interpolated image I. 0 and interpolated image I 1 Derive the gradient image of the current block (Ix) 0 , Ix 1 ,Iy 0 ,Iy 1 (Step Sy_3). Note that the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 ,Iy 1 The gradient image derivator 126c can derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image can indicate the amount of spatial change in pixel values along the horizontal direction, along the vertical direction, or along both directions.
[0578] Next, the optical flow deriver 126d uses an interpolated image (I 0 I 1 ) and gradient image (Ix 0 , Ix 1 ,Iy 0 ,Iy 1 For each sub-block of the current block, derive the optical flow (vx, vy) as a velocity vector (step Sy_4). The optical flow indicates the coefficients used to correct the spatial pixel movement and can be referred to as a local motion estimate, a corrected motion vector, or a corrected weighted vector. As an example, a sub-block can be a 4×4 pixel sub-CU. Note that the optical flow derivation can be performed on a per-pixel unit basis, rather than on a per-sub-block basis.
[0579] Next, the inter-frame predictor 126 uses optical flow (vx, vy) to correct the predicted image for the current block. For example, the correction value derivator 126e uses optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_5). The predicted image corrector 126f can then use the correction values to correct the predicted image for the current block (step Sy_6). Note that the correction values can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis.
[0580] Note that the BIO process flow is not limited to Figure 64 The process is publicly disclosed. For example, it may be possible to execute only the... Figure 64 It is part of the publicly disclosed process, or different processes can be added or different processes can be used as alternatives, or processes can be executed in different processing orders, etc.
[0581] (Motion compensation > LIC)
[0582] Next, an example of a pattern used to generate a predicted image (prediction) using the Local Illumination Compensation (LIC) process is described.
[0583] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses an illumination correction process performed by a LIC. Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.
[0584] First, the inter-frame predictor 126 derives the MV from the encoded reference image and obtains the reference image corresponding to the current block (step Sz_1).
[0585] Next, the inter-frame predictor 126 extracts information indicating how the luminance value changes between the current block and the reference image (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region) in the current image, as well as the luminance pixel values at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 126 uses the information indicating how the luminance value changes to calculate the illumination correction parameters (step Sz_3).
[0586] The inter-frame predictor 126 generates a predicted image for the current block by performing an illumination correction process in which illumination correction parameters are applied to a reference image in a reference picture specified by the MV (step Sz_4). In other words, the predicted image, which serves as a reference image in a reference picture specified by the MV, is corrected based on the illumination correction parameters. This correction may adjust illumination, chroma, or both. In other words, chroma correction parameters can be calculated using information indicating how chroma changes, and a chroma correction process can be performed.
[0587] Notice, Figure 66A The shape of the surrounding reference area shown is an example; another shape can be used.
[0588] Furthermore, although the process of generating a predicted image from a single reference image has been described herein, the case of generating a predicted image from multiple reference images can be described in the same manner. The predicted image can be generated after performing an illumination correction process on the reference image obtained from the reference images, in the same way as described above.
[0589] An example of a method for determining whether to apply LIC is a method using lic_flag as a signal indicating whether LIC is applied. As a specific example, encoder 100 determines whether the current block belongs to an area with a change in illuminance. When the block belongs to an area with a change in illuminance, encoder 100 sets lic_flag to the value "1" and applies LIC during encoding; when the block does not belong to an area with a change in illuminance, encoder 100 sets lic_flag to the value "0" and performs encoding without applying LIC. Decoder 200 can decode the lic_flag written to the stream and decode the current block by switching between applying and not applying LIC based on the flag value.
[0590] One example of a different method for determining whether to apply the LIC procedure is based on whether the LIC procedure has already been applied to surrounding blocks. As a concrete example, when the current block is already being processed in merge mode, the inter-frame predictor 126 determines whether the encoded surrounding blocks selected in the MV derivation in merge mode have already been encoded using LIC. The inter-frame predictor 126 performs encoding by switching between applying and not applying LIC based on the result. Note that the same procedure is also applied on the decoder 200 side in this example.
[0591] The illuminance correction (LIC) process has been referenced Figure 66A and Figure 66B It has been described, and will be further described below.
[0592] First, the inter-frame predictor 126 derives the MV of the reference image corresponding to the current block to be encoded from the reference image, which is the encoded image.
[0593] Next, the inter-frame predictor 126 uses the luminance pixel values of the encoded surrounding reference regions adjacent to the left and top of the current block and the luminance values at the corresponding positions in the reference image specified by MV to extract information indicating how the luminance values of the reference image change to the luminance values of the current image, and calculates illumination correction parameters. For example, suppose the luminance pixel value of a given pixel in the surrounding reference region of the current image is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region of the reference image is p1. The inter-frame predictor 126 calculates coefficients A and B to optimize A×p1+B=p0 as illumination correction parameters for multiple pixels in the surrounding reference region.
[0594] Next, the inter-frame predictor 126 performs an illumination correction process using illumination correction parameters for the reference image in the reference picture specified by MV to generate a predicted image for the current block. For example, suppose the luminance pixel value in the reference image is p2, and the illumination-corrected luminance pixel value in the predicted image is p3. The inter-frame predictor 126 generates the predicted image after undergoing the illumination correction process by calculating A×p2+B=p3 for each pixel in the reference image.
[0595] For example, a region having a defined number of pixels extracted from each of its upper and left adjacent pixels can be used as a surrounding reference region. Furthermore, the surrounding reference region is not limited to regions adjacent to the current block; it can also be regions not adjacent to the current block. Figure 66A In the example shown, the surrounding reference region in the reference image can be a region specified by another MV in the current image from the surrounding reference region in the current image. For example, the other MV can be an MV in the surrounding reference region of the current image.
[0596] Although the operations performed by encoder 100 have been described here, it should be noted that decoder 200 performs similar operations.
[0597] Note that LIC can be applied not only to luminance but also to chrominance. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any of Y, Cb, and Cr.
[0598] Alternatively, the LIC procedure can be applied on a sub-block basis. For example, correction parameters can be derived using the surrounding reference regions in the current sub-block and the surrounding reference regions in the reference sub-block of the reference image specified by the MV of the current sub-block.
[0599] (Predictive Controller)
[0600] Prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from intra-frame predictor 124) and the inter-frame prediction signal (the image or signal output from inter-frame predictor 126), and outputs the selected prediction image to subtractor 104 and adder 116 as the prediction signal.
[0601] (Prediction parameter generator)
[0602] Prediction parameter generator 130 can output information related to intra-frame prediction, inter-frame prediction, and the selection of prediction images in prediction controller 128 as prediction parameters to entropy encoder 110. Entropy encoder 110 can generate a stream based on the prediction parameters input from prediction parameter generator 130 and quantized coefficients input from quantizer 108. The prediction parameters can be used in decoder 200. Decoder 200 can receive and decode the stream and perform the same process as the prediction process performed by intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128. Prediction parameters may include, for example, (i) selecting a prediction signal (e.g., MV, prediction type, or prediction mode used by intra-frame predictor 124 or inter-frame predictor 126), or (ii) optional indices, flags, or values indicating the prediction process based on the prediction process performed in each of intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128.
[0603] (Decoder)
[0604] Next, a decoder 200 capable of decoding the stream output from the encoder 100 described above will be described. Figure 67 This is a block diagram illustrating the functional configuration of the decoder 200 according to this embodiment. The decoder 200 is an apparatus for decoding a stream of encoded images on a block-by-block basis.
[0605] like Figure 67 As shown, decoder 200 includes entropy decoder 202, inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, loop filter 212, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, prediction parameter generator 222, and split determiner 224. Note that intra-frame predictor 216 and inter-frame predictor 218 are configured as part of the prediction executor.
[0606] (Decoder installation example)
[0607] Figure 68 This is a functional block diagram illustrating an example installation of decoder 200. Decoder 200 includes a processor b1 and a memory b2. For example, Figure 67The decoder 200 shown in the figure has multiple components mounted on it. Figure 68 The processor b1 and memory b2 are shown in the diagram.
[0608] Processor b1 is a circuit that performs information processing and is coupled to memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 can be a processor such as a CPU. Alternatively, processor b1 can be an assembly of multiple electronic circuits. Additionally, for example, processor b1 can perform... Figure 67 The roles of two or more constituent elements other than the constituent element used for storing information, etc., in the multiple constituent elements of the decoder 200 shown in the figure.
[0609] Memory b2 is a dedicated or general-purpose memory used by processor b1 to decode streams. Memory b2 can be an electronic circuit and can be connected to processor b1. Alternatively, memory b2 can be included within processor b1. Alternatively, memory b2 can be an assembly of multiple electronic circuits. Alternatively, memory b2 can be a disk, optical disk, etc., or can be represented as a storage device, recording medium, etc. Additionally, memory b2 can be non-volatile memory or volatile memory.
[0610] For example, memory b2 can store images or streams. Additionally, memory b2 can store programs used by processor b1 to decode the streams.
[0611] Alternatively, for example, memory b2 can serve as... Figure 67 The decoder 200 shown herein has multiple constituent elements, including the roles of two or more constituent elements used for storing information. More specifically, memory b2 can serve as... Figure 67 The roles of block memory 210 and frame memory 214 are shown in the diagram. More specifically, memory b2 can store reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.).
[0612] Note that this may not be implemented in decoder 200. Figure 67 All of the constituent elements indicated herein may not perform all the processes described herein. Figure 67 A portion of the constituent elements indicated herein may be included in another device, or a portion of the process described herein may be performed by another device.
[0613] The following describes the overall flow of the process performed by decoder 200, followed by a description of each of the constituent elements included in decoder 200. Note that some of the constituent elements included in decoder 200 perform the same processes as some of those performed in encoder 100, and therefore the same processes will not be described in detail again. For example, the inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, and loop filter 212 included in decoder 200 perform similar processes to those performed by the inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and loop filter 120 included in encoder 100.
[0614] (Overall flow of the decoding process)
[0615] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by decoder 200.
[0616] First, the split determiner 224 in decoder 200 determines the splitting pattern (step Sp_1) for each of the multiple fixed-size blocks (e.g., 128×128 pixels) included in the image based on parameters input from entropy decoder 202. This splitting pattern is selected by encoder 100. Decoder 200 then performs steps Sp_2 to Sp_6 for each of the multiple blocks in the splitting pattern.
[0617] The entropy decoder 202 decodes the encoded and quantized coefficients and the prediction parameters of the current block (specifically, performs entropy decoding) (step Sp_2).
[0618] Next, the inverse quantizer 204 performs inverse quantization on the multiple quantized coefficients, and the inverse transformer 206 performs inverse transformation on the result to recover the prediction residual (i.e., the difference block) (step Sp_3).
[0619] Next, all or some of the prediction executors, including the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220, generate the prediction signal for the current block (step Sp_4).
[0620] Next, adder 208 adds the predicted image to the predicted residual to generate the reconstructed image of the current block (also known as the decoded image block) (step Sp_5).
[0621] When the reconstructed image is generated, the loop filter 212 performs filtering on the reconstructed image (step Sp_6).
[0622] Decoder 200 then determines whether decoding of the entire image has ended (step Sp_7). If it is determined that decoding has not ended ("No" in step Sp_7), decoder 200 repeats the process that started from step Sp_1.
[0623] Note that these steps Sp_1 through Sp_7 can be executed sequentially by the decoder 200, or two or more steps in the process can be executed in parallel. The processing order of two or more steps in the process can be modified.
[0624] (Split Determiner)
[0625] Figure 70 This is a conceptual diagram used to illustrate the relationship between the split determiner 224 and other constituent elements in an embodiment. As an example, the split determiner 224 may perform the following process.
[0626] For example, the split determiner 224 collects block information from block memory 210 or frame memory 214 and further obtains parameters from entropy decoder 202. The split determiner 224 can then determine the splitting pattern for fixed-size blocks based on the block information and parameters. The split determiner 224 can then output information indicating the determined splitting pattern to inverse transformer 206, intra-frame predictor 216, and inter-frame predictor 218. Inverse transformer 206 can perform an inverse transform on the transform coefficients based on the splitting pattern indicated by the information from split determiner 224. Intra-frame predictor 216 and inter-frame predictor 218 can generate predicted images based on the splitting pattern indicated by the information from split determiner 224.
[0627] (Entropy Decoder)
[0628] Figure 71 This is a block diagram illustrating an example of the functional configuration of the entropy decoder 202.
[0629] Entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the splitting pattern by entropy decoding the stream. For example, CABAC is used for entropy decoding. More specifically, entropy decoder 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a debinarizer 202c. Binary arithmetic decoder 202a arithmetically decodes the stream into a binary signal using context values derived by context controller 202b. Context controller 202b derives context values based on the characteristics of the syntactic elements or the surrounding state (i.e., the probability of occurrence of the binary signal) in the same manner as context controller 110b of encoder 100. Debinarizer 202c performs debinarization to transform the binary signal output from binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. This binarization can be performed according to the binarization method described above.
[0630] Thus, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can then output the prediction parameters included in the stream (see...). Figure 1 The output is sent to intra-frame predictor 216, inter-frame predictor 218, and prediction controller 220. Intra-frame predictor 216, inter-frame predictor 218, and prediction controller 220 are capable of performing the same prediction process as those performed by intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128 on the encoder 100 side.
[0631] Figure 72 This is a conceptual diagram illustrating the flow of an example CABAC procedure in the entropy decoder 202.
[0632] First, initialization is performed in the CABAC within the entropy decoder 202. During initialization, initialization and setting of the initial context value are performed in the binary arithmetic decoder 202a. The binary arithmetic decoder 202a and debinarizer 202c then perform arithmetic decoding and debinarization of the encoded data, such as a CTU. Meanwhile, the context controller 202b updates the context value each time line arithmetic decoding is performed. The context controller 202b then saves the context value for post-processing. For example, the saved context value is used to initialize the context value for the next CTU.
[0633] (Inverse quantizer)
[0634] Inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block, which are input from entropy decoder 202. More specifically, inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. Inverse quantizer 204 then outputs the inverse-quantized transform coefficients (i.e., transform coefficients) of the current block to inverse transform 206.
[0635] Figure 73 This is a block diagram illustrating an example of the functional configuration of the inverse quantizer 204.
[0636] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage device 204d, and an inverse quantization executor 204e.
[0637] Figure 74 This is a flowchart illustrating an example of the dequantization process performed by the dequantizer 204.
[0638] As an example, the inverse quantizer 204 can be based on Figure 74 Each CU in the flow shown performs an inverse quantization process. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined that inverse quantization should be performed ("yes" in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameters for the current block from the entropy decoder 202 (step Sv_12).
[0639] Next, the predicted quantization parameter generator 204b obtains the quantization parameters for the processing unit different from the current block from the quantization parameter storage device 204d (step Sv_13). The predicted quantization parameter generator 204b generates the predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_14).
[0640] The quantization parameter generator 204a then generates quantization parameters for the current block based on the differential quantization parameters for the current block obtained from the entropy decoder 202 and the predicted quantization parameters for the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the differential quantization parameters for the current block obtained from the entropy decoder 202 and the predicted quantization parameters for the current block generated by the predicted quantization parameter generator 204b can be added to generate quantization parameters for the current block. Additionally, the quantization parameter generator 204a stores the quantization parameters for the current block in the quantization parameter storage device 204d (step Sv_16).
[0641] Next, the inverse quantization executor 204e uses the quantization parameters generated in step Sv_15 to inverse quantize the quantized coefficients of the current block into transform coefficients (step Sv_17).
[0642] Note that differential quantization parameters can be decoded at the bit sequence level, image level, slice level, brick level, or CTU level. Additionally, the initial values of the quantization parameters can be decoded at the sequence level, image level, slice level, brick level, or CTU level. In this case, the initial values of the quantization parameters and the differential quantization parameters can be used to generate the quantization parameters.
[0643] Note that the inverse quantizer 204 may include multiple inverse quantizers and may use an inverse quantization method selected from a variety of inverse quantization methods to inverse quantize the quantized coefficients.
[0644] (Inverse Transformer)
[0645] The inverse transformer 206 recovers the prediction residual by performing an inverse transformation on the transform coefficients, which are inputs from the inverse quantizer 204.
[0646] For example, when the information parsed from the stream indicates that EMT or AMT should be applied (e.g., when the AMT flag is true), the inverse transformer 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the transform type parsed.
[0647] Furthermore, for example, when the information parsed from the stream indicates that NSST should be applied, the inverse transformer 206 applies a second inverse transform to the transform coefficients.
[0648] Figure 75 This is a flowchart illustrating an example of the process performed by the inverse converter 206.
[0649] For example, inverse transformer 206 determines whether there is information in the stream indicating that an orthogonal transformation is not performed (step St_11). Here, when it is determined that such information does not exist ("No" in step St_11) (e.g., there is no indication of whether an orthogonal transformation is performed; there is an indication that an orthogonal transformation will be performed), inverse transformer 206 obtains information indicating the transformation type decoded by entropy decoder 202 (step St_12). Next, based on this information, inverse transformer 206 determines the transformation type used for the orthogonal transformation in encoder 100 (step St_13). Inverse transformer 206 then performs an inverse orthogonal transformation using the determined transformation type (step St_14). Figure 75 As shown, when it is determined that there is information indicating that an orthogonal transformation should not be performed ("Yes" in step St_11) (e.g., an explicit instruction not to perform an orthogonal transformation; no instruction to perform an orthogonal transformation), the orthogonal transformation is not performed.
[0650] Figure 76 This is a flowchart illustrating an example of the process performed by the inverse converter 206.
[0651] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a predetermined value (step Su_11). This predetermined value may be pre-determined. Here, when it is determined that the transform size is less than or equal to the predetermined value ("yes" in step Su_11), the inverse transformer 206 obtains information from the entropy decoder 202 indicating which transform type the encoder 100 used in at least one transform type included in the first transform type group (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transformer 206.
[0652] Based on this information, the inverse transformer 206 determines the transformation type for the orthogonal transformation in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transformation on the transformation coefficients of the current block using the determined transformation type (step Su_14). When it is determined that the transformation size is not less than or equal to a determined value ("No" in step Su_11), the inverse transformer 206 performs an inverse transformation on the transformation coefficients of the current block using the second transformation type group (step Su_15).
[0653] Note that, as an example, the inverse orthogonal transform performed by inverse transform 206 can be based on Figure 75 or Figure 76 The process shown is executed for each TU. Alternatively, the inverse orthogonal transform can be performed by using a defined transform type without decoding information indicating the transform type used for the orthogonal transform. The defined transform type can be a predefined transform type or a default transform type. Specifically, the transform type can be DST7, DCT8, etc. In the inverse orthogonal transform, the inverse transform basis functions corresponding to the transform type are used.
[0654] (Adder)
[0655] Adder 208 reconstructs the current block by adding the prediction residual, which is input from inverse transformer 206, and the prediction image, which is input from prediction controller 220. In other words, it generates a reconstructed image of the current block. Adder 208 then outputs the reconstructed image of the current block to block memory 210 and loop filter 212.
[0656] (Block memory)
[0657] Block memory 210 is a storage device for storing blocks included in the current image and referential to in intra-frame prediction. More specifically, block memory 210 stores the reconstructed image output from adder 208.
[0658] (Loop filter)
[0659] The loop filter 212 applies a loop filter to the reconstructed image generated by the adder 208, outputs the filtered reconstructed image to the frame memory 214, and provides the output of the decoder 200, for example, to a display device, etc.
[0660] When the information indicating that ALF is on or off from the stream resolution indicates that ALF is on, a filter can be selected from multiple filters, for example, based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0661] Figure 77 This is a block diagram illustrating an example of the functional configuration of loop filter 212. Note that loop filter 212 has a configuration similar to that of loop filter 120 of encoder 100.
[0662] For example, such as Figure 77 As shown, the loop filter 212 includes a deblocking filter executor 212a, a SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs a deblocking filter process on the reconstructed image. The SAO executor 212b performs a SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs an ALF process on the reconstructed image after the SAO process. Note that the loop filter 212 does not always need to include... Figure 77 All constituent elements disclosed herein may be included, but only a portion thereof may be included. Additionally, the loop filter 212 may be configured to interact with... Figure 77 The above process can be executed in a different order than the one disclosed in the document, or it may be omitted. Figure 77 All the processes shown in the diagram, etc.
[0663] (Frame Memory)
[0664] Frame memory 214 is, for example, a storage device for storing reference images used for inter-frame prediction, and may also be referred to as a frame buffer. More specifically, frame memory 214 stores the reconstructed image filtered by loop filter 212.
[0665] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))
[0666] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of decoder 200. Note that the prediction executor may include all or some of the following constituent elements: intra-frame predictor 216; inter-frame predictor 218; and prediction controller 220. The prediction executor includes, for example, intra-frame predictor 216 and inter-frame predictor 218.
[0667] The predictor generates a predicted image for the current block (step Sq_1). This predicted image can also be referred to as the predicted signal or the predicted block. It should be noted that the predicted signal is, for example, an intra-frame predicted image or an inter-frame predicted image. More specifically, the predictor generates the predicted image for the current block using a reconstructed image already obtained for another block, through the generation of the predicted image, the recovery of the predicted residual, and the addition of the predicted image. The predictor of decoder 200 generates the same predicted image as the predictor of encoder 100. In other words, the predicted image is generated according to a common method or a mutually corresponding method between the predictors.
[0668] The reconstructed image can be, for example, an image in a reference image, or an image of a decoded block in the current image that includes the current block (i.e., other blocks described above). A decoded block in the current image is, for example, a neighboring block of the current block.
[0669] Figure 79 This is a flowchart illustrating another example of the process performed by the predictor of decoder 200.
[0670] The predictor determines the method or pattern used to generate the predicted image (step Sr_1). For example, the method or pattern can be determined based on, for example, prediction parameters.
[0671] When the first method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the third method (step Sr_2c).
[0672] The first, second, and third methods can be different from each other for generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0673] Figure 80 This is a flowchart illustrating another example of the process performed by the predictor of decoder 200.
[0674] As an example, the predictor can be based on Figure 80 The flow shown in the diagram is used to perform the prediction process. Note that... Figure 80 The intra-block copying shown is a mode of inter-frame prediction, where the block included in the current image is called the reference image or reference block. In other words, intra-block copying does not reference images different from the current image. Furthermore, Figure 80The PCM mode shown is a type of intra-frame prediction mode, in which no transformation or quantization is performed.
[0675] (Intra-frame predictor)
[0676] Intra-predictor 216 performs intra-prediction based on the intra-prediction mode parsed from the stream by referencing blocks in the current image stored in block memory 210 to generate a predicted image of the current block (i.e., the intra-prediction block). More specifically, intra-predictor 216 performs intra-prediction by referencing pixel values (e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-prediction image, and then outputs the intra-prediction image to prediction controller 220.
[0677] Note that when the intra-prediction mode that references the luma block in the intra-prediction of the chroma block is selected, the intra-predictor 216 can predict the chroma component of the current block based on the luma component of the current block.
[0678] Furthermore, when the information parsed from the stream indicates that PDPC will be applied, the intra-predictor 216 corrects the intra-predicted pixel values based on the horizontal / vertical reference pixel gradients.
[0679] Figure 81 This is a diagram illustrating an example of the process performed by the intra-frame predictor 216 of the decoder 200.
[0680] The intra-frame predictor 216 first determines whether to use MPM. For example... Figure 81 As shown, the intra predictor 216 determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that an MPM flag indicating 1 exists ("Yes" in step Sw_11), the intra predictor 216 obtains information from the entropy decoder 202 indicating the intra prediction mode selected in the encoder 100 within the MPM. Note that this information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_13). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode included among the multiple intra prediction modes included in the MPM and indicated by the information obtained in step Sw_12 (step Sw_14).
[0681] When it is determined that the MPM flag indicating 1 does not exist (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information from entropy decoder 202 indicating the intra prediction mode selected in encoder 100 from at least one intra prediction mode that is never included in the MPM. Note that this information is decoded by entropy decoder 202 and output to intra predictor 216. Intra predictor 216 then determines the intra prediction mode that is not included among the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15 (step Sw_17).
[0682] Intra-predictor 216 generates a prediction image based on the intra-prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0683] (Inter-frame predictor)
[0684] The inter-frame predictor 218 predicts the current block by referring to a reference image stored in the frame memory 214. Prediction is performed on a unit basis: the current block or the current sub-block within the current block. Note that the sub-block is included within the block and is a unit smaller than the block. The size of the sub-block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub-block can be switched for units such as slices, bricks, images, etc.
[0685] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from a stream (e.g., prediction parameters output from the entropy decoder 202) and outputs the inter-frame prediction image to the prediction controller 220.
[0686] When the information parsed from the stream indicates that the OBMC mode should be applied, the inter-frame predictor 218 uses motion information of neighboring blocks to generate an inter-frame predicted image, in addition to the motion information of the current block obtained through motion estimation.
[0687] Furthermore, when the information parsed from the stream indicates that the FRUC mode should be applied, the inter-frame predictor 218 derives motion information by performing motion estimation based on the mode matching method parsed from the stream (e.g., bilateral matching or template matching). The inter-frame predictor 218 then uses the derived motion information to perform motion compensation (prediction).
[0688] Furthermore, when applying BIO mode, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. Additionally, when information from the stream parsing indicates that affine mode should be applied, the inter-frame predictor 218 derives the MV for each sub-block based on the MVs of multiple adjacent blocks.
[0689] (MV Derivation Process)
[0690] Figure 82 This is a flowchart illustrating an example of the MV derivation process in decoder 200.
[0691] For example, inter-frame predictor 218 determines whether to decode motion information (e.g., motion video). For example, inter-frame predictor 218 may make this determination based on prediction modes included in the stream, or based on other information included in the stream. Here, when it is determined that motion information should be decoded, inter-frame predictor 218 derives the motion video for the current block in a mode where the motion information is decoded. When it is determined that motion information should not be decoded, inter-frame predictor 218 derives the motion video in a mode where the motion information is not decoded.
[0692] Here, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., which will be described later. Among these modes, those where motion information is decoded include the normal inter-frame mode, normal merging mode, and affine mode (specifically, affine inter-frame mode and affine merging mode). Note that motion information may include not only MV but also MV predictor selection information, which will be described later. Modes where motion information is not decoded include FRUC mode, etc. The inter-frame predictor 218 selects a mode from multiple modes for deriving the MV for the current block and uses the selected mode to derive the MV for the current block.
[0693] Figure 83 This is a flowchart illustrating an example of the MV derivation process in decoder 200.
[0694] For example, the inter-frame predictor 218 can determine whether to decode the MV difference, i.e., based on the prediction mode included in the stream, or based on other information included in the stream. Here, when it is determined that the MV difference should be decoded, the inter-frame predictor 218 can derive the MV for the current block in the mode in which the MV difference is decoded. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.
[0695] When it is determined that no MV difference will be decoded, the inter-frame predictor 218 derives the MV in a mode where the MV difference is not decoded. In this case, the encoded MV difference is not included in the stream.
[0696] Here, as described above, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., which will be described later. Modes in which the MV difference is encoded include the normal inter-frame mode and the affine mode (specifically, the affine inter-frame mode). Modes in which the MV difference is not encoded include the FRUC mode, the normal merging mode, and the affine mode (specifically, the affine merging mode). The inter-frame predictor 218 selects a mode from the various modes for deriving the MV for the current block, and uses the selected mode to derive the MV for the current block.
[0697] (MV derivation > Normal inter-frame mode)
[0698] For example, when the information parsed from the stream indicates that a normal inter-frame mode should be applied, the inter-frame predictor 218 derives the MV based on the information parsed from the stream and uses the MV to perform motion compensation (prediction).
[0699] Figure 84 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using normal inter-frame mode.
[0700] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple decoded blocks surrounding the current block in time or space) (step Sg_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0701] Next, the inter-frame predictor 218 extracts N (2 or larger integers) MV candidates from the multiple MV candidates obtained in step Sg_11 as motion vector predictor candidates (also referred to as MV predictor candidates) according to the ranking in the determined priority order (step Sg_12). Note that the priority order ranking can be predetermined for the corresponding N MV predictor candidates, and this ranking can be predetermined.
[0702] Next, the inter-frame predictor 218 decodes the MV predictor selection information from the input stream and uses the decoded MV predictor selection information to select one MV predictor candidate from N MV predictor candidates as the MV predictor for the current block (step Sg_13).
[0703] Next, the inter-frame predictor 218 decodes the MV difference from the input stream and derives the MV for the current block by adding the difference, which is the decoded MV difference, to the selected MV predictor (step Sg_14).
[0704] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference image (step Sg_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed on each block in all blocks of a slice, the inter-frame prediction of the slice using normal inter-frame mode ends. Similarly, when the processes in steps Sg_11 to Sg_15 are performed on each block in all blocks of an image, the inter-frame prediction of the image using normal inter-frame mode ends. Note that not all blocks included in a slice can undergo the processes in steps Sg_11 to Sg_15, and the inter-frame prediction of the slice using normal inter-frame mode can end when only a portion of a block undergoes the process. This also applies to the images in steps Sg_11 to Sg_15. When the process is performed on only a portion of a block in an image, the inter-frame prediction of the image using normal inter-frame mode can end.
[0705] (MV derivation > Normal merge mode)
[0706] For example, when the information parsed from the stream indicates that a normal merging mode should be applied, the inter-frame predictor 218 derives the MV and uses the MV to perform motion compensation (prediction).
[0707] Figure 85 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using normal merging mode.
[0708] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple decoded blocks surrounding the current block in time or space) (step Sh_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0709] Next, the inter-frame predictor 218 selects an MV candidate from the multiple MV candidates obtained in step Sh_11, thereby deriving the MV for the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains MV selection information included in the stream as prediction parameters, and selects the MV candidate identified by the MV selection information as the MV for the current block.
[0710] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference image (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed for each block in all blocks of a slice, the inter-frame prediction of the slice using the normal merging mode ends. Similarly, when the processes in steps Sh_11 to Sh_13 are performed for each block in all blocks of an image, the inter-frame prediction of the image using the normal merging mode ends. Note that not all blocks included in a slice undergo the processes in steps Sh_11 to Sh_13, and the inter-frame prediction of the slice using the normal merging mode can end when only a portion of a block undergoes the process. This also applies to the images in steps Sh_11 to Sh_13. When the process is performed on only a portion of a block in an image, the inter-frame prediction of the image using the normal merging mode can end.
[0711] (MV Derivation > Radiation Combination Pattern)
[0712] For example, when information parsed from the stream indicates that FRUC mode will be applied, the inter-frame predictor 218 derives the motion MV in FRUC mode and uses the MV to perform motion compensation (prediction). In this case, the motion information is derived on the decoder 200 side, without being signaled from the encoder 100 side. For example, the decoder 200 can derive motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values in the current block.
[0713] Figure 86 This is a flowchart illustrating an example of the inter-frame prediction process performed in decoder 200 via FRUC mode.
[0714] First, the inter-frame predictor 218 generates a list of MVs indicating decoded blocks spatially or temporally adjacent to the current block by referencing MVs as MV candidates (this list is an MV candidate list, and can also be used, for example, as an MV candidate list for normal merge mode) (step Si_11). Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates an evaluation value for each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV for the current block (step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block can be derived using pattern matching in the surrounding area included in the reference image and corresponding to the position of the selected best MV candidate. In other words, pattern matching and estimation using the reference image can be performed in the region surrounding the best MV candidate, and when an MV that produces a better evaluation value exists, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV for the current block. In an embodiment, updating to the MV that produces a better evaluation value may not be performed.
[0715] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference image (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of a slice, the inter-frame prediction of the slice using the FRUC mode ends. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of an image, the inter-frame prediction of the image using the FRUC mode ends. Each sub-block can be processed similarly to the case of each block.
[0716] (MV Derivation > FRUC Pattern)
[0717] For example, when the information parsed from the stream indicates that an affine merge mode should be applied, the inter-frame predictor 218 derives the MV in the affine merge mode and uses the MV to perform motion compensation (prediction).
[0718] Figure 87 This is a flowchart illustrating an example of the inter-frame prediction process performed in decoder 200 using affine merging mode.
[0719] In affine merging mode, firstly, the inter-frame predictor 218 derives the MV for the current block at the corresponding control point (step Sk_11). For example... Figure 46A As shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 46B As shown, the control points are the top-left corner, the top-right corner, and the bottom-left corner of the current block.
[0720] For example, when using Figures 47A to 47C When the MV derivation method is shown in the figure, such as Figure 47A As shown, the inter-frame predictor 218 sequentially examines the decoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), and identifies the first valid block decoded according to the affine pattern. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47B As shown, the inter-frame predictor 218 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the decoded block including block A. In this way, the MV at each control point is derived.
[0721] Note that, as Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, and as follows: Figure 49B As shown, when block A is identified and when block A has three control points, the MV at two control points can be calculated.
[0722] Additionally, when MV selection information is included in the stream as a prediction parameter, the inter-frame predictor 218 can use the MV selection information to derive the MV for the current block at each control point.
[0723] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to compute an affine MV for each of the multiple sub-blocks (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and an encoded reference image to perform motion compensation for the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merging mode for the current block ends. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.
[0724] Note that the MV candidate list described above can be generated in step Sk_11. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can be, for example... Figures 47A to 47C The MV derivation method shown in the figure Figure 48A and Figure 48B The MV derivation method shown in the figure Figure 49A and Figure 49B The MV derivation method shown in the figure, and any combination of other MV derivation methods.
[0725] Note that the MV candidate list can include MV candidates in modes that perform prediction on a sub-block basis, other than affine mode.
[0726] Note that, for example, an MV candidate list can be generated that includes MV candidates in an affine merge pattern using two control points and an affine merge pattern using three control points. Alternatively, an MV candidate list can be generated separately that includes MV candidates in an affine merge pattern using two control points and an MV candidate list that includes MV candidates in an affine merge pattern using three control points. Alternatively, an MV candidate list can be generated that includes MV candidates in one of the affine merge patterns using two control points and an affine merge pattern using three control points.
[0727] (MV Derivation > Affine Inter-Frame Mode)
[0728] For example, when the information parsed from the stream indicates that an affine inter-frame mode will be applied, the inter-frame predictor 218 derives the MV in the affine inter-frame mode and uses the MV to perform motion compensation (prediction).
[0729] Figure 88 This is a flowchart illustrating an example of the process of inter-frame prediction in decoder 200 using affine inter-frame modes.
[0730] In affine inter-frame mode, firstly, the inter-frame predictor 218 derives the MV predictor (v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_11). The control points are the top-left corner, top-right corner, and bottom-left corner of the current block, such as... Figure 46A or Figure 46B As shown in the image.
[0731] Inter-frame predictor 218 obtains MV predictor selection information included in the stream as prediction parameters, and uses the MV identified by the MV predictor selection information to derive the MV predictor at each control point of the current block. For example, when using Figure 48A and Figure 48BWhen the MV derivation method shown in the figure is used, the inter-frame predictor 218 selects... Figure 48A or Figure 48B The MV predictor in the decoded block near the corresponding control point of the current block is identified by the block's MV selection information to deduce the motion vector predictor (v0, v1) or (v0, v1, v2) at the control point of the current block.
[0732] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block to the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV for the current block at each control point is derived.
[0733] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to compute an affine MV for each of the multiple sub-blocks (step Sk_13). The inter-frame predictor 218 then uses these affine MVs and the encoded reference image to perform motion compensation for the sub-blocks (step Sk_14). When the processes in steps Sk_13 and Sk_14 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merging mode for the current block ends. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.
[0734] Note that the MV candidate list described above can be generated in step Sj_11 as in step Sk_11.
[0735] (MV Derivation > Triangle Pattern)
[0736] For example, when the information parsed from the stream indicates that a triangle pattern will be applied, the inter-frame predictor 218 derives the MV in the triangle pattern and uses the MV to perform motion compensation (prediction).
[0737] Figure 89 This is a flowchart illustrating an example of the process of inter-frame prediction using a triangle pattern in decoder 200.
[0738] In triangle mode, firstly, the inter-frame predictor 218 splits the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as prediction parameters; this partition information is related to the splitting. The inter-frame predictor 218 can then split the current block into a first partition and a second partition based on the partition information.
[0739] Next, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information (e.g., the MVs of multiple decoded blocks surrounding the current block in time or space) (step Sx_12). In other words, the inter-frame predictor 218 generates a list of MV candidates.
[0740] Inter-frame predictor 218 then selects the MV candidate for the first partition and the MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this point, inter-frame predictor 218 can obtain MV selection information from the stream to identify each selected MV candidate as prediction parameters. Inter-frame predictor 218 can then select the first MV and the second MV based on the MV selection information.
[0741] Next, the inter-frame predictor 218 performs motion compensation using the selected first MV and the decoded reference image to generate a first predicted image (step Sx_14). Similarly, the inter-frame predictor 218 performs motion compensation using the selected second MV and the decoded reference image to generate a second predicted image (step Sx_15).
[0742] Finally, the inter-frame predictor 218 generates a prediction image for the current block by performing a weighted summation of the first and second prediction images (step Sx_16).
[0743] (MV estimate > DMVR)
[0744] For example, information from stream parsing indicates that DMVR will be applied, and inter-frame predictor 218 uses DMVR to perform motion estimation.
[0745] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed by DMVR in decoder 200.
[0746] Inter-frame predictor 218 derives the MV for the current block based on the merging pattern (step Sl_11). Next, inter-frame predictor 218 derives the final MV for the current block by searching the region around the reference image indicated by the MV derived in Sl_11 (step Sl_12). In other words, in this case, the MV of the current block is determined according to the DMVR.
[0747] Figure 91 This is a flowchart illustrating an example of the motion estimation process performed via DMVR in decoder 200, and... Figure 58B same.
[0748] First of all, Figure 58AIn step 1 shown, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also known as the start point) and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost is minimized at each of the search positions other than the start point. Here, when it is determined that the cost is minimized at one of the search positions other than the start point, the inter-frame predictor 218 changes its target to the search position where the minimum cost is obtained and performs... Figure 58A The process in step 2 is shown in the diagram. When the cost at the starting point is minimized, the inter-frame predictor 218 skips... Figure 58A The process in step 2 is shown in the figure, and the process in step 3 is executed.
[0749] exist Figure 58A In step 2 shown, the inter-frame predictor 218 performs a search similar to that in step 1, taking the search position after the target change as the new starting point based on the result of the process in step 1. Then, the inter-frame predictor 218 determines whether the cost is minimized at each search position other than the starting point. Here, when it is determined that the cost is minimized at one of the search positions other than the starting point, the inter-frame predictor 218 performs the process in step 4. When the cost is minimized at the starting point, the inter-frame predictor 218 performs the process in step 3.
[0750] In step 4, the inter-frame predictor 218 considers the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0751] exist Figure 58A In step 3 shown, the inter-frame predictor 218 determines the pixel position with the minimum cost and sub-pixel accuracy based on the cost of four points located above, below, to the left and to the right of the starting point in step 1 or step 2, and regards the pixel position as the final search position.
[0752] The pixel position with sub-pixel accuracy is determined by performing a weighted summation of each of the four vectors ((0,1), (0,-1), (-1,0), and (1,0)) for the top, bottom, left, and right sides using the cost corresponding to one of the four search positions as a weight. The inter-frame predictor 218 then determines the vector difference as the difference between the position indicated by the initial MV and the final search position.
[0753] (Motion compensation > BIO / OBMC / LIC)
[0754] For example, when the information parsed from the stream indicates that correction of the predicted image should be performed, the inter-frame predictor 218 corrects the predicted image based on a mode used for correction when generating the predicted image. This mode is, for example, one of BIO, OBMC, and LIC described above.
[0755] Figure 92 This is a flowchart illustrating an example of the process of generating a predicted image in decoder 200.
[0756] Inter-frame predictor 218 generates a predicted image (step Sm_11) and corrects the predicted image according to any of the modes described above (step Sm_12).
[0757] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in decoder 200.
[0758] Inter-frame predictor 218 derives the MV for the current block (step Sn_11). Next, inter-frame predictor 218 uses the MV to generate a predicted image (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, inter-frame predictor 218 obtains prediction parameters included in the stream and determines whether to perform a correction process based on these prediction parameters. For example, the prediction parameters are flags indicating whether to apply one or more of the modes described above. Here, when it is determined that a correction process should be performed ("Yes" in step Sn_13), inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). Note that in LIC, illumination and chroma can be corrected in step Sn_14. When it is determined that a correction process should not be performed ("No" in step Sn_13), inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).
[0759] (Motion compensation > OBMC)
[0760] For example, when the information parsed from the stream indicates that OBMC should be performed, the inter-frame predictor 218 corrects the predicted image according to OBMC when generating the predicted image.
[0761] Figure 94 This is a flowchart illustrating an example of the process of correcting the predicted image via OBMC in decoder 200. Note that... Figure 94 The flowchart in the document indicates the use of Figure 62 The current image and reference image shown in the diagram represent the correction process for the predicted image.
[0762] First, such as Figure 62 As shown, the inter-frame predictor 218 uses the MV assigned to the current block to obtain the predicted image (Pred) by performing normal motion compensation.
[0763] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) already derived for the encoded block adjacent to the left of the current block to the current block (reusing the motion vector used for the current block). Then, the inter-frame predictor 218 performs a first correction to the predicted image by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.
[0764] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) already derived for the coded block adjacent to the current block (reusing the MV for the current block) to the current block. Then, the inter-frame predictor 218 performs a second correction on the predicted image by overlapping the predicted image Pred_U with the predicted images (e.g., Pred and Pred_L) that have already undergone the first correction. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained through the second correction is one in which the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.
[0765] (Motion compensation > BIO)
[0766] For example, when the information parsed from the stream indicates that a BIO should be performed, the inter-frame predictor 218 corrects the predicted image based on the BIO when generating the predicted image.
[0767] Figure 95 This is a flowchart illustrating an example of the process of BIO correcting the predicted image in decoder 200.
[0768] like Figure 63 As shown, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference images (Ref0, Ref1) that are different from the image (Cur Pic) including the current block. Then, the inter-frame predictor 218 uses the two motion vectors (M0, M1) to derive the predicted image for the current block (step Sy_11). Note that motion vector M0 is the motion vector (MV) corresponding to the reference image Ref0. x0 MV y0 Furthermore, motion vector M1 is the motion vector (MV) corresponding to the reference image Ref1. x1 MV y1 ).
[0769] Next, the inter-frame predictor 218 uses the motion vector M0 and the reference image L0 to derive the interpolated image I for the current block. 0Additionally, the inter-frame predictor 218 uses motion vector M1 and reference image L1 to derive the interpolated image I for the current block. 1 (Step Sy_12). Here, the interpolated image I 0 It is the image included in the reference image Ref0 and derived for the current block, and the interpolated image I 1 It is the image included in reference image Ref1 and derived for the current block. Interpolated image I 0 and interpolated image I 1 Each element in the interpolated image can be the same size as the current block. Alternatively, the interpolated image I... 0 and interpolated image I 1 Each of the elements can be an image larger than the current block. Furthermore, the interpolated image I... 0 and interpolated image I 1 This can include a predicted image obtained by using motion vectors (M0, M1) and a reference image (L0, L1) and applying a motion compensation filter.
[0770] Additionally, the inter-frame predictor 218 determines the interpolated image I based on... 0 and interpolated image I 1 Derive the gradient image of the current block (Ix) 0 , Ix 1 ,Iy 0 ,Iy 1 (Step Sy_13). Note that the gradient image in the horizontal direction is (Ix 0 , Ix 1 ), and the gradient image in the vertical direction is (Iy 0 ,Iy 1 The inter-frame predictor 218 can derive a gradient image, for example, by applying a gradient filter to the interpolated image. The gradient image can be an image indicating either the spatial change of pixel values along the horizontal direction or the spatial change of pixel values along the vertical direction.
[0771] Next, the inter-frame predictor 218 uses the interpolated image (I 0 I 1 ) and gradient image (Ix 0 , Ix 1 ,Iy 0 ,Iy 1 For each sub-block of the current block, derive the optical flow (vx, vy) as a velocity vector (step Sy_14). As an example, a sub-block can be a 4×4 pixel sub-CU.
[0772] Next, the inter-frame predictor 218 uses optical flow (vx, vy) to correct the predicted image for the current block. For example, the inter-frame predictor 218 uses optical flow (vx, vy) to derive correction values for the values of the pixels included in the current block (step Sy_15). The inter-frame predictor 218 can then use the correction values to correct the predicted image for the current block (step Sy_16). Note that the correction values can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis.
[0773] Note that the BIO process flow is not limited to Figure 95 The publicly disclosed process. It is possible to execute only... Figure 95 It is part of the publicly disclosed process, or different processes can be added or different processes can be used as alternatives, or processes can be executed in different processing orders, etc.
[0774] (Motion compensation > LIC)
[0775] For example, when the information parsed from the stream indicates that LIC should be performed, the inter-frame predictor 218 corrects the predicted image according to the LIC when generating the predicted image.
[0776] Figure 96 This is a flowchart illustrating an example of the process by which the LIC corrects the predicted image in the decoder 200.
[0777] First, the inter-frame predictor 218 uses MV to obtain a reference image corresponding to the current block from the decoded reference image (step Sz_11).
[0778] Next, the inter-frame predictor 218 extracts information indicating how the illumination value changes between the current image and the reference image for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region), as well as the luminance pixel values at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 218 uses the information indicating how the luminance value changes to calculate the illumination correction parameters (step Sz_13).
[0779] The inter-frame predictor 218 generates a predicted image for the current block by performing an illumination correction process in which illumination correction parameters are applied to a reference image in a reference picture specified by the MV (step Sz_14). In other words, the predicted image is corrected based on the illumination correction parameters, which serves as a reference image in the reference picture specified by the MV. In this correction, either illumination or chroma can be corrected.
[0780] (Predictive Controller)
[0781] Prediction controller 220 selects an intra-frame predicted image or an inter-frame predicted image and outputs the selected image to adder 208. In general, the configuration, function, and procedure of prediction controller 220, intra-frame predictor 216, and inter-frame predictor 218 on the decoder 200 side can correspond to the configuration, function, and procedure of prediction controller 128, intra-frame predictor 124, and inter-frame predictor 126 on the encoder 100 side.
[0782] (Decoding using predicted chroma samples)
[0783] In the first aspect, it is determined whether a luminance sample can be used to predict a chroma sample block for the current block, wherein the predicted chroma sample is used to decode the block. For example, an embodiment may employ a process of using the decoding result of the illuminance signal in a decoding or encoding method to determine whether to enable a tool such as CCLM to predict the chroma signal.
[0784] Figure 97 This is a flowchart illustrating an example of a process 1000 for decoding a block using predicted chroma samples, which can be, for example, by... Figure 7 encoder 100 or Figure 67 The decoder 200 is executed. For convenience, refer to... Figure 67 Decoder 200 to describe Figure 97 .
[0785] At S1001, decoder 200 determines whether the current chroma block is inside an M×N non-overlapping region aligned with the M×N grid of the chroma sample. Figure 99 and Figure 100 This is a conceptual diagram illustrating an example of determining whether the current chroma block is within an M×N non-overlapping region aligned with an M×N grid of chroma samples. In some formats (e.g., YUV420), a 16×16 pixel region of chroma corresponds to a 32×32 pixel region of illuminance. For example... Figure 99 and Figure 100 As shown, chroma blocks within a 32×32 illuminance region aligned with a 16×16 chroma grid are identified as being within an M×N non-overlapping region aligned with an M×N grid of chroma samples. Chroma blocks not within the 32×32 illuminance region are not identified as being within an M×N non-overlapping region aligned with an M×N grid of chroma samples. Even if the current chroma block crosses the boundary of the corresponding luminance block, if it is included in the same VPDU, the chroma sample can be obtained using the luminance sample. For example, Figure 99 The chroma block A uses the corresponding sample in the luminance block B to predict the sample in the chroma block A-1, and uses the corresponding sample in the luminance block C to predict the sample in the chroma block A-2.
[0786] like Figure 100As shown, the chromaticity sample of the chromaticity block can be predicted using the luminance sample for the chromaticity block, since the chromaticity block is included in a grid (as shown, a 16×16 grid), and the co-located luminance block is also within the co-located 32×32 area.
[0787] In some embodiments, for example, by default, when other conditions, such as those discussed below with reference to S1002, are met, luminance samples may not be used to predict chroma sample blocks that are not determined to be inside an M×N non-overlapping region aligned with an M×N grid of chroma samples, but luminance samples may be used to predict chroma sample blocks that are determined to be inside an M×N non-overlapping region aligned with an M×N grid of chroma samples.
[0788] like Figure 97 As shown, when it is not determined at S1001 that the current chroma block is within an M×N non-overlapping region aligned with the M×N grid of the chroma samples, process 1000 proceeds from S1001 to S1004. In S1004, decoder 200 predicts the chroma sample block without using luminance samples. Process 1000 proceeds from S1004 to S1005. In S1005, decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1001 that the current chroma block is within an M×N non-overlapping region, process 1000 proceeds from S1001 to S1002.
[0789] At S1002, decoder 200 determines whether to split the current luminance VPDU into smaller blocks. A VPDU is a unit that is processed in parallel during encoding or decoding, for example, a 64×64 block. The size of a VPDU can be determined by standards, or it can be encoded in a stream.
[0790] There are various ways to determine whether to split the current luminance VPDU into smaller blocks, and the following references are available. Figure 102 and Figure 103 Let's discuss some examples in more detail.
[0791] If it is not determined at S1002 that the current luminance VPDU should be split into smaller blocks, process 1000 proceeds from S1002 to S1004, in which decoder 200 predicts chroma sample blocks without using luminance samples. Process 1000 proceeds from S1004 to S1005, in which decoder 200 decodes the blocks using the predicted chroma samples. If it is determined at S1002 that the current luminance VPDU should be split into smaller blocks, process 1000 proceeds from S1002 to S1003, in which decoder 200 uses luminance samples to predict chroma sample blocks. Process 1000 proceeds from S1003 to S1005, in which decoder 200 decodes the blocks using the predicted chroma samples. In some embodiments, additional considerations may be taken into account to determine whether to use luminance samples to decode chroma sample blocks, for example, as referenced below. Figures 104 to 110 The subject of discussion.
[0792] Figure 98 This is a flowchart illustrating another example of the process 2000 for decoding blocks using predicted chroma samples, a process that can be, for example, by... Figure 7 encoder 100 or Figure 67 The decoder 200 is used for execution. For convenience, refer to... Figure 67 Decoder 200 to describe Figure 98 .
[0793] At S2001, decoder 200 determines whether to split the first VPDU and the second VPDU into smaller blocks. This can be determined in various ways, and is discussed below. Figure 102 and Figure 103 Let's discuss some examples in more detail.
[0794] When it is determined at S2001 that the first VPDU should not be split into smaller blocks and the second VPDU should be split into smaller blocks, process 2000 proceeds from S2001 to S2002. In S2002, decoder 200 predicts chroma sample blocks without using luminance samples. Process 2000 proceeds from S2002 to S2004. In S2004, decoder 200 decodes the blocks using the predicted chroma samples.
[0795] If at S2001 it is not determined whether the first luminance VPDU should be split into smaller blocks and whether the second VPDU should be split into smaller blocks, process 2000 proceeds from S2001 to S2003, in which decoder 200 uses luminance samples to predict chroma sample blocks. Process 2000 proceeds from S2003 to S2004, in which decoder 200 uses the predicted chroma samples to decode the blocks. In some embodiments, additional considerations may be taken into account to determine whether to use luminance samples to decode chroma sample blocks, for example, as referenced below. Figures 104 to 110 The subject of discussion.
[0796] Figure 101 This is a conceptual diagram used to illustrate a VPDU. A VPDU is non-overlapping and represents the buffer size of a pipeline stage. Figure 101 The left side (labeled a) shows an example of a 128×128 CTU with four 64×64 VPDUs. Figure 101 The right side (labeled b) shows an example of a 128×128 CTU with 16 32×32 VPDUs. For example, if the VPDU is 64×64, then both M and N are set to 16. When the VPDU is further subdivided, the size of the subdivided CU becomes 2M×2N (32×32) or smaller. In the YUV420 format, a 16×16 chroma region corresponds to a 32×32 illuminance region, so pixels in the 16×16 chroma grid can be predicted based on the pixels in the corresponding 32×32 grid in the illuminance. Therefore, when the decoding of the 32×32 illuminance region is complete, the decoding process for predicting chroma difference can begin in the 16×16 chroma difference region. In the case of the YUV444 format, the illuminance M×N region corresponds to the chroma difference M×N region. If 2M×2N is half the size of the VPDU in both the horizontal and vertical directions, then in Figure 97 In step S1002 or Figure 98 In step S2001, it can be determined whether the VPDU is further divided into one or more layers, but this is done in the case of 1 / 4 of the VPDU. An embodiment can determine whether the divided CU becomes 2M×2N or smaller, for example, whether the CU is further divided into two or more layers.
[0797] Figure 102 This is a conceptual diagram illustrating an example of determining whether a current VPDU can be used to predict a chromaticity sample block based on whether the luminance VPDU is split into blocks. The left side shows the luminance CTU, and the right side shows the corresponding chromaticity CTU. As shown, luminance VPDU0 will be split into blocks, while luminance VPDU1 will not. Therefore, refer to... Figure 97In process 1000, luminance samples can be used to predict the chromaticity samples of VPDU0, and luminance samples can be used to predict the chromaticity samples of VPDU1 without using luminance samples.
[0798] Figure 103 This is a conceptual diagram illustrating two example ways to determine whether a luminance VPDU will be split into smaller blocks. Figure 103 In the first example shown on the left (labeled a), whether the luminance VPDU should be split can be determined based on the split flag associated with the luminance VPDU. As shown, when the split flag has a value of 1, the VPDU should be split (and, see reference...). Figure 97 The process 1000 uses luminance samples to predict the chrominance samples of the blocks. When the split flag has a value of 0, the VPDU is not split (and, see reference...). Figure 97 The process 1000 does not use luminance samples to predict the chrominance samples of the block. Other splitting flag values can be used to determine whether the luminance VPDU has been split.
[0799] exist Figure 103 In the second example shown on the right (labeled b), whether a luminance VPDU should be split can be determined based on the quadtree splitting depth of the luminance block of the VPDU. As shown, the quadtree splitting depth of the luminance block of VPDU0 is greater than 1, therefore, referring to... Figure 97 In process 1000, when decoding the VPDU0 block, luminance samples can be used to predict chrominance samples. In contrast, the quadtree split depth of the VPDU1 block is less than or equal to 1; therefore, referencing... Figure 97 In process 1000, when decoding the VPDU0 block, luminance samples may not be used to predict chrominance samples. Other split depth values can be used to determine whether the luminance VPDU has been split.
[0800] Figure 104 This is a conceptual diagram used to illustrate additional considerations that can be taken into account when determining whether to use luminance samples to predict the chrominance samples of a block. As shown, whether the current block size is equal to or less than a threshold block size can be used as an additional consideration in determining whether to use luminance samples to predict the chrominance samples of a block.
[0801] The threshold block size can be a default block size, a block size notified by a signal, or a predetermined block size, and can be either a luma or chroma block size. For example, if the threshold block size is a 16×16 luma block size, then the luma block size of VPDU0 is greater than 16×16, therefore it can be determined that luma samples are not used to determine the chroma samples for the block. Figure 97 At S1002 or Figure 98 At S2001, the threshold block size is used to determine whether to split the luminance VPDU into smaller blocks.
[0802] Figure 97 The process of 1000 aspects and Figure 98 The process 2000 can be modified in various ways. For example, process 1000 or 2000 can be modified to perform more actions than shown, can be modified to perform fewer actions than shown, can be modified to perform actions in various orders, can be modified to combine or split actions, etc. For example, before S1001 or S1002, process 1000 can be modified based on other considerations (e.g., reference). Figure 103 The size of the current block (discussed) determines whether to use luminance samples to predict chrominance samples for the block. In another example, process 1000 can be modified to omit S1001. In another example, Figure 98 An embodiment of process 2000 can be modified to perform step S1001 before performing step S2001. In another example, S2001 may determine whether the first VPDU and the second VPDU are split into smaller blocks.
[0803] Figure 105 This is a conceptual diagram illustrating an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (e.g.) Figure 105 The example combination shown is based on whether both the luma VPDU and the corresponding chroma VPDU have a quadtree split depth greater than or equal to 2. Luma VPDU0 has a quadtree depth greater than or equal to 2, and chroma VPDU0 also has a quadtree split depth greater than or equal to 2, so a luma sample can be used to predict a chroma sample for chroma VPDU0. However, luma VPDU1 has a quadtree split depth not greater than or equal to 2, therefore one of the conditions is not met, and a chroma sample for chroma VPDU1 will be predicted without using a luma sample.
[0804] Figure 106 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (See diagram for example.) Figure 106The example combination of conditions shown is: (i) whether the quadtree split depth of the luminance VPDU is greater than or equal to 2; (ii) whether the quadtree split depth of the corresponding chrominance VPDU is equal to 1; and (iii) whether the chrominance split threshold condition of 32×32 is met (e.g., when the chrominance size is 32×32, the block is not split). Luminance VPDU0 has a quadtree depth greater than or equal to 2, satisfying condition (i); chrominance VPDU0 has a quadtree split depth equal to 1, satisfying condition (ii), and the chrominance VPDU is not split into blocks smaller than 32×32. Therefore, all three conditions are met, and the chrominance sample for VPDU0 can be predicted using the luminance sample. However, the block size of chrominance VPDU1 is smaller than the 32×32 threshold, therefore condition (iii) is not met, and the chrominance sample for VPDU1 will be predicted without using the luminance sample.
[0805] Figure 107 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (See diagram for example.) Figure 107 The example combination of conditions shown is: (i) whether the luminance VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chrominance VPDU quadtree split depth is equal to 1; and (iii) whether the chrominance split threshold condition of 32×32 is met (e.g., when the chrominance size is 32×32, the block is not split). Figure 107 In the code, qtDepthC indicates the chroma quadtree splitting depth, and Figure 107 `mtDepthC` indicates the chroma multi-way tree split depth. A quadtree split can be followed by another quadtree split or a multi-way tree split (binary or ternary split). To specify that the chroma quadtree split terminates at depth 1, add the condition `chromaSplit32×32 == CU_DONT_SPLIT` (see reference). Figure 106 Condition iii) under discussion implies no further splitting at the chroma 32×32 level. Assuming the luminance quadtree split depth qtDepthl is greater than or equal to 2, only chroma VPDU0 satisfies all three conditions, and chroma samples for VPDU0 can be predicted using luminance samples. Chroma VPDU1 has a chroma quadtree split depth of 2, and its blocks are split into smaller than 32×32 blocks; therefore, chroma samples for VPDU1 will be predicted without using luminance samples. Chroma VPDU2 has a chroma quadtree split depth of 1, but its blocks are split into smaller than 32×32 blocks; therefore, chroma samples for VPDU2 will be predicted without using luminance samples. Chroma VPDU3 has a chroma quadtree split depth of 1, but its blocks are split into smaller than 32×32 blocks; therefore, chroma samples for VPDU3 will be predicted without using luminance samples.
[0806] Figure 108 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chroma samples for a block. In the example, the combination of conditions is: (i) whether the luminance VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDU quadtree split depth is equal to 1; and (iii) whether there is no vertical or horizontal ternary split after a horizontal chroma split of size 32×32. For VPDU0, the conditions are met; the VPDU is horizontally split into two 16×32 blocks, and these blocks are not further split using horizontal or vertical ternary splits. Therefore, luminance samples can be used to predict chroma samples in all blocks of VPDU0. For VPDU1, the lower 16×32 block meets the conditions, no further ternary splits are performed, and luminance samples can be used to predict the chroma samples of the lower 16×32 block. The upper 16×32 block of VPDU1 does not meet the conditions because there is a further vertical ternary split, and the chroma samples of the upper 16×32 block of VPDU1 will be predicted without using luminance samples.
[0807] Figure 109 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (See diagram for example.) Figure 109 The example combination of conditions shown is: (i) whether the luminance VPDU quadtree split depth is equal to 1; (ii) whether a luminance split threshold of 64×64 is met (e.g., when the luminance size is 64×64, the block is not split); (iii) whether the corresponding chrominance VPDU quadtree split depth is equal to 1; and (iv) whether a chrominance split threshold of 32×32 is met (e.g., when the chrominance size is 32×32, the block is not split). VPDU0 meets all four conditions, and chrominance samples can be used to predict CVD samples for VPDU0. The chrominance quadtree split depth for chrominance VPDU1 is 2, and the block is split into blocks smaller than 32×32, so chrominance samples for VPDU1 will be predicted without using luminance samples.
[0808] Figure 110 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (See diagram for example.) Figure 110As shown, if any one of the conditions is true, the luminance sample can be used to predict the chrominance sample in the block. Example combinations of conditions are: (i) whether the luminance VPDU quadtree split depth is greater than or equal to 2, and whether the chrominance VPDU quadtree split depth is greater than or equal to 2; (ii) whether the luminance VPDU quadtree split depth is equal to 1, satisfying the 64×64 luminance splitting threshold condition (e.g., when the luminance size is 64×64, the block is not split), and the corresponding chrominance VPDU quadtree split depth is equal to 1, satisfying the 32×32 chrominance splitting threshold condition (e.g., when the chrominance size is 32×32, the block is not split); (iii) Whether the luminance VPDU quadtree splitting depth is greater than or equal to 2, and the corresponding chrominance VPDU quadtree splitting depth is equal to 1, and satisfies the 32×32 chrominance splitting threshold condition (e.g., when the chrominance size is 32×32, the block is not split); and (iv) whether the luminance VPDU quadtree splitting depth is greater than or equal to 2, and the corresponding chrominance VPDU quadtree splitting depth is equal to 1, the chrominance splitting of the 32×32 block is horizontal, and chrominance blocks smaller than 32×32 are either not split or vertically split. The block of VPDU1 violates all four conditions, and therefore the chrominance sample of VPDU1 will be predicted without using the luminance sample. Considering the scan order, a... Figure 110 The example conditions are used to limit the chroma prediction latency (based on the luminance samples) within a 32×32 sample. For example, in VPDU1, chroma block 0 must wait for luminance block 0 reconstruction before prediction. Chroma block 1 must wait for both luminance blocks 0 and 1 reconstruction before prediction. To avoid this latency, chroma prediction can be performed without using luminance samples.
[0809] The blocks described in each aspect can be replaced with rectangular or non-rectangular shaped partitions. Figure 111 Examples of non-rectangular shape partitions are shown, such as triangular, L-shaped, pentagonal, hexagonal, and polygonal shape partitions. Other non-rectangular shape partitions can be used, and various combinations of shapes can be employed. The term "partition" described in each aspect can be replaced by the term "prediction unit." The term "partition" described in each aspect can also be replaced by the term "sub-prediction unit." The term "partition" described in each aspect can also be replaced by the term "decoding unit."
[0810] Other conditions may be used. For example, in one embodiment, when a decoding mode for predicting color difference based on illuminance (e.g., CCLM) is enabled, the first partition of the VPDU may always be a quaternary partition. In another embodiment, CCLM may be disabled in all VPDUs of the CTU when a defined number of quaternary partitions are not applied to VPDUs in at least one VPDU in the CTU (e.g., the head VPDU of the CTU in the scan sequence).
[0811] CCLM can be defined as an intra-prediction mode that uses mode information such as `intra_chroma_pred_mode`. The index number indicating the intra-prediction mode and each mode can be associated one-to-one in a table; however, when CCLM is disabled, entries for the table for CCLM are unnecessary, so the index number is encoded, thereby facilitating a reduction in the number of bits used to encode the signal. In an embodiment, the table indicating the intra-prediction mode can be switched depending on whether CCLM is valid or invalid. For example, referring to the partition flag information of the quaternion indicating illumination, if the illumination is not partitioned into a defined size or smaller in the VPDU, it can be determined that CCLM is invalid, and the table corresponding to the invalid CCLM case can be used. Otherwise, the corresponding table used when CCLM is available can be used. In an embodiment, a table including entries for CCLM can be used without switching the table, but when CCLM is invalid, entries for CCLM can be ignored.
[0812] refer t...
Claims
1. An encoder, comprising: Circuit; as well as A memory coupled to the circuit; The circuit performs the following operations during operation: Determine whether the first virtual pipelined decoding unit (VPDU) is split into smaller blocks, and whether the second virtual pipelined decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and determining that the second virtual pipeline decoding unit has been split into smaller blocks, predict chroma sample blocks without using luminance samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, luminance samples are used to predict chrominance sample blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and that the second virtual pipeline decoding unit has not been split into smaller blocks, luminance samples are used to predict chrominance sample blocks; and The blocks are encoded using predicted chroma samples. The circuit determines whether to split the virtual pipeline decoding unit into smaller blocks based on a split flag, block split depth, or threshold block size during operation.
2. A decoder, comprising: Circuit; A memory coupled to the circuit; The circuit performs the following operations during operation: Determine whether the first virtual pipelined decoding unit (VPDU) is split into smaller blocks, and whether the second virtual pipelined decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and determining that the second virtual pipeline decoding unit has been split into smaller blocks, predict chroma sample blocks without using luminance samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, luminance samples are used to predict chrominance sample blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and that the second virtual pipeline decoding unit has not been split into smaller blocks, luminance samples are used to predict chrominance sample blocks; and The blocks are decoded using predicted chroma samples. The circuit determines whether to split the virtual pipeline decoding unit into smaller blocks based on a split flag, block split depth, or threshold block size during operation.
3. A non-transitory medium that stores a bit stream and can be read by a computer. The bitstream includes syntax for instructing the computer to perform a decoding process. The decoding process includes: Determine whether the first virtual pipelined decoding unit (VPDU) is split into smaller blocks, and whether the second virtual pipelined decoding unit is split into smaller blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and determining that the second virtual pipeline decoding unit has been split into smaller blocks, predict chroma sample blocks without using luminance samples; In response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, luminance samples are used to predict chrominance sample blocks; In response to determining that the first virtual pipeline decoding unit has not been split into smaller blocks and that the second virtual pipeline decoding unit has not been split into smaller blocks, luminance samples are used to predict chrominance sample blocks; as well as The blocks are decoded using predicted chroma samples. The determination of whether to split the virtual pipeline decoding unit into smaller blocks is based on a split flag, block split depth, or threshold block size.