Systems and methods for video coding
By determining whether to use luminance samples to predict chrominance sample blocks based on the splitting of virtual pipeline decoding units during video decoding, the problems of low chrominance sample encoding efficiency and large circuit size in existing technologies are solved, achieving a more efficient encoding and decoding process.
Patent Information
- Application Number
- CN202511243177.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-18
- Publication Date
- 2025-11-07
AI Technical Summary
Existing video decoding technologies are inefficient when processing chroma samples and have large circuit sizes, making it difficult to meet the demands of ever-increasing digital video data volumes.
By using predicted chroma sample blocks in video decoding, the decision to use luminance samples for prediction is made based on whether the virtual pipeline decoding unit is split into smaller blocks, thus optimizing the encoding and decoding process.
It improves encoding efficiency, enhances image quality, reduces the utilization of processing resources and circuit size, and increases encoding/decoding speed.
Smart Images

Figure CN120915944A_ABST
Abstract
Description
[0001] This application is a divisional application of the same named patent application filed on June 18, 2020 under application number 202080044292.2. TECHNICAL FIELD
[0002] The present disclosure relates to video coding, and in particular, to video encoding and decoding systems, components, and methods in video coding and decoding, e.g., for performing encoding of a block using predicted chroma samples. BACKGROUND
[0003] As video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding), there is an ongoing need to provide improvements and optimizations to video coding technology to keep up with the ever-increasing amount of digital video data in various applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in performing encoding of a block using predicted chroma samples. SUMMARY
[0004] In one aspect, an encoder includes circuitry; and a memory coupled to the circuitry. The circuitry determines whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks, and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, chroma sample blocks are predicted without using luma samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, chroma sample blocks are predicted using luma samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, chroma sample blocks are predicted using luma samples. The block is encoded using the predicted chroma samples.
[0005] In one aspect, an encoder includes a block splitter that, in operation, splits a first image into a plurality of blocks; an intra predictor that, in operation, predicts a block included in the first image using a reference block included in the first image; an inter predictor that, in operation, predicts the block included in the first image using a reference block included in a second image different from the first image; a loop filter that, in operation, filters the block included in the first image; a transformer that, in operation, transforms a prediction error between an original signal and a predicted signal generated by the intra predictor or the inter predictor to generate transform coefficients; a quantizer that, in operation, quantizes the transform coefficients to generate quantized coefficients; and an entropy encoder that, in operation, variable encodes the quantized coefficients to generate an encoded bitstream including the encoded quantized coefficients and control information. The prediction block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chrominance samples without using luminance samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the block of chrominance samples using luminance samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting the block of chrominance samples using luminance samples. The block is encoded using the predicted chrominance samples.
[0006] In one aspect, a decoder includes circuitry; and a memory coupled to the circuitry. The circuitry determines whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chrominance samples without using luminance samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the block of chrominance samples using luminance samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting the block of chrominance samples using luminance samples. The block is decoded using the predicted chrominance samples.
[0007] In one aspect, a decoding device includes a decoder that, in operation, decodes an encoded bitstream to output quantized coefficients, an inverse quantizer that, in operation, inverse quantizes the quantized coefficients to output transform coefficients, an inverse transformer that, in operation, inverse transforms the transform coefficients to output prediction errors, an intra predictor that, in operation, predicts a block included in a first picture using a reference block included in the first picture, an inter predictor that, in operation, predicts the block included in the first picture using a reference block included in a second picture different from the first picture, a loop filter that, in operation, filters the block included in the first picture, and an output that, in operation, outputs a picture including the first picture. The prediction block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chrominance samples without using luminance samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the block of chrominance samples using luminance samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting the block of chrominance samples using luminance samples. Decoding the block using the predicted chrominance samples.
[0008] In one aspect, a decoding device includes a decoder that, in operation, decodes an encoded bitstream to output quantized coefficients, an inverse quantizer that, in operation, inverse quantizes the quantized coefficients to output transform coefficients, an inverse transformer that, in operation, inverse transforms the transform coefficients to output prediction errors, an intra predictor that, in operation, predicts a block included in a first picture using a reference block included in the first picture, an inter predictor that, in operation, predicts the block included in the first picture using a reference block included in a second picture different from the first picture, a loop filter that, in operation, filters the block included in the first picture, and an output that, in operation, outputs a picture including the first picture. The prediction block includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second VPDU is split into smaller blocks. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting a block of chrominance samples without using luminance samples. In response to determining that the first VPDU is split into smaller blocks and determining that the second VPDU is split into smaller blocks, predicting the block of chrominance samples using luminance samples. In response to determining that the first VPDU is not split into smaller blocks and determining that the second VPDU is not split into smaller blocks, predicting the block of chrominance samples using luminance samples. Decoding the block using the predicted chrominance samples.
[0009] In one aspect, a decoding method includes determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller chunks and whether a second VPDU is split into smaller chunks. In response to determining that the first VPDU is not split into smaller chunks and determining that the second VPDU is split into smaller chunks, predicting a block of chroma samples without using luma samples. In response to determining that the first VPDU is split into smaller chunks and determining that the second VPDU is split into smaller chunks, predicting the block of chroma samples using luma samples. In response to determining that the first VPDU is not split into smaller chunks and determining that the second VPDU is not split into smaller chunks, predicting the block of chroma samples using luma samples. Decoding the block using the predicted chroma samples.
[0010] In video coding techniques, it is desirable to propose new methods in order to improve coding efficiency, enhance image quality, and reduce circuit size. Some implementations of embodiments of the present disclosure, including constituent elements of embodiments of the present disclosure considered individually or in various combinations, can facilitate one or more of the following: improving coding efficiency; enhancing image quality; reducing utilization of processing resources associated with encoding / decoding; reducing circuit size; improving processing speed of encoding / decoding; etc.
[0011] Additionally, some implementations of embodiments of the present disclosure, including constituent elements of embodiments of the present disclosure considered individually or in various combinations, can facilitate appropriate selection of one or more elements (e.g., filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations) in encoding and decoding. Note that the present disclosure includes disclosure of configurations and methods that can provide advantages in addition to those described above. Examples of such configurations and methods include configurations or methods for improving coding efficiency while reducing increases in processing resource usage.
[0012] Additional benefits and advantages of the disclosed embodiments will become apparent from the description and drawings. Benefits and / or advantages can be realized independently and / or in combination with one or more of the embodiments and features described in the specification and drawings.
[0013] It should be noted that generic or specific embodiments can be implemented as a system, method, integrated circuit, computer program, storage medium or any selective combination thereof. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is a schematic diagram showing one example of a functional configuration of a transmission system according to an embodiment.
[0015] Figure 2 is a conceptual diagram for showing one example of a hierarchical structure of data in a stream.
[0016] Figure 3 FIG. 1 is a conceptual diagram for illustrating one example of a slice configuration.
[0017] Figure 4 FIG. 2 is a conceptual diagram for illustrating one example of a tile configuration.
[0018] Figure 5 FIG. 3 is a conceptual diagram for illustrating one example of an encoding structure in scalable coding.
[0019] Figure 6 FIG. 4 is a conceptual diagram for illustrating one example of an encoding structure in scalable coding.
[0020] Figure 7 FIG. 5 is a block diagram illustrating a functional configuration of an encoder according to an embodiment.
[0021] Figure 8 FIG. 6 is a functional block diagram illustrating an installation example of an encoder.
[0022] Figure 9 FIG. 7 is a flowchart indicating one example of an overall encoding process performed by an encoder.
[0023] Figure 10 FIG. 8 is a conceptual diagram for illustrating one example of block splitting.
[0024] Figure 11 FIG. 9 is a block diagram illustrating one example of a functional configuration of a splitter according to an embodiment.
[0025] Figure 12 FIG. 10 is a conceptual diagram for illustrating an example of a splitting mode.
[0026] Figure 13A FIG. 11 is a conceptual diagram for illustrating one example of a syntax tree of a splitting mode.
[0027] Figure 13B FIG. 12 is a conceptual diagram for illustrating another example of a syntax tree of a splitting mode.
[0028] Figure 14 FIG. 13 is a graph indicating example transform basis functions for various transform types.
[0029] Figure 15 FIG. 14 is a conceptual diagram for illustrating an example spatial variation transform (SVT).
[0030] Figure 16 FIG. 15 is a flowchart illustrating one example of a process performed by a transformer.
[0031] Figure 17 FIG. 16 is a flowchart illustrating another example of a process performed by a transformer.
[0032] Figure 18 is a block diagram illustrating one example of a functional configuration of a quantizer according to an embodiment.
[0033] Figure 19 is a flowchart illustrating one example of a quantization process performed by a quantizer.
[0034] Figure 20 is a block diagram illustrating one example of a functional configuration of an entropy encoder according to an embodiment.
[0035] Figure 21 is a conceptual diagram for illustrating one example of a flow of a context-based adaptive binary arithmetic coding (CABAC) process in an entropy encoder.
[0036] Figure 22 is a block diagram illustrating one example of a functional configuration of a loop filter according to an embodiment.
[0037] Figure 23A is a conceptual diagram for illustrating one example of a filter shape used in an adaptive loop filter (ALF).
[0038] Figure 23B is a conceptual diagram for illustrating another example of a filter shape used in an ALF.
[0039] Figure 23C is a conceptual diagram for illustrating another example of a filter shape used in an ALF.
[0040] Figure 23D is a conceptual diagram for illustrating one example of a flow of a cross-component ALF (CC-ALF).
[0041] Figure 23E is a conceptual diagram for illustrating one example of a filter shape used in a CC-ALF.
[0042] Figure 23F is a conceptual diagram for illustrating one example of a flow of a joint chroma CCALF (JC-CCALF).
[0043] Figure 23G is a table indicating example weighting index candidates that can be employed in a JC-CCALF.
[0044] Figure 24 is a block diagram indicating one example of a specific configuration of a loop filter used as a deblocking filter (DBF).
[0045] Figure 25 is a conceptual diagram for illustrating one example of a deblocking filter having a symmetric filtering characteristic with respect to a block boundary.
[0046] Figure 26is a conceptual diagram for illustrating a block boundary on which a deblocking filtering process is performed.
[0047] Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0048] Figure 28 is a flowchart illustrating one example of a process performed by a predictor of an encoder.
[0049] Figure 29 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0050] Figure 30 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0051] Figure 31 is a conceptual diagram for illustrating sixty-seven intra prediction modes used in intra prediction in an embodiment.
[0052] Figure 32 is a flowchart illustrating one example of a process performed by an intra predictor.
[0053] Figure 33 is a conceptual diagram for illustrating an example of a reference picture.
[0054] Figure 34 is a conceptual diagram for illustrating an example of a reference picture list.
[0055] Figure 35 is a flowchart illustrating an example basic process flow of inter prediction.
[0056] Figure 36 is a flowchart illustrating one example of a process of derivation of a motion vector.
[0057] Figure 37 is a flowchart illustrating another example of a process of derivation of a motion vector.
[0058] Figure 38A is a conceptual diagram for illustrating an example representation of modes for MV derivation.
[0059] Figure 38B is a conceptual diagram for illustrating an example representation of modes for MV derivation.
[0060] Figure 39 is a flowchart illustrating an example of a process of inter prediction in normal inter mode.
[0061] Figure 40 is a flowchart illustrating an example of a process of inter prediction in normal merge mode.
[0062] Figure 41 is a conceptual diagram for showing one example of a motion vector derivation process in merge mode.
[0063] Figure 42 is a conceptual diagram for showing one example of a MV derivation process for a current picture by HMVP merge mode.
[0064] Figure 43 is a flowchart showing one example of a frame rate up conversion (FRUC) process.
[0065] Figure 44 is a conceptual diagram for showing one example of a pattern matching between two blocks along a motion trajectory (bi-lateral matching).
[0066] Figure 45 is a conceptual diagram for showing one example of a pattern matching between a template in a current picture and a block in a reference picture (template matching).
[0067] Figure 46A is a conceptual diagram for showing one example of deriving a motion vector for each sub-block based on motion vectors of multiple neighboring blocks.
[0068] Figure 46B is a conceptual diagram for showing one example of deriving a motion vector for each sub-block in an affine mode where three control points are used.
[0069] Figure 47A is a conceptual diagram for showing one example of example MV derivation at control points in an affine mode.
[0070] Figure 47B is a conceptual diagram for showing one example of example MV derivation at control points in an affine mode.
[0071] Figure 47C is a conceptual diagram for showing one example of example MV derivation at control points in an affine mode.
[0072] Figure 48A is a conceptual diagram for showing an affine mode where two control points are used.
[0073] Figure 48B is a conceptual diagram for showing an affine mode where three control points are used.
[0074] Figure 49A is a conceptual diagram for showing one example of a method for MV derivation at control points when the number of control points used for an encoded block and the number of control points used for a current block are different from each other.
[0075] Figure 49Bis a conceptual diagram for illustrating another example of a process for MV derivation at a control point when the number of control points for the coded block and the number of control points for the current block are different from each other.
[0076] Figure 50 is a flowchart illustrating one example of a process in the affine merge mode.
[0077] Figure 51 is a flowchart illustrating one example of a process in the affine inter mode.
[0078] Figure 52A is a conceptual diagram for illustrating generation of two triangle prediction pictures.
[0079] Figure 52B is a conceptual diagram for illustrating a first portion of a first partition that overlaps a second partition and a first and second set of samples that can be weighted as part of a correction process.
[0080] Figure 52C is a conceptual diagram for illustrating a first portion of a first partition that is a portion of the first partition that overlaps a portion of a neighboring partition.
[0081] Figure 53 is a flowchart illustrating one example of a process in the triangle mode.
[0082] Figure 54 is a conceptual diagram for illustrating one example of an advanced temporal motion vector prediction (ATMVP) mode in which MVs are derived in sub-block units.
[0083] Figure 55 is a flowchart illustrating a relationship between the merge mode and dynamic motion vector refresh (DMVR).
[0084] Figure 56 is a conceptual diagram for illustrating one example of DMVR.
[0085] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining an MV.
[0086] Figure 58A is a conceptual diagram for illustrating one example of motion estimation in DMVR.
[0087] Figure 58B is a flowchart illustrating one example of a process of motion estimation in DMVR.
[0088] Figure 59 is a flowchart illustrating one example of a process of prediction picture generation.
[0089] Figure 60 is a flowchart showing another example of a process of generating a prediction image.
[0090] Figure 61 is a flowchart showing one example of a process of correcting a prediction image through overlapping block motion compensation (OBMC).
[0091] Figure 62 is a conceptual diagram for showing one example of a process of correcting a prediction image through OBMC.
[0092] Figure 63 is a conceptual diagram for showing a model assuming uniform straight-line motion.
[0093] Figure 64 is a flowchart showing one example of a process of inter prediction according to BIO.
[0094] Figure 65 is a functional block diagram showing one example of a functional configuration of an inter predictor that can perform inter prediction according to BIO.
[0095] Figure 66A is a conceptual diagram for showing one example of a process of a prediction image generation method using a luminance correction process performed by LIC.
[0096] Figure 66B is a flowchart showing one example of a process of a prediction image generation method using LIC.
[0097] Figure 67 is a block diagram showing a functional configuration of a decoder according to an embodiment.
[0098] Figure 68 is a functional block diagram showing an installation example of a decoder.
[0099] Figure 69 is a flowchart showing one example of an overall decoding process performed by a decoder.
[0100] Figure 70 is a conceptual diagram for showing a relationship between a split determiner and other constituent elements.
[0101] Figure 71 is a block diagram showing one example of a functional configuration of an entropy decoder.
[0102] Figure 72 is a conceptual diagram for showing an example flow of a CABAC process in an entropy decoder.
[0103] Figure 73 is a block diagram showing one example of a functional configuration of an inverse quantizer.
[0104] Figure 74 FIG. 38 is a flowchart illustrating one example of a process performed by an inverse quantizer.
[0105] Figure 75 FIG. 39 is a flowchart illustrating one example of a process performed by an inverse transformer.
[0106] Figure 76 FIG. 40 is a flowchart illustrating another example of a process performed by an inverse transformer.
[0107] Figure 77 FIG. 41 is a block diagram illustrating one example of a functional configuration of a loop filter.
[0108] Figure 78 FIG. 42 is a flowchart illustrating one example of a process performed by a predictor of a decoder.
[0109] Figure 79 FIG. 43 is a flowchart illustrating another example of a process performed by a predictor of a decoder.
[0110] Figure 80 FIG. 44 is a flowchart illustrating another example of a process performed by a predictor of a decoder.
[0111] Figure 81 FIG. 45 is a diagram illustrating one example of a process performed by an intra predictor of a decoder.
[0112] Figure 82 FIG. 46 is a flowchart illustrating one example of a process of MV derivation in a decoder.
[0113] Figure 83 FIG. 47 is a flowchart illustrating another example of a process of MV derivation in a decoder.
[0114] Figure 84 FIG. 48 is a flowchart illustrating an example of a process of inter prediction by normal inter mode in a decoder.
[0115] Figure 85 FIG. 49 is a flowchart illustrating an example of a process of inter prediction by normal merge mode in a decoder.
[0116] Figure 86 FIG. 50 is a flowchart illustrating an example of a process of inter prediction by FRUC mode in a decoder.
[0117] Figure 87 FIG. 51 is a flowchart illustrating an example of a process of inter prediction by affine merge mode in a decoder.
[0118] Figure 88 FIG. 52 is a flowchart illustrating an example of a process of inter prediction by affine inter mode in a decoder.
[0119] Figure 89 is a flowchart illustrating an example of a process of inter prediction by triangle mode in a decoder.
[0120] Figure 90 is a flowchart illustrating an example of a process of motion estimation by DMVR in a decoder.
[0121] Figure 91 is a flowchart illustrating an example process of motion estimation by DMVR in a decoder.
[0122] Figure 92 is a flowchart illustrating an example of a process of generation of a prediction image in a decoder.
[0123] Figure 93 is a flowchart illustrating another example of a process of generation of a prediction image in a decoder.
[0124] Figure 94 is a flowchart illustrating an example of a process of correction of a prediction image by OBMC in a decoder.
[0125] Figure 95 is a flowchart illustrating an example of a process of correction of a prediction image by BIO in a decoder.
[0126] Figure 96 is a flowchart illustrating an example of a process of correction of a prediction image by LIC in a decoder.
[0127] Figure 97 is a flowchart illustrating an example of a process of decoding a block using predicted chroma samples.
[0128] Figure 98 is a flowchart illustrating an example of a process of decoding a block using predicted chroma samples.
[0129] Figure 99 is a conceptual diagram for illustrating an example of determining whether a current chroma block is inside an MxN non-overlapping region aligned with an MxN grid of chroma samples.
[0130] Figure 100 is a conceptual diagram for illustrating an example of determining whether a current chroma block is inside an MxN non-overlapping region aligned with an MxN grid of chroma samples.
[0131] Figure 101 is a conceptual diagram for illustrating a virtual pipeline decoding unit (VPDU).
[0132] Figure 102 is a conceptual diagram for illustrating an example of determining whether a current VPDU can use luma samples to predict a chroma sample block.
[0133] Figure 103 is a conceptual diagram for illustrating an example way of determining whether a luma VPDU is to be split into smaller blocks.
[0134] Figure 104 is a conceptual diagram for illustrating additional considerations that can be taken into account to determine whether to use luma samples to predict chroma samples of a block.
[0135] Figure 105 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0136] Figure 106 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0137] Figure 107 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0138] Figure 108 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0139] Figure 109 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0140] Figure 110 is a conceptual diagram for illustrating an example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block.
[0141] Figure 111 is a conceptual diagram for illustrating an example of a partition of a non-rectangular shape.
[0142] Figure 112 is a diagram showing an example overall configuration of a content providing system for implementing a content distribution service.
[0143] Figure 113 is a conceptual diagram for illustrating an example of a display screen of a web page.
[0144] Figure 114 is a conceptual diagram for illustrating an example of a display screen of a web page.
[0145] Figure 115 is a block diagram showing one example of a smartphone.
[0146] Figure 116 is a block diagram showing an example of a functional configuration of a smartphone. DETAILED DESCRIPTION
[0147] In the drawings, like reference numerals refer to like elements throughout. The sizes and relative positions of elements in the drawings are not necessarily drawn to scale.
[0148] Hereinafter, embodiments will be described with reference to the accompanying drawings. Note that each of the embodiments described below shows a general or specific example. The numerical values, shapes, materials, components, arrangement and connection of components, steps, the relationship and order of steps, and the like indicated in the following embodiments are merely examples and do not intend to limit the scope of the claims.
[0149] Embodiments of an encoder and a decoder will be described below. The embodiments are examples of an encoder and a decoder to which processes and / or configurations presented in the description of aspects of the present disclosure are applied. The processes and / or configurations can also be implemented in an encoder and a decoder different from the encoder and the decoder according to the embodiments. For example, any of the following can be implemented with respect to the processes and / or configurations applied to the embodiments:
[0150] (1) Any of the components of the encoder or the decoder according to the embodiments presented in the description of aspects of the present disclosure can be replaced with or combined with another component presented anywhere in the description of aspects of the present disclosure.
[0151] (2) In the encoder or the decoder according to the embodiments, any change can be made to the functions or processes performed by one or more components of the encoder or the decoder, such as addition, replacement, removal, and the like of the functions or processes. For example, any function or process can be replaced with or combined with another function or process presented anywhere in the description of aspects of the present disclosure.
[0152] (3) In the method implemented by the encoder or the decoder according to the embodiments, any change can be made, such as addition, replacement, and removal of one or more of the processes included in the method. For example, any process in the method can be replaced with or combined with another process presented anywhere in the description of aspects of the present disclosure.
[0153] (4) One or more components included in the encoder or the decoder according to the embodiments can be combined with a component presented anywhere in the description of aspects of the present disclosure, can be combined with a component including one or more functions presented anywhere in the description of aspects of the present disclosure, and can be combined with a component implementing one or more processes implemented by a component presented in the description of aspects of the present disclosure.
[0154] (5) A component comprising one or more functions of an encoder or decoder according to an embodiment, or implementing one or more processes of an encoder or decoder according to an embodiment, can be combined with or replaced by a component presented anywhere in the description of aspects of the present disclosure, comprising one or more functions presented anywhere in the description of aspects of the present disclosure, or implementing one or more processes presented anywhere in the description of aspects of the present disclosure.
[0155] (6) In a method implemented by an encoder or decoder according to an embodiment, any of the processes included in the method can be replaced by or combined with a process presented anywhere in the description of aspects of the present disclosure, or any corresponding or equivalent process.
[0156] (7) One or more processes included in a method implemented by an encoder or decoder according to an embodiment can be combined with a process presented anywhere in the description of aspects of the present disclosure.
[0157] (8) Implementations of processes and / or configurations presented in the description of aspects of the present disclosure are not limited to encoders or decoders according to embodiments. For example, processes and / or configurations can be implemented in devices for different purposes than mobile picture encoders or mobile picture decoders disclosed in embodiments.
[0158] (Definitions of Terms)
[0159] Respective terms can be defined as indicated below as examples.
[0160] An image is a data unit configured with a set of pixels, is a picture, or includes a block smaller than a pixel. An image includes still images in addition to video.
[0161] A picture is an image processing unit configured with a set of pixels, and can also be referred to as a frame or a field. For example, a picture can take the form of an array of luma samples in monochrome format or the form of an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0162] A block is a processing unit that is a set of a determined number of pixels. A block can have any number of different shapes. For example, a block can have a rectangular shape of M x N pixels, a square shape of M x M pixels, a triangular shape, a circular shape, and so on. Examples of blocks include a slice, a tile, a brick, a CTU, a super block, a basic splitting unit, a VPDU, a processing splitting unit for hardware, a CU, a processing block unit, a prediction block unit (PU), an orthogonal transform block unit (TU), a unit, and a sub-block. A block can take the form of an M x N array of samples or an M x N array of transform coefficients. For example, a block can be a square or rectangular region of pixels including one luma matrix and two chroma matrices.
[0163] A pixel or a sample is a smallest point of an image. A pixel or a sample includes a pixel at an integer position, and a pixel at a sub-pixel position, e.g., a pixel generated based on a pixel at an integer position.
[0164] A pixel value or a sample value is a characteristic value of a pixel. A pixel value or a sample value can include one or more of a luma value, a chroma value, an RGB gradation level, a depth value, a binary value of zero or 1, and so on.
[0165] Chroma or chrominance is an intensity of a color, typically denoted by the symbols Cb and Cr, which specify a value of a sample array or a single sample value representing one of two color difference signals related to primary colors.
[0166] Luma or luminance is a brightness of an image, typically denoted by the symbols or subscripts Y or L, which specify a value of a sample array or a single sample value representing a value of a monochrome signal related to primary colors.
[0167] A flag includes one or more bits indicating a value of, e.g., a parameter or an index. A flag can be a binary flag indicating a binary value of the flag, or can indicate a non-binary value of a parameter.
[0168] A signal conveys information, which is signalled or encoded into the signal. A signal includes a discrete digital signal and a continuous analog signal.
[0169] A stream or a bitstream is a string of digital data of a digital data stream. A stream or a bitstream can be one stream, or can be configured with multiple streams having multiple hierarchical layers. A stream or a bitstream can be transmitted in serial communication using a single transmission path, or can be transmitted in packet communication using multiple transmission paths.
[0170] Difference refers to various mathematical differences, e.g., simple difference (x-y), absolute value of difference (|x-y|), squared difference (x^2-y^2), square root of difference (sqrt(x-y)), weighted difference (ax-by: a and b are constants), offset difference (x-y+a: a is an offset), etc. In the case of scalars, simple difference can be sufficient and includes difference computation.
[0171] Sum refers to various mathematical sums, e.g., simple sum (x+y), absolute value of sum (|x+y|), squared sum (x^2+y^2), square root of sum (sqrt(x+y)), weighted sum (ax+by: a and b are constants), offset sum (x+y+a: a is an offset), etc. In the case of scalars, simple sum can be sufficient and includes sum computation.
[0172] A frame is a combination of a top field and a bottom field, where sample lines 0, 2, 4,... originate from the top field and sample lines 1, 3, 5,... originate from the bottom field.
[0173] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) up to the next independent slice segment (if any) within the same access unit.
[0174] A tile is a rectangular region of coding tree blocks within a particular tile column and a particular tile row in a picture. A tile can be a rectangular region of a frame intended to be able to be independently decoded and encoded, but loop filtering across tile edges can still be applied.
[0175] A coding tree unit (CTU) can be a coding tree block of luma samples of a picture with three sample arrays, or two corresponding coding tree blocks of chroma samples. Alternatively, a CTU can be a coding tree block of samples of one of a monochrome picture and a picture coded using three separate color planes and syntax structures for coding samples. A superblock can be a 64x64-pixel square block composed of 1 or 2 mode information blocks, or recursively divided into four 32x32 blocks, which themselves can be further divided.
[0176] (System configuration)
[0177] First, a transmission system according to an embodiment will be described. Figure 1 is a schematic diagram showing one example of a configuration of a transmission system 400 according to an embodiment.
[0178] The transmission system 400 is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. As shown, the transmission system 400 includes an encoder 100, a network 300, and a decoder 200 as shown in Figure 1 .
[0179] An image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image, and outputs the stream to the network 300. The stream includes, for example, an encoded image and control information for decoding the encoded image. The image is compressed by encoding.
[0180] It should be noted that the image before encoding by the encoder 100 is also referred to as an original image, an original signal, or an original sample. The image can be a video or a still image. The image is a general concept of a sequence, a picture, and a block, and thus the image is not limited to a spatial region having a specific size and a temporal region having a specific size unless otherwise specified. The image is an array of pixels or pixel values, and a signal representing the image or the pixel values is also referred to as a sample. The stream can be referred to as a bitstream, an encoded bitstream, a compressed bitstream, or an encoded signal. Furthermore, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.
[0181] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 300 can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a bidirectional communication network, and can be a unidirectional communication network that transmits a broadcast wave of digital terrestrial broadcasting, satellite broadcasting, or the like. Alternatively, the network 300 can be replaced by a recording medium (for example, a digital versatile disc (DVD) and a Blu-ray disc (BD), or the like) on which the stream is recorded.
[0182] The decoder 200 generates a decoded image that is an uncompressed image, for example, by decoding the stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method employed by the encoder 100.
[0183] It should be noted that the decoder 200 can also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 can also be referred to as a decoding method, an image decoding method, or a video decoding method.
[0184] (Data structure)
[0185] Figure 2 is a conceptual diagram for illustrating a hierarchical structure of data in a stream. For convenience, the transmission system 400 of Figure 1 will be described. Figure 2 The stream includes, for example, a video sequence. As Figure 2As shown in (a), the video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplementary enhancement information (SEI), and multiple pictures.
[0186] In a video with multiple layers, a VPS can include decoding parameters that are common between some of the layers, as well as decoding parameters that are associated with some of the layers included in the video or with a single layer.
[0187] The SPS includes parameters for the sequence, that is, decoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the decoding parameters may indicate the width or height of the image. It should be noted that multiple SPSs may exist.
[0188] PPS includes parameters for the images, that is, decoding parameters that the decoder 200 refers to in order to decode each of the images in the sequence. For example, decoding parameters may include reference values for the quantization width used to decode the images and flags indicating the application of weighted predictions. It should be noted that multiple PPSs may exist. Each of the SPS and PPS can be simply referred to as a parameter set.
[0189] like Figure 2 As shown in (b), the image may include an image header and one or more slices. The image header includes decoding parameters, which the decoder 200 refers to to decode the one or more slices.
[0190] like Figure 2 As shown in (c), the slice includes a slice header and one or more bricks. The slice header includes decoding parameters that the decoder 200 refers to in order to decode the one or more bricks.
[0191] like Figure 2 As shown in (d), the bricks comprise one or more decoding tree units (CTUs).
[0192] It should be noted that an image may not include any slices and may include groups of slices instead of slices. In this case, a group of slices includes at least one slice. Alternatively, bricks may include slices.
[0193] CTU is also known as a superblock or basic split unit. For example... Figure 2 As shown in (e), the CTU includes a CTU header and at least one decoding unit (CU). As shown, the CTU includes four decoding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes decoding parameters that the decoder 200 refers to in order to decode at least one CU.
[0194] A CU can be split into multiple smaller CUs. As shown, CU (10) is not split into smaller coding units; CU (11) is split into four smaller coding units CU (110), CU (111), CU (112), and CU (113); CU (12) is not split into smaller coding units; and CU (13) is split into seven smaller coding units CU (1310), CU (1311), CU (1312), CU (1313), CU (132), CU (133), and CU (134). As shown, a CU is a square block of pixels. However, a CU can be a non-square block of pixels. Figure 2 As shown in (f), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information used to predict the CU, and the residual coefficient information is information indicating a prediction residual, which is described later. Although a CU is substantially the same as a prediction unit (PU) and a transform unit (TU), it should be noted that, for example, sub-block transform (SBT), which is described later, can include multiple TUs smaller than a CU. Additionally, a CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. A VPDU is, for example, a fixed unit, which can be processed in one stage when pipeline processing is performed in hardware.
[0195] It should be noted that a stream can not include all the hierarchical layers shown in Figure 3 Here, a picture that is a target of a process to be performed by a device (e.g., the encoder 100 or the decoder 200) is referred to as a current picture. The current picture represents a current picture to be coded when the process is an encoding process, and the current picture represents a current picture to be decoded when the process is a decoding process. Likewise, a CU or a CU block that is a target of a process to be performed by a device (e.g., the encoder 100 or the decoder 200) is referred to as a current block. The current block represents a current block to be coded when the process is an encoding process, and the current block represents a current block to be decoded when the process is a decoding process.
[0196] (Picture structure: slice / tile)
[0197] A picture can be configured with one or more slice units or one or more tile units to facilitate parallel coding / decoding of the picture.
[0198] A slice is a basic coding unit included in a picture. A picture can include, for example, one or more slices. Additionally, a slice includes one or more coding tree units (CTUs).
[0199] Figure 3 is a conceptual diagram for illustrating one example of a slice configuration. For example, in Figure 4In the example of FIG. 6A, a picture includes 11x8 CTUs, and is split into four slices (slices 1 to 4). Slice 1 includes sixteen CTUs, slice 2 includes twenty-one CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is a shape obtained by horizontally splitting the picture. The boundary of each slice does not need to coincide with the image endpoints, and can coincide with any one of the boundaries between the CTUs in the image. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, a raster scan order. A slice includes a slice header and encoded data. Characteristics of the slice can be written in the slice header. These characteristics can include the CTU address of the top CTU in the slice, the slice type, and the like.
[0200] A tile is a unit rectangular region included in a picture. Tiles of a picture can be assigned numbers called Tileld in a raster scan order.
[0201] Figure 4 is a conceptual diagram for illustrating one example of a tile configuration. For example, in Figure 4 In the example of FIG. 6A, a picture includes 11x8 CTUs, and is split into four slices (slices 1 to 4). Slice 1 includes sixteen CTUs, slice 2 includes twenty-one CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is a shape obtained by horizontally splitting the picture. The boundary of each slice does not need to coincide with the image endpoints, and can coincide with any one of the boundaries between the CTUs in the image. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, a raster scan order. A slice includes a slice header and encoded data. Characteristics of the slice can be written in the slice header. These characteristics can include the CTU address of the top CTU in the slice, the slice type, and the like. Figure 7 In the example of FIG. 6A, a picture includes 11x8 CTUs, and is split into four slices (slices 1 to 4). Slice 1 includes sixteen CTUs, slice 2 includes twenty-one CTUs, slice 3 includes twenty-nine CTUs, and slice 4 includes twenty-two CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is a shape obtained by horizontally splitting the picture. The boundary of each slice does not need to coincide with the image endpoints, and can coincide with any one of the boundaries between the CTUs in the image. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, a raster scan order. A slice includes a slice header and encoded data. Characteristics of the slice can be written in the slice header. These characteristics can include the CTU address of the top CTU in the slice, the slice type, and the like.
[0202] It should be noted that one tile can include one or more slices, and one slice can include one or more tiles.
[0203] It should be noted that a picture can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more tiles. A picture can be configured with one of a tile set, a tile group, and a tile. For example, assume that the order in which a plurality of tiles are scanned for each tile set in a raster scan order is the basic encoding order of the tiles. Assume that the set of one or more tiles that are consecutive in the basic encoding order in each tile set is a tile group. Such a picture can be configured by the splitter 102 (see Figure 5 ) described later.
[0204] (Scalable coding)
[0205] Figure 6 andFigure 1 is a conceptual diagram showing an example of a scalable stream structure, and will be described with reference to Figure 5 .
[0206] As shown in Figure 6 , the encoder 100 can generate a time / space scalable stream by dividing each of a plurality of pictures into any one of a plurality of layers and encoding the pictures in the layers. For example, the encoder 100 encodes the pictures of each layer, thereby realizing scalability in the case where an enhancement layer exists above a base layer. Such encoding of each picture is also called scalable encoding. In this way, the decoder 200 is able to switch the image quality of an image displayed by decoding the stream. In other words, the decoder 200 can determine which layer to decode based on internal factors such as the processing capability of the decoder 200 and external factors such as the state of the communication bandwidth. As a result, the decoder 200 is able to decode the content while freely switching between low resolution and high resolution. For example, a user of a stream watches half of a stream video using a smartphone on the way home, and continues to watch the video on a device such as a television connected to the Internet at home. It should be noted that each of the smartphone and the device described above includes a decoder 200 having the same or different performance. In this case, when the device decodes up to a layer of a higher layer in the stream, the user can watch a high-quality video at home. In this way, the encoder 100 does not need to generate a plurality of streams having different image qualities of the same content, and thus can reduce the processing load.
[0207] In addition, the enhancement layer can include meta information based on statistical information about the image. The decoder 200 can generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the meta data. Super-resolution imaging can include, for example, improvement of the SN ratio, increase in resolution, and the like at the same resolution. The meta data can include, for example, information for identifying linear or nonlinear filter coefficients as used in the super-resolution process, or information identifying parameter values in a filtering process, machine learning, or least squares used in the super-resolution process, and the like.
[0208] In an embodiment, a configuration can be provided in which a picture is divided into, for example, slices according to the meaning of an object in the picture. In this case, the decoder 200 can decode only a partial area in the picture by selecting a slice to be decoded. Additionally, the attribute of the object (person, car, ball, and the like) and the position of the object in the picture (coordinates in the same image) can be stored as meta data. In this case, the decoder 200 is able to identify the position of a desired object based on the meta data, and determine a slice including the object. For example, as shown in Figure 7As shown in FIG. 1, the metadata can be stored using a data storage structure different from the image data (e.g., SEI (Supplemental Enhancement Information) messages in HEVC). The metadata indicates, for example, the position, size, or color of a main object.
[0209] The metadata can be stored in units of a plurality of pictures (e.g., stream, sequence, random access unit). In this way, the decoder 200 can obtain, for example, the time when a certain person appears in a video, and by fitting the time information with the picture unit information, can identify the picture in which the object (person) is present and determine the position of the object in the picture.
[0210] (Encoder)
[0211] An encoder according to an embodiment will be described. Figure 7 is a block diagram showing a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes a video in units of blocks.
[0212] As Figure 8 As shown in FIG. 1, the encoder 100 is an apparatus that encodes an image in units of blocks, and includes a splitter 102, a subtracter 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra predictor 124, an inter predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra predictor 124 and the inter predictor 126 are part of the prediction controller.
[0213] The encoder 100 is realized as, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitter 102, the subtracter 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra predictor 124, the inter predictor 126, and the prediction controller 128. Alternatively, the encoder 100 can be realized as one or a plurality of dedicated electronic circuits corresponding to the splitter 102, the subtracter 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra predictor 124, the inter predictor 126, and the prediction controller 128.
[0214] (Installation example of encoder)
[0215] Figure 7 is a functional block diagram showing an installation example of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example,Figure 8 The plurality of constituent elements of the encoder 100 shown in FIG. 1 are mounted on Figure 7 The processor a1 and the memory a2 shown in FIG. 1.
[0216] The processor a1 is a circuit that performs information processing and is coupled to the memory a2. For example, the processor a1 is a dedicated or general electronic circuit that encodes an image. The processor a1 can be a processor such as a CPU. Additionally, the processor a1 can be a collection of a plurality of electronic circuits. Additionally, for example, the processor a1 can assume the role of two or more of the constituent elements of the encoder 100 shown in FIG. 1, and the like. Figure 7
[0217] The memory a2 is a dedicated or general memory that stores information used by the processor a1 to encode an image. The memory a2 can be an electronic circuit and can be connected to the processor a1. Additionally, the memory a2 can be included in the processor a1. Additionally, the memory a2 can be a collection of a plurality of electronic circuits. Additionally, the memory a2 can be a magnetic disk, an optical disk, or the like, or can be denoted as a storage device, a recording medium, or the like. Additionally, the memory a2 can be a non-volatile memory or a volatile memory.
[0218] For example, the memory a2 can store an image to be encoded or a bitstream corresponding to an encoded image. Additionally, the memory a2 can store a program for causing the processor a1 to encode an image.
[0219] Additionally, for example, the memory a2 can assume the role of two or more of the constituent elements of the encoder 100 shown in FIG. 1 for storing information, and the like. For example, the memory a2 can assume the role of the block memory 118 and the frame memory 122 shown in FIG. 1. More specifically, the memory a2 can store a reconstructed block, a reconstructed picture, or the like. Figure 7 Figure 7 It should be noted that, in the encoder 100, all of the constituent elements indicated in FIG. 1, and the like, can not be implemented, and all of the processes described herein can not be performed.
[0220] It should be noted that, in the encoder 100, all of the constituent elements indicated in FIG. 1, and the like, can not be implemented, and all of the processes described herein can not be performed. Figure 7 Part of the constituent elements indicated in FIG. 1, and the like, can be included in another device, or part of the processes described herein can be performed by another device. Figure 9
[0221] Hereinafter, the overall flow of the processes performed by the encoder 100 is described, and then each of the constituent elements included in the encoder 100 will be described.
[0222] (Overall flow of encoding processes)
[0223] Figure 7 is a flowchart indicating one example of the overall encoding process performed by the encoder 100, and will be described with reference to Figure 10 described.
[0224] First, the splitter 102 of the encoder 100 splits each of the pictures included in the input image into a plurality of blocks having a fixed size (e.g., 128 x 128 pixels) (step Sa_1). The splitter 102 then selects a split mode for the fixed-size block (also referred to as a block shape) (step Sa_2). In other words, the splitter 102 further splits the fixed-size block into a plurality of blocks forming the selected split mode. For each of the plurality of blocks, the encoder 100 performs steps Sa_3 to Sa_9 with respect to the block (i.e., the current block to be encoded).
[0225] The prediction controller 128 and the prediction executers (including the intra predictor 124 and the inter predictor 126) generate a prediction image of the current block (step Sa-3). The prediction image can also be referred to as a prediction signal, a prediction block, or a prediction sample.
[0226] Next, the subtracter 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual can also be referred to as a prediction error.
[0227] Next, the transformer 106 transforms the prediction image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients can sometimes be referred to as a coefficient block.
[0228] Next, the entropy encoder 110 encodes (specifically, entropy-encodes) the plurality of quantized coefficients and prediction parameters related to the generation of the prediction image to generate a stream (step Sa_6). The stream can sometimes be referred to as an encoded bitstream or a compressed bitstream.
[0229] Next, the inverse quantizer 112 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 114 performs inverse transformation on the result to restore the prediction residual (step Sa_7).
[0230] Next, the adder 116 adds the prediction image and the restored prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image can also be referred to as a reconstructed block or a decoded image block.
[0231] When the reconstructed image is generated, the loop filter 120 performs filtering on the reconstructed image as necessary (step Sa_9).
[0232] Then, the encoder 100 determines whether the encoding of the entire picture has ended (step Sa_10). When it is determined that the encoding has not ended (NO in step Sa_10), the process starting with step Sa_2 is repeatedly performed for the next block of the picture.
[0233] While the encoder 100 selects one split mode for the fixed-size block and encodes each block according to the split mode in the example described above, it should be noted that each block can be encoded according to a corresponding one of a plurality of split modes. In this case, the encoder 100 can evaluate the cost of each of the plurality of split modes, and for example, can select the stream obtained by encoding according to the split mode that yields the smallest cost as the output stream.
[0234] As illustrated, the processes in steps Sa_1 to Sa_10 are sequentially performed by the encoder 100. Alternatively, two or more of the processes can be performed in parallel, the processes can be reordered, and the like.
[0235] The encoding process employed by the encoder 100 is hybrid encoding using predictive encoding and transform encoding. Additionally, the predictive encoding is performed by an encoding loop configured with the subtractor 104, the transformer 106, the quantizer 108, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the block memory 118, the frame memory 122, the intra predictor 124, the inter predictor 126, and the prediction controller 128. In other words, the prediction performer configured with the intra predictor 124 and the inter predictor 126 is part of the encoding loop.
[0236] (Splitter)
[0237] The splitter 102 splits each picture included in the original image into a plurality of blocks, and outputs each block to the subtractor 104. For example, the splitter 102 first splits a picture into blocks of a fixed size (e.g., 128 x 128 pixels). Other fixed block sizes can be employed. The fixed-size block is also referred to as a coding tree unit (CTU). Then, the splitter 102 splits each fixed-size block into blocks of a variable size (e.g., 64 x 64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a split mode. The variable-size block can also be referred to as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). It should be noted that, in various categories of processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or some of the blocks in a picture can be processed in units of CUs, PUs, or TUs.
[0238] Figure 10 is a conceptual diagram for illustrating one example of block splitting according to an embodiment. In this example, the fixed-size block is split into four variable-size blocks.Figure 10 In the figure, solid lines indicate block boundaries of blocks split by quad-tree block splitting, and dashed lines indicate block boundaries of blocks split by binary-tree block splitting.
[0239] Here, the block 10 is a square block having 128 x 128 pixels (128 x 128 block). The 128 x 128 block 10 is first split into four square 64 x 64 pixel blocks (quad-tree block splitting).
[0240] The 64 x 64 pixel block at the upper left is further vertically split into two rectangular 32 x 64 pixel blocks, and the 32 x 64 pixel block at the left is further vertically split into two rectangular 16 x 64 pixel blocks (binary-tree block splitting). As a result, the 64 x 64 pixel block at the upper left is split into two 16 x 64 pixel blocks 11 and 12 and one 32 x 64 pixel block 13.
[0241] The 64 x 64 pixel block at the upper right is horizontally split into two rectangular 64 x 32 pixel blocks 14 and 15 (binary-tree block splitting).
[0242] The 64 x 64 pixel block at the lower left is first split into four square 32 x 32 pixel blocks (quad-tree block splitting). The block at the upper left and the block at the lower right among the four square 32 x 32 pixel blocks are further split. The square 32 x 32 pixel block at the upper left is vertically split into two rectangular 16 x 32 pixel blocks, and the 16 x 32 pixel block at the right is further horizontally split into two 16 x 16 pixel blocks (binary-tree block splitting). The 32 x 32 pixel block at the lower right is horizontally split into two 32 x 16 pixel blocks (binary-tree block splitting). The square 32 x 32 pixel block at the upper right is horizontally split into two rectangular 32 x 16 pixel blocks (binary-tree block splitting). As a result, the 64 x 64 pixel block at the lower left is split into a rectangular 16 x 32 pixel block 16, two square 16 x 16 pixel blocks 17 and 18, two square 32 x 32 pixel blocks 19 and 20, and two rectangular 32 x 16 pixel blocks 21 and 22.
[0243] The 64 x 64 pixel block 23 at the lower right is not split.
[0244] As described above, in the case where the block 10 is split into thirteen variable-size blocks 11 to 23 based on the recursive quad-tree and binary-tree block splitting, the 64 x 64 pixel block at the upper left is split into two 16 x 64 pixel blocks 11 and 12 and one 32 x 64 pixel block 13, the 64 x 64 pixel block at the upper right is split into two rectangular 64 x 32 pixel blocks 14 and 15, the 64 x 64 pixel block at the lower left is split into a rectangular 16 x 32 pixel block 16, two square 16 x 16 pixel blocks 17 and 18, two square 32 x 32 pixel blocks 19 and 20, and two rectangular 32 x 16 pixel blocks 21 and 22, and the 64 x 64 pixel block 23 at the lower right is not split. Figure 10 In the figure, solid lines indicate block boundaries of blocks split by quad-tree block splitting, and dashed lines indicate block boundaries of blocks split by binary-tree block splitting.
[0245] It should be noted that, in the case where the block 10 is split into thirteen variable-size blocks 11 to 23 based on the recursive quad-tree and binary-tree block splitting, the 64 x 64 pixel block at the upper left is split into two 16 x 64 pixel blocks 11 and 12 and one 32 x 64 pixel block 13, the 64 x 64 pixel block at the upper right is split into two rectangular 64 x 32 pixel blocks 14 and 15, the 64 x 64 pixel block at the lower left is split into a rectangular 16 x 32 pixel block 16, two square 16 x 16 pixel blocks 17 and 18, two square 32 x 32 pixel blocks 19 and 20, and two rectangular 32 x 16 pixel blocks 21 and 22, and the 64 x 64 pixel block 23 at the lower right is not split. Figure 11In the HEVC, one block is split into four or two blocks (quad-tree or binary-tree block splitting), but the splitting is not limited to these examples. For example, one block can be split into three blocks (ternary block splitting). The splitting including such ternary block splitting is also referred to as multi-type tree (MBT) splitting.
[0246] Figure 11 is a block diagram showing one example of a functional configuration of the splitter 102 according to one embodiment. As shown in Figure 12 As shown in the HEVC, the splitter 102 can include a block split determiner 102a. As an example, the block split determiner 102a can perform the following process.
[0247] For example, the block split determiner 102a can obtain or retrieve block information from the block memory 118 and / or the frame memory 122, and determine a split pattern (e.g., the split patterns described above) based on the block information. The splitter 102 splits the original image according to the split pattern, and outputs at least one block obtained by the splitting to the subtracter 104.
[0248] Additionally, for example, the block split determiner 102a outputs one or more parameters indicating the determined split pattern (e.g., the split patterns described above) to the transformer 106, the inverse transformer 114, the intra predictor 124, the inter predictor 126, and the entropy encoder 110. The transformer 106 can transform the prediction residual based on the one or more parameters. The intra predictor 124 and the inter predictor 126 can generate the prediction image based on the one or more parameters. Additionally, the entropy encoder 110 can entropy-encode the one or more parameters.
[0249] As indicated below, the parameters related to the split pattern can be written in a stream, as one example.
[0250] Figure 13A is a conceptual diagram for showing examples of the split pattern. The examples of the split pattern include: split into four regions (QT) in which a block is split into two regions horizontally and into two regions vertically; split into three regions (HT or VT) in which a block is split in a 1:2:1 ratio in the same direction; split into two regions (HB or VB) in which a block is split in a 1:1 ratio in the same direction; and no split (NS).
[0251] It should be noted that the split pattern does not have a block split direction in the case of split into four regions and no split, and the split pattern has split direction information in the case of split into two regions or three regions.
[0252] Figure 13B is a conceptual diagram for showing one example of a syntax tree of the split pattern.
[0253] Figure 13A is a conceptual diagram for illustrating another example of a syntax tree of a split mode.
[0254] Figure 13B and Figure 13A is a conceptual diagram for illustrating an example of a syntax tree of a split mode. In the example of Figure 13A , first, there is information indicating whether or not to perform splitting (S: split flag), and next, there is information indicating whether or not to perform splitting into four regions (QT: QT flag). Next, there is information indicating which one of splitting into three regions and splitting into two regions is to be performed (TT: TT flag, or BT: BT flag), and then, there is information indicating a division direction (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each of at least one block obtained by splitting according to such a split mode can be further repeatedly split in a similar procedure. In other words, as one example, it can be recursively determined whether or not to perform splitting, whether or not to perform splitting into four regions, which one of a horizontal direction and a vertical direction is a direction of a splitting method to be performed, which one of splitting into three regions and splitting into two regions is to be performed, and the determination results can be encoded in a stream according to an encoding order disclosed by the syntax tree illustrated in Figure 13A .
[0255] Additionally, although the information items indicating S, QT, TT, and Ver, respectively, are arranged in the listed order in the syntax tree illustrated in Figure 13B , the information items indicating S, QT, Ver, and BT, respectively, can be arranged in the listed order. In other words, in the example of Figure 14 , first, there is information indicating whether or not to perform splitting (S: split flag), and next, there is information indicating whether or not to perform splitting into four regions (QT: QT flag). Next, there is information indicating a division direction (Ver: vertical flag, or Hor: horizontal flag), and then, there is information indicating which one of splitting into two regions and splitting into three regions is to be performed (BT: BT flag, or TT: TT flag).
[0256] It should be noted that the split modes described above are examples, and a split mode other than the described split modes can be used, or a part of the described split modes can be used.
[0257] (subtracter)
[0258] The subtractor 104 subtracts a prediction image (a prediction sample input from the prediction controller 128 indicated below) from the original image input from the splitter 102 and split by the splitter 102 in units of blocks. In other words, the subtractor 104 calculates a prediction residual (also referred to as an error) of a current block. The subtractor 104 then outputs the calculated prediction residual to the transformer 106.
[0259] The original image can be an image for which a signal (e.g., a luminance signal and two chrominance signals) representing the image has been input into the encoder 100 as a picture included in a video. The signal representing the image can also be referred to as a sample.
[0260] (transformer)
[0261] The transformer 106 transforms the prediction residual in the spatial domain into a transform coefficient in the frequency domain, and outputs the transform coefficient to the quantizer 108. More specifically, the transformer 106 applies a discrete cosine transform (DCT) or a discrete sine transform (DST) defined, for example, to the prediction residual in the spatial domain. The DCT or DST defined can be predefined.
[0262] It should be noted that the transformer 106 can adaptively select a transform type from a plurality of transform types, and transform the prediction residual into a transform coefficient by using a transform basis function corresponding to the selected transform type. This way of transform is also referred to as explicit multi-kernel transform (EMT) or adaptive multi- transform (AMT). The transform basis function can also be referred to as a basis.
[0263] The transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. It should be noted that these transform types can be denoted as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 is a graph indicating an example transform basis function for an example transform type. In Figure 15 In, N indicates the number of input pixels. For example, the selection of the transform type from a plurality of transform types can depend on a prediction type (one of intra prediction and inter prediction), and can depend on an intra prediction mode.
[0264] Information indicating whether to apply such EMT or AMT (referred to as, for example, an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. It should be noted that signaling such information does not necessarily need to be performed at the CU level, and can be performed at another level (e.g., a sequence level, a picture level, a slice level, a tile level, or a CTU level).
[0265] Additionally, the transformer 106 can perform a retransformation on the transform coefficients, which are the result of the transformation. This retransformation is also referred to as adaptive secondary transform (AST) or non-separable secondary transform (NSST). For example, the transformer 106 performs the retransformation in units of sub-blocks (e.g., 4x4 pixel sub-blocks) included in a transform coefficient block corresponding to an intra prediction residual. Information indicating whether to apply the NSST and information related to a transform matrix used for the NSST are generally signaled at the CU level. It should be noted that signaling of such information does not necessarily need to be performed at the CU level and can be performed at another level (e.g., sequence level, picture level, slice level, tile level, or CTU level).
[0266] The transformer 106 can employ separable transform and non-separable transform. The separable transform is a method in which a transform is performed multiple times by individually performing a transform for each of a plurality of directions according to the number of dimensions of an input. The non-separable transform is a method of performing a collective transform in which two or more dimensions in a multi-dimensional input are collectively regarded as a single dimension.
[0267] In one example of the non-separable transform, when the input is a 4x4 pixel block, the 4x4 pixel block is regarded as a single array including sixteen elements, and the transform applies a 16x16 transform matrix to the array.
[0268] In another example of the non-separable transform, an input block of 4x4 pixels is regarded as a single array including sixteen elements, and then a transform in which the array is given multiple rotations (hypercube given transform) can be performed.
[0269] In the transform in the transformer 106, the transform type to be transformed into a transform basis function in the frequency domain according to the region in the CU can be switched. Examples include a spatially varying transform (SVT).
[0270] Figure 15 is a conceptual diagram for illustrating one example of the SVT.
[0271] In the SVT, as illustrated in Figure 15 the CU is split horizontally or vertically into two equal regions, and only one of the regions is transformed into the frequency domain. The transform base type can be set for each region. For example, DST7 and DST8 are used. For example, among two regions obtained by splitting the CU vertically into two equal regions, DST7 and DCT8 can be used for the region at position 0. Alternatively, among the two regions, DST7 is used for the region at position 1. Similarly, among two regions obtained by splitting the CU horizontally into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, among the two regions, DST7 is used for the region at position 1. Although in the above examples, the DSTs are used, the DCTs can be used instead of the DSTs.Figure 16 In the example shown in the middle, only one of the two regions in the CU is transformed and the other is not, but each of the two regions can be transformed. Additionally, the splitting method can include not only splitting into two regions, but also splitting into four regions. Additionally, the splitting method can be more flexible. For example, information indicating the splitting method can be encoded and can be signaled in the same manner as the CU split. It should be noted that SVT can also be referred to as sub-block transform (SBT).
[0272] The AMT and EMT described above can be referred to as MTS (multiple transform selection). When MTS is applied, a transform type such as DST7, DCT8, etc. can be selected, and information indicating the selected transform type can be encoded as index information for each CU. There is another process called IMTS (implicit MTS) as a process for selecting a transform type to be used for orthogonal transform performed without encoding index information. When IMTS is applied, for example, when a CU has a rectangular shape, orthogonal transform of the rectangular shape can be performed using DST7 for the short side and DST2 for the long side. Additionally, for example, when a CU has a square shape, orthogonal transform of the rectangular shape can be performed using DCT2 when MTS is active in the sequence and DST7 when MTS is inactive in the sequence. DCT2 and DST7 are only examples. Other transform types can be used, and the combination of changing transform types can also be changed for different combinations of transform types. IMTS can be used only for intra prediction blocks, or can be used for both intra prediction blocks and inter prediction blocks.
[0273] The three processes of MTS, SBT, and IMTS have been described above as selection processes for selectively switching transform types for a transform. However, all three selection processes can be employed, or only a portion of the selection processes can be selectively employed. For example, whether one or more of the selection processes is employed can be identified based on flag information in a header such as an SPS, or the like. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and a transform of the CU is performed. It should be noted that the selection processes for selectively switching transform types can be different selection processes from the three selection processes above, or each of the three selection processes can be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is performed. Function [1] is a function for performing a transform of an entire CU and encoding information indicating a transform type used in the transform. Function [2] is a function for performing a transform of an entire CU and determining a transform type based on a determined rule without encoding information indicating the transform type. Function [3] is a function for performing a transform of a partial region of a CU and encoding information indicating a transform type used in the transform. Function [4] is a function for performing a transform of a partial region of a CU and determining a transform type based on a determined rule without encoding information indicating the transform type used in the transform. The determined rule can be predetermined.
[0274] It should be noted that whether MTS, IMTS, and / or SBT is applied can be determined for each processing unit. For example, whether MTS, IMTS, and / or SBT is applied can be determined for each sequence, picture, tile, slice, CTU, or CU.
[0275] It should be noted that the tools for selectively switching transform types in this disclosure can be described as methods, selection processes, or processes for selecting a basis used in a transform process for selectively selecting. Additionally, the tools for selectively switching transform types can be described as modes for adaptively selecting a transform type.
[0276] Figure 7 is a flowchart showing one example of a process performed by the transformer 106, and will be described with reference to Figure 17 for convenience.
[0277] For example, the transformer 106 determines whether to perform the orthogonal transform (step St_l). Here, when it is determined that the orthogonal transform is to be performed (Yes in step St_l), the transformer 106 selects a transform type for the orthogonal transform from among a plurality of transform types (step St_2). Next, the transformer 106 performs the orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transform type to the entropy encoder 110 so as to allow the entropy encoder 110 to encode the information (step St_4). On the other hand, when it is determined that the orthogonal transform is not to be performed (No in step St_l), the transformer 106 outputs information indicating that the orthogonal transform is not performed so as to allow the entropy encoder 110 to encode the information (step St_5). It should be noted that whether or not the orthogonal transform is performed can be determined in step St_l based on, for example, the size of the transform block, the prediction mode applied to the CU, and the like. Alternatively, the orthogonal transform can be performed using a defined transform type without encoding information indicating the transform type used in the orthogonal transform. The defined transform type can be predefined.
[0278] Figure 7 is a flowchart showing one example of a process performed by the transformer 106, and will be described with reference to Figure 17 It should be noted that Figure 16 the example shown in Figure 18 is an example of the orthogonal transform in a case where the transform type used in the orthogonal transform is selectively switched (as in the case of the example shown in
[0279] As one example, the first transform type group can include DCT2, DST7, and DCT8. As another example, the second transform type group can include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group can partially overlap each other, or can be completely different from each other.
[0280] The transformer 106 determines whether the transform size is less than or equal to a determined value (step Su_1). Here, when the transform size is determined to be less than or equal to the determined value (Yes in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using a transform type included in a first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one transform type included in the first transform type group to the entropy encoder 110 in order to allow the entropy encoder 110 to encode the information (step Su_3). On the other hand, when the transform size is determined not to be less than or equal to the predetermined value (No in step Su_1), the transformer 106 performs an orthogonal transform on the prediction residual of the current block using a second transform type group (step Su_4). The determined value can be a threshold value, and can be a predetermined value.
[0281] In step Su_3, the information indicating the transform type used in the orthogonal transform can be information indicating a combination of a transform type to be applied vertically in the current block and a transform type to be applied horizontally in the current block. The first type group can include only one transform type, and the information indicating the transform type used in the orthogonal transform can not be encoded. The second transform type group can include a plurality of transform types, and the information indicating the transform type used in the orthogonal transform among one or more transform types included in the second transform type group can be encoded.
[0282] Alternatively, the transform type can be indicated based on the transform size without encoding the information indicating the transform type. It should be noted that such determination is not limited to the determination of whether the transform size is less than or equal to the determined value, and other processes can also be used to determine the transform type used in the orthogonal transform based on the transform size.
[0283] (Quantizer)
[0284] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in the determined scan order, and quantizes the scanned transform coefficients based on a quantization parameter (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients (hereinafter, also referred to as quantized coefficients) of the current block to the entropy encoder 110 and the inverse quantizer 112. The determined scan order can be predetermined.
[0285] The determined scan order is an order for quantizing / inverse quantizing the transform coefficients. For example, the determined scan order can be defined as an ascending order of frequency (from low frequency to high frequency) or a descending order of frequency (from high frequency to low frequency).
[0286] A quantization parameter (QP) is a parameter that defines a quantization step size (quantization width). For example, when the value of the quantization parameter increases, the quantization step size also increases. In other words, when the value of the quantization parameter increases, the error (quantization error) of the quantized coefficient increases.
[0287] Additionally, a quantization matrix can be used for quantization. For example, several kinds of quantization matrices can be used corresponding to a frequency transform size such as 4x4, 8x8, a prediction mode such as intra prediction and inter prediction, a pixel component such as a luminance and a chrominance pixel component. It should be noted that quantization means to digitize a value that is sampled with a determined interval corresponding to a determined level. In this technical field, quantization can be referred to using other expressions (e.g., rounding and scaling), and rounding and scaling can be employed. The determined interval and the determined level can be predetermined.
[0288] Methods using a quantization matrix can include a method of using a quantization matrix that has been directly set at the encoder 100 side, and a method of using a quantization matrix that has been set as a default (default matrix). At the encoder 100 side, a quantization matrix that is suitable for the characteristics of an image can be set by directly setting the quantization matrix. However, this case can have a disadvantage of increasing the amount of coding for encoding the quantization matrix. It should be noted that a quantization matrix for quantizing a current block can be generated based on a default quantization matrix or an encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.
[0289] There is a method for quantizing high frequency coefficients and low frequency coefficients without using a quantization matrix. It should be noted that this method can be regarded as a method equivalent to using a quantization matrix whose coefficients have the same value (flat matrix).
[0290] A quantization matrix can be encoded, for example, at a sequence level, a picture level, a slice level, a tile level, or a CTU level. A quantization matrix can be specified using, for example, a sequence parameter set (SPS) or a picture parameter set (PPS). The SPS includes parameters for a sequence, and the PPS includes parameters for a picture. Each of the SPS and the PPS can be simply referred to as a parameter set.
[0291] When a quantization matrix is used, the quantizer 108 scales the quantization width that can be calculated based on a quantization parameter or the like, for each transform coefficient, using the values of the quantization matrix. The quantization process performed without using a quantization matrix can be a process for quantizing transform coefficients according to a quantization width calculated based on a quantization parameter or the like. It should be noted that in the quantization process performed without using any quantization matrix, the quantization width can be multiplied by a determined value that is common to all transform coefficients in a block. The determined value can be predetermined.
[0292] Figure 19 is a block diagram illustrating one example of a functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a delta quantization parameter generator 108a, a predicted quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter storage 108d, and a quantization executor 108e.
[0293] Figure 7 is a flowchart illustrating one example of a quantization process performed by the quantizer 108, and will be described with reference to Figure 18 and Figure 19 for convenience.
[0294] As one example, the quantizer 108 can perform quantization for each CU based on Figure 20 the flowchart illustrated in FIG. 13. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined that quantization is to be performed (Yes in step Sv_1), the quantization parameter generator 108c generates a quantization parameter for the current block (step Sv_2), and stores the quantization parameter to the quantization parameter storage 108d (step Sv_3).
[0295] Next, the quantization executor 108e quantizes the transform coefficients of the current block using the quantization parameter generated in step Sv_2 (step Sv_4). The predicted quantization parameter generator 108b then obtains a quantization parameter for a processing unit different from the current block from the quantization parameter storage 108d (step Sv_5). The predicted quantization parameter generator 108b generates a predicted quantization parameter for the current block based on the obtained quantization parameter (step Sv_6). The delta quantization parameter generator 108a calculates a difference between the quantization parameter of the current block generated by the quantization parameter generator 108c and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 108b (step Sv_7). The delta quantization parameter can be generated by calculating the difference. The delta quantization parameter generator 108a outputs the delta quantization parameter to the entropy encoder 110 so as to allow the entropy encoder 110 to encode the delta quantization parameter (step Sv_8).
[0296] It should be noted that the delta quantization parameter can be encoded at, for example, a sequence level, a picture level, a slice level, a tile level, or a CTU level. Additionally, an initial value of the quantization parameter can be encoded at a sequence level, a picture level, a slice level, a tile level, or a CTU level. At the initialization, the initial value of the quantization parameter and the delta quantization parameter can be used to generate the quantization parameter.
[0297] It should be noted that the quantizer 108 can include a plurality of quantizers, and dependent quantization can be applied in which a quantization method selected from a plurality of quantization methods is used to quantize the transform coefficients.
[0298] (Entropy encoder)
[0299] Figure 7 is a block diagram showing one example of a functional configuration of the entropy encoder 110 according to an embodiment, and will be described with reference to Figure 21 The entropy encoder 110 generates a stream by entropy-encoding the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy encoding. More specifically, the illustrated entropy encoder 110 includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization in which a multi-level signal such as a quantized coefficient and a prediction parameter is transformed into a binary signal. Examples of the binarization method include truncated Rice binarization, exponential Golomb code, and fixed length binarization. The context controller 110b derives a context value according to a characteristic of a syntax element or a surrounding state (i.e., an occurrence probability of a binary signal). Examples of the method for deriving a context value include bypass, reference to a syntax element, reference to an above and left neighboring block, reference to hierarchical information, and the like. The binary arithmetic encoder 110c performs arithmetic encoding of a binary signal using the derived context.
[0300] Figure 22 is a conceptual diagram for showing an example flow of a CABAC process in the entropy encoder 110. First, initialization in CABAC is performed in the entropy encoder 110. In the initialization, initialization in the binary arithmetic encoder 110c and setting of an initial context value are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic encoding of a plurality of quantized coefficients in a CTU. Each time arithmetic encoding is performed, the context controller 110b can update a context value. The context controller 110b can then save the context value as post-processing. For example, the saved context value can be used to initialize a context value for a next CTU.
[0301] (Inverse quantizer)
[0302] The inverse quantizer 112 inverse quantizes the quantized coefficients that have been input from the quantizer 108. More specifically, the inverse quantizer 112 inverse quantizes the quantized coefficients of the current block in the determined scan order. The inverse quantizer 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transformer 114. The determined scan order can be predetermined.
[0303] (Inverse transformer)
[0304] The inverse transformer 114 recovers the prediction residual by performing an inverse transform on the transform coefficients that have been input from the inverse quantizer 112. More specifically, the inverse transformer 114 recovers the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the recovered prediction residual to the adder 116.
[0305] It should be noted that the recovered prediction residual does not match the prediction residual calculated by the subtractor 104 due to the information typically lost in quantization. In other words, the recovered prediction residual typically includes quantization errors.
[0306] (addition)
[0307] The adder 116 reconstructs the current block by adding the prediction residual that has been input from the inverse transformer 114 and the prediction image that has been input from the prediction controller 128. Thus, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block can also be referred to as a locally decoded block.
[0308] (block memory)
[0309] The block memory 118 is a storage device for storing blocks in the current picture (e.g., for intra prediction). More specifically, the block memory 118 stores the reconstructed image output from the adder 116.
[0310] (frame memory)
[0311] The frame memory 122 is a storage device, e.g., for storing reference pictures used in inter prediction, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.
[0312] (loop filter)
[0313] The loop filter 120 applies a loop filter to the reconstructed image output by the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter used in the encoding loop (loop filter). Examples of the loop filter include, e.g., an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, etc.
[0314] Figure 22 is a block diagram illustrating one example of a functional configuration of the loop filter 120 according to an embodiment. For example, as Figure 22As shown in FIG. 1, the in-loop filter 120 includes a deblocking filter enforcer 120a, an SAO enforcer 120b, and an ALF enforcer 120c. The deblocking filter enforcer 120a performs a deblocking filter process on the reconstructed picture. The SAO enforcer 120b performs an SAO process on the reconstructed picture after the deblocking filter process. The ALF enforcer 120c performs an ALF process on the reconstructed picture after the SAO process. The ALF and the deblocking filter will be described later in detail. The SAO process is a process for enhancing the image quality by reducing ringing (a phenomenon in which pixel values are distorted like a wave around an edge) and correcting bias of pixel values. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that, in some embodiments, the in-loop filter 120 can not include all of the constituent elements disclosed in FIG. 1, and can include some of the constituent elements, and can include additional elements. Additionally, the in-loop filter 120 can be configured to perform the above-described processes in a processing order different from the processing order disclosed in FIG. 1, can not perform all of the processes, and the like. Figure 22 Figure 23A to 23C
[0315] (Adaptive loop filter)
[0316] In the ALF, a least square error filter for removing compression artifacts is applied. For example, one filter is selected from a plurality of filters based on a direction and activity of local gradient, for each 2x2 pixel sub-block in a current block.
[0317] More specifically, first, each sub-block (e.g., each 2x2 pixel sub-block) is classified into one of a plurality of classes (e.g., fifteen or twenty-five classes). The classification of the sub-block can be based on, for example, gradient directionality and activity. In an example, a class index C (e.g., C=5D+A) is calculated or determined based on a gradient directionality D (e.g., 0 to 2 or 0 to 4) and a gradient activity A (e.g., 0 to 4). Then, based on the class index C, each sub-block is classified into one of the plurality of classes.
[0318] The gradient directionality D is calculated, for example, by comparing gradients of a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Further, the gradient activity A is calculated, for example, by adding the gradients of the plurality of directions and quantizing the added result.
[0319] A filter to be used for each sub-block can be determined from a plurality of filters based on a result of such classification.
[0320] A filter shape to be used in the ALF is, for example, a circularly symmetric filter shape. Figure 23A is a conceptual diagram for showing an example of a filter shape used in the ALF. Figure 23B A 5x5 diamond filter is shown, Figure 23C A 7x7 diamond filter is shown, and Figure 23D A 9x9 diamond filter is shown. Information indicating the filter shape is typically signaled at the picture level. It should be noted that signaling such information indicating the filter shape does not necessarily need to be performed at the picture level, but can also be performed at another level (e.g., sequence level, slice level, tile level, CTU level, or CU level).
[0321] For example, the turning on or off of ALF can be determined at the picture level or the CU level. For example, it can be decided at the CU level whether to apply ALF for luma, and it can be decided at the picture level whether to apply ALF for chroma. Information indicating the turning on or off of ALF is typically signaled at the picture level or the CU level. It should be noted that signaling such information indicating the turning on or off of ALF does not necessarily need to be performed at the picture level or the CU level, but can also be performed at another level (e.g., sequence level, slice level, tile level, or CTU level).
[0322] Additionally, as described above, one filter is selected from the plurality of filters, and the ALF process of the subblock is performed. A coefficient set of coefficients to be used for each of the plurality of filters (e.g., up to the fifteenth or twenty-fifth filter) is typically signaled at the picture level. It should be noted that signaling the coefficient set does not necessarily need to be performed at the picture level, and can be performed at another level (e.g., sequence level, slice level, tile level, CTU level, CU level, or subblock level).
[0323] (Cross-component adaptive loop filter)
[0324] Figure 23E is a conceptual diagram for illustrating an example flow of cross-component ALF (CC-ALF). Figure 23D is a conceptual diagram for illustrating filter shapes used in CC-ALF (e.g., Figure 23D of CC-ALF). Figure 23E and Figure 23D Example shaped CC-ALF operates by applying a linear diamond filter to the luma channel of each chroma component. For example, filter coefficients can be sent in an APS, scaled by a factor of 2A10, and rounded for fixed-point representation. For example, in Figure 23F In, Y samples (first component) are used for CCALF for Cb and CCALF for Cr (component different from the first component).
[0325] The application of filters can be controlled on a variable block size and is signaled by a flag received for context decoding for each sample block. The block size, along with the CC-ALF enable flag, can be received at the slice level for each chroma component. CC-ALF can support various block sizes, such as (in chroma samples) 16×16 pixels, 32×32 pixels, 64×64 pixels, and 128×128 pixels.
[0326] (Loop filter > Joint chroma cross-component adaptive loop filter)
[0327] An example of Union Chromaticity-CCALF is in Figure 23G and Figure 23F As shown in the image. Figure 23G This is a conceptual diagram used to illustrate an example process for Joint Chromaticity CCALF. Figure 24 This is a table showing example weighted index candidates. As shown, a CCALF filter is used to generate a CCALF-filtered output as a chroma refinement signal for one color component, while a weighted version of the same chroma refinement signal is applied to another color component. In this way, the complexity of the existing CCALF is reduced by approximately half. The weights can be decoded into a symbol flag and a weighted index. The weighted index (denoted as weight_index) can be decoded into 3 bits, specifying the magnitude of the JC-CCALF weight JcCcWeight, which is a non-zero magnitude. For example, the magnitude of JcCcWeight can be determined as follows:
[0328] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index>>2;
[0329] Otherwise, JcCcWeight equals 4 / (weight_index-4).
[0330] The block-level on / off control for ALF filtering of Cb and Cr can be separate. This is the same as in CCALF, and two separate sets of block-level on / off control flags can be decoded. Unlike CCALF, the block sizes for Cb and Cr on / off control are the same in this paper, so only one block size variable can be decoded.
[0331] (Loop filter > Deblocking filter)
[0332] During the deblocking filtering process, the loop filter 120 performs a filtering process on the block boundaries in the reconstructed image in order to reduce the distortion that occurs at the block boundaries.
[0333] Figure 7 This shows a loop filter 120 used as a deblocking filter (see [link]).Figure 22 and Figure 24 a block diagram of one example of a specific configuration of the deblocking filter enforcer 120a of FIG. 12.
[0334] The deblocking filter enforcer 120a includes a boundary determiner 1201, a filter determiner 1203, a filter enforcer 1205, a process determiner 1208, a filter characteristic determiner 1207, and switches 1202, 1204, and 1206.
[0335] The boundary determiner 1201 determines whether a pixel to be subjected to deblocking filtering (i.e., a current pixel) exists around a block boundary. The boundary determiner 1201 then outputs the determination result to the switch 1202 and the process determiner 1208.
[0336] In a case where the boundary determiner 1201 has determined that the current pixel exists around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1204. In a case where the boundary determiner 1201 has determined that the current pixel does not exist around the block boundary, on the contrary, the switch 1202 outputs the unfiltered image to the switch 1206. It should be noted that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.
[0337] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0338] In a case where the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the filter enforcer 1205. In a case where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel, on the contrary, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the switch 1206.
[0339] When the unfiltered image is obtained through the switches 1202 and 1204, the filter enforcer 1205 performs deblocking filtering on the current pixel with a filter characteristic determined by the filter characteristic determiner 1207. The filter enforcer 1205 then outputs the filtered pixel to the switch 1206.
[0340] The switch 1206 selectively outputs one of a pixel that has not been subjected to deblocking filtering and a pixel that has been subjected to deblocking filtering by the filter enforcer 1205 under the control of the process determiner 1208.
[0341] The processing determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 has determined that the current pixel exists around the block boundary, and when the filter determiner 1203 has determined that deblocking filtering is to be performed on the current pixel, the processing determiner 1208 causes the switch 1206 to output the pixel on which deblocking filtering has been performed. Additionally, the processing determiner 1208 causes the switch 1206 to output the pixel on which deblocking filtering has not been performed, except for the above-described case. By repeating the output of the pixel in this way, a filtered image is output from the switch 1206. It should be noted that, Figure 25 The configuration illustrated in FIG. 12 is one example of the configuration in the deblocking filter executor 120a. The deblocking filter executor 120a can have various configurations.
[0342] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having a symmetric filter characteristic with respect to a block boundary.
[0343] In the deblocking filtering process, one of two deblocking filters having different characteristics (i.e., a strong filter and a weak filter) can be selected using a pixel value and a quantization parameter. In the case of the strong filter, when the pixels p0 to p2 and the pixels q0 to q2 exist across the block boundary, as illustrated in Figure 26 By performing, for example, a calculation according to the following expression, the pixel values of the respective pixels q0 to q2 are changed to the pixel values q'0 to q'2, as illustrated in
[0344] q'0 = (p1 + 2 x p0 + 2 x q0 + 2 x q1 + q2 + 4) / 8
[0345] q'1 = (p0 + q0 + q1 + q2 + 2) / 4
[0346] q'2 = (p0 + q0 + q1 + 3 x q2 + 2 x q3 + 4) / 8
[0347] It should be noted that, in the above expressions, p0 to p2 and q0 to q2 are the pixel values of the respective pixels p0 to p2 and the pixels q0 to q2. Additionally, q3 is the pixel value of the neighboring pixel q3 located on the opposite side of the block boundary from the pixel q2. Additionally, on the right side of each of the expressions, the coefficients multiplied by the respective pixel values of the pixels to be used for deblocking filtering are filter coefficients.
[0348] Further, in the deblocking filtering, clipping can be performed so that the calculated pixel value change does not exceed a threshold value. For example, in the clipping process, the pixel value calculated according to the above expression can be clipped to a value obtained according to "calculated pixel value ± 2 x threshold value", using a threshold value determined based on the quantization parameter. In this way, excessive smoothing can be prevented.
[0349] Figure 27 is a conceptual diagram for illustrating a block boundary on which a deblocking filtering process is performed. Figure 26 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0350] A block boundary on which a deblocking filtering process is performed is, for example, a boundary between CUs, Pus, or TUs having 8x8 pixel blocks, as illustrated in Figure 26 The deblocking filtering process can be performed, for example, in units of four rows or four columns. First, a boundary strength (Bs) value is determined for a block P and a block Q, as illustrated in Figure 27 Figure 27 as indicated in
[0351] According to the Bs value in Figure 27 , it can be determined whether to perform a deblocking filtering process for a block boundary belonging to the same picture using different strengths. When the Bs value is 2, a deblocking filtering process for a chrominance signal is performed. When the Bs value is 1 or more and a determined condition is satisfied, a deblocking filtering process for a luminance signal is performed. The determined condition can be predetermined. It should be noted that the conditions for determining the Bs value are not limited to those indicated in Figure 28 , and the Bs value can be determined based on another parameter.
[0352] (predictor (intra predictor, inter predictor, prediction controller))
[0353] Figure 29 is a flowchart illustrating one example of a process performed by a predictor of the encoder 100. It should be noted that the predictor includes all or a part of the following constituent elements: the intra predictor 124; the inter predictor 126; and the prediction controller 128. The prediction performer includes, for example, the intra predictor 124 and the inter predictor 126.
[0354] The predictor generates a prediction image of the current block (step Sb_1). The prediction image can also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra prediction image (image prediction signal) or an inter prediction image (inter prediction signal). The predictor generates the prediction image of the current block using a reconstructed image that has been obtained through another block by prediction image generation, prediction residual generation, quantized coefficient generation, prediction residual restoration, and prediction image addition.
[0355] The reconstructed image can be, for example, an image in a reference picture, or an image of an encoded block in a current picture that includes the picture of the current block (i.e., the other block described above). The encoded block in the current picture is, for example, a neighboring block of the current block.
[0356] Figure 29 is a flowchart showing another example of a process performed by the predictor of the encoder 100.
[0357] The predictor generates a prediction image using a first method (step Sc la), generates a prediction image using a second method (step Sc lb), and generates a prediction image using a third method (step Sc lc). The first method, the second method, and the third method can be mutually different methods for generating a prediction image. Each of the first method to the third method can be an inter prediction method, an intra prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.
[0358] Next, the prediction processor evaluates the prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 2). For example, the predictor calculates costs C for the prediction images generated in steps Sc la, Sc lb, and Sc lc, and evaluates the prediction images by comparing the costs C of the prediction images. Note that the cost C can be calculated, for example, according to an expression of an R-D optimization model (e.g., C = D + λ x R). In this expression, D represents a compression artifact of a prediction image, and is represented as, for example, a sum of absolute differences between pixel values of a current block and pixel values of a prediction image. Additionally, R represents a bit rate of a stream. Additionally, λ represents a multiplier according to, for example, a method of a Lagrange multiplier.
[0359] Then, the predictor selects one of the prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 3). In other words, the predictor selects a method or a mode for obtaining a final prediction image. For example, the predictor selects a prediction image having a minimum cost C based on the costs C calculated for the prediction images. Alternatively, the evaluation in step Sc 2 and the selection of the prediction image in step Sc 3 can be made based on parameters used in the encoding process. The encoder 100 can transform information for identifying the selected prediction image, method, or mode into a stream. The information can be, for example, a flag or the like. In this way, the decoder 200 is able to generate a prediction image according to the method or the mode selected by the encoder 100 based on the information. Note that in the example shown in Figure 30 In the example shown in FIG. 10, the predictor selects any one of the prediction images after generating the prediction images using the respective methods. However, the predictor can select a method or a mode based on parameters used in the encoding process described above before generating the prediction images, and can generate the prediction images according to the selected method or mode.
[0360] For example, the first method and the second method can be intra prediction and inter prediction, respectively, and the predictor can select a final prediction image for a current block from the prediction images generated according to the prediction methods.
[0361] Figure 31 is a flowchart showing another example of a process performed by the predictor of the encoder 100.
[0362] First, the predictor generates a prediction image using intra prediction (step Sd la), and generates a prediction image using inter prediction (step Sd lb). Note that the prediction image generated by the intra prediction is also referred to as an intra prediction image, and the prediction image generated by the inter prediction is also referred to as an inter prediction image.
[0363] Next, the predictor evaluates each of the intra prediction image and the inter prediction image (step Sd 2). The cost C described above can be used in the evaluation. The predictor can then select the prediction image for which the minimum cost C has been calculated, among the intra prediction image and the inter prediction image, as the final prediction image for the current block (step Sd 3). In other words, the prediction method or mode used to generate the prediction image for the current block is selected.
[0364] (Intra predictor)
[0365] The intra predictor 124 generates a prediction signal (i.e., an intra prediction image) by performing intra prediction (also referred to as intra-frame prediction) of the current block by referring to one or more blocks in the current picture and stored in the block memory 118. More specifically, the intra predictor 124 generates an intra prediction image by performing intra prediction by referring to pixel values (e.g., luma and / or chroma values) of one or more blocks neighboring the current block, and then outputs the intra prediction image to the prediction controller 128.
[0366] For example, the intra predictor 124 performs the intra prediction by using one of a plurality of intra prediction modes that have been defined. The intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes. The defined modes can be predefined.
[0367] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / High Efficiency Video Coding (HEVC) standard.
[0368] The plurality of directional prediction modes include, for example, thirty-three directional prediction modes defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes can include thirty-two directional prediction modes in addition to the thirty-three directional prediction modes (a total of sixty-five directional prediction modes). Figure 31is a conceptual diagram showing a total of sixty-seven kinds of intra prediction modes (two kinds of non-directional prediction modes and sixty-five kinds of directional prediction modes) that can be used in intra prediction. The solid arrows indicate thirty-three directions defined in the H.265 / HEVC standard, and the dotted arrows indicate additional thirty-two directions Figure 32
[0369] In various types of processing examples, a luma block can be referenced in intra prediction of a chroma block. In other words, a chroma component of a current block can be predicted based on a luma component of the current block. Such intra prediction is also referred to as cross-component linear model (CCLM) prediction. An intra prediction mode for a chroma block in which such a luma block is referenced (also referred to as, for example, a CCLM mode) can be added as one of the intra prediction modes for the chroma block.
[0370] The intra predictor 124 can correct pixel values of intra prediction based on horizontal / vertical reference pixel gradients. Intra prediction accompanied by such correction is also referred to as position-dependent intra prediction combination (PDPC). Information indicating whether to apply PDPC (for example, referred to as a PDPC flag) is generally signaled at the CU level. Note that signaling such information does not necessarily need to be performed at the CU level, and can be performed at another level (for example, a sequence level, a picture level, a slice level, a slice segment level, or a CTU level).
[0371] Figure 33 is a flowchart showing one example of a process performed by the intra predictor 124.
[0372] The intra predictor 124 selects one of a plurality of intra prediction modes (step Sw_1). The intra predictor 124 then generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra predictor 124 determines a most probable mode (MPM) (step Sw_3). The MPM includes, for example, six kinds of intra prediction modes. For example, two of the six kinds of intra prediction modes can be a planar mode and a DC prediction mode, and the other four modes can be directional prediction modes. The intra predictor 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).
[0373] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra predictor 124 sets an MPM flag to 1 (step Sw_5), and generates information indicating the selected intra prediction mode in the MPM (step Sw_6). Note that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110.
[0374] When it is determined that the selected intra prediction mode is not included in the MPM (NO in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flag. The intra predictor 124 then generates information indicating the selected intra prediction mode among the at least one intra prediction mode that is not included in the MPM (step Sw_8). Note that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded as the prediction parameter by the entropy encoder 110. The information indicating the intra prediction mode indicates, for example, any one of 0 to 60.
[0375] (inter-predictor)
[0376] The inter-predictor 126 generates a predicted image (inter-predicted image) by performing inter-prediction (also referred to as inter-frame prediction) on the current block by referring to one or more blocks in a reference picture that is different from the current picture and is stored in the frame memory 122. Inter-prediction is performed in units of the current block or a current sub-block (e.g., 4x4 block) in the current block. A sub-block is included in a block and is a unit smaller than the block. The size of the sub-block can be in the form of a slice, a tile, a picture, or the like.
[0377] For example, the inter-predictor 126 performs motion estimation in the reference picture for the current block or the current sub-block and finds a reference block or a reference sub-block that best matches the current block or the current sub-block. The inter-predictor 126 then obtains motion information (e.g., a motion vector) that compensates for motion or change from the reference block or the reference sub-block to the current block or the sub-block. The inter-predictor 126 generates an inter-predicted image of the current block or the sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter-predictor 126 outputs the generated inter-predicted image to the prediction controller 128.
[0378] The motion information used in the motion compensation can be signaled as the inter-prediction signal in various forms. For example, a motion vector can be signaled. As another example, a difference between a motion vector and a motion vector predictor can be signaled.
[0379] (reference picture list)
[0380] Figure 34 is a conceptual diagram for illustrating an example of a reference picture. Figure 33 is a conceptual diagram for illustrating an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. Note that in the example of the reference picture list illustrated in FIG. 8, the reference picture list is a list indicating two reference pictures. However, the reference picture list can be a list indicating one or more reference pictures. Figure 33In the diagram, each image within a rectangle indicates its reference relationship, each arrow indicates its reference relationship, the horizontal axis indicates time, and the I, P, and B symbols within the rectangles indicate intra-frame predicted images, single-predicted images, and double-predicted images, respectively. The numbers within the rectangles also indicate the decoding order. For example... Figure 34 As shown, the decoding order of the images is I0, P1, B2, B3, and B4, and the display order of the images is I0, B3, B2, B4, and P1. Figure 33 As shown, the reference image list is a list representing candidate reference images. For example, an image (or slice) may include at least one reference image list. For instance, one reference image list is used when the current image is a single-prediction image, and two reference image lists are used when the current image is a double-prediction image. Figure 34 and Figure 34 In the example, image B3, which is the current image currPic, has two lists of reference images: the L0 list and the L1 list. When the current image currPic is image B3, the candidate reference images for the current image currPic are I0, P1, and B2, and the list of reference images (i.e., the L0 list and the L1 list) indicates these images. The inter-frame predictor 126 or the prediction controller 128 specifies which image in each list to actually reference in the form of a reference image index refidxLx. Figure 35 In the text, reference images P1 and B2 are specified by reference image indices refIdxL0 and refIdxL1.
[0381] Such a list of reference images can be generated for each unit, such as a sequence, picture, slice, brick, CTU, or CU. Furthermore, the reference image index indicating the reference image to be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, brick level, CTU level, or CU level. Additionally, a common list of reference images can be used in multiple inter-frame prediction modes.
[0382] (Basic process of inter-frame prediction)
[0383] Figure 9 This is a flowchart illustrating the basic processing flow of an example of inter-frame prediction.
[0384] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates the difference between the current block and the prediction image as the prediction residual (step Se_4).
[0385] Here, when generating the prediction image, the inter prediction unit 126 generates the prediction image by determining a motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Further, in determining the MV, the inter prediction unit 126 determines the MV by selection of a motion vector candidate (MV candidate) (step Se_1) and derivation of the MV (step Se_2). The selection of the MV candidate is performed by, for example, the inter prediction unit 126 generating a list of MV candidates and selecting at least one MV candidate from the list of MV candidates. Note that a past-derived MV can be added to the list of MV candidates. Alternatively, in deriving the MV, the inter prediction unit 126 can also select at least one MV candidate from the at least one MV candidate and determine the selected at least one MV candidate as the MV for the current block. Alternatively, the inter prediction unit 126 can determine the MV for the current block by performing estimation in a reference picture region specified by each of the selected at least one MV candidate. Note that the estimation in the reference picture region can be referred to as motion estimation.
[0386] Additionally, although in the example described above, the steps Se_1 to Se_3 are performed by the inter prediction unit 126, the processes of, for example, the step Se_1, the step Se_2, and the like can be performed by another constituent element included in the encoder 100.
[0387] Note that a list of MV candidates can be generated for each process in the inter prediction mode, or a common list of MV candidates can be used in multiple inter prediction modes. The processes in the steps Se_3 and Se_4 correspond to the processes in the steps Sa_3 and Sa_4, respectively, illustrated in FIG. 8. Figure 30 Figure 36
[0388] (Motion vector derivation flow)
[0389] Figure 37 is a flowchart illustrating one example of a process of derivation of a motion vector.
[0390] The inter prediction unit 126 can derive the MV of the current block in a mode in which motion information (e.g., the MV) is encoded. In this case, for example, the motion information can be encoded as a prediction parameter and can be signaled. In other words, the encoded motion information is included in a stream.
[0391] Alternatively, the inter prediction unit 126 can derive the MV in a mode in which the motion information is not encoded. In this case, the motion information is not included in the stream.
[0392] Here, the MV derivation mode can include a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, and the like, which are described later. The modes in which the motion information is coded include the normal inter mode, the affine mode (specifically, an affine inter mode and an affine merge mode), and the like. Note that the motion information can include not only the MV but also motion vector predictor selection information, which is described later. The modes in which the motion information is not coded include the FRUC mode, and the like. The inter predictor 126 selects a mode for deriving the MV of the current block from among the plurality of modes, and derives the MV of the current block using the selected mode.
[0393] Figure 38A is a flowchart illustrating another example of the derivation of the motion vector.
[0394] The inter predictor 126 can derive the MV for the current block in a mode in which the MV difference is coded. In this case, for example, the MV difference can be coded as a prediction parameter, and can be signaled. In other words, the coded MV difference is included in the stream. The MV difference is a difference between the MV of the current block and the MV predictor. Note that the MV predictor is a motion vector predictor.
[0395] Alternatively, the inter predictor 126 can derive the MV in a mode in which the MV difference is not coded. In this case, the coded MV difference is not included in the stream.
[0396] Here, as described above, the MV derivation mode includes a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, and the like, which are described later. The modes in which the MV difference is coded include the normal inter mode, the affine mode (specifically, an affine inter mode), and the like. The modes in which the MV difference is not coded include the FRUC mode, the normal merge mode, the affine mode (specifically, an affine merge mode), and the like. The inter predictor 126 selects a mode for deriving the MV of the current block from among the plurality of modes, and derives the MV of the current block using the selected mode.
[0397] (Motion vector derivation mode)
[0398] Figure 38B and Figure 38A is a conceptual diagram for illustrating an example classification of the modes for the MV derivation. For example, as illustrated in Figure 38B , the MV derivation mode is roughly classified into three modes according to whether or not the motion information is coded and whether or not the MV difference is coded. The three modes are an inter mode, a merge mode, and a frame rate up conversion (FRUC) mode. The inter mode is a mode in which motion estimation is performed, and in which the motion information and the MV difference are coded. For example, as illustrated in Figure 38BAs illustrated in FIG. 1, the inter mode includes an affine inter mode and a normal inter mode. The merge mode is a mode in which motion estimation is not performed and in which an MV is selected from coded surrounding blocks and the MV is used to derive an MV for a current block. The merge mode is a mode in which motion information is substantially coded and an MV difference is not coded. For example, as illustrated in FIG. 1, the MV difference is not coded in the merge mode. Figure 38A As illustrated in FIG. 1, the merge mode includes a normal merge mode (also referred to as a normal merge mode or a regular merge mode), a motion vector difference merge (MMVD) mode, a combined inter merge / intra prediction (CIIP) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, among the modes included in the merge mode, the MV difference is exceptionally coded in the MMVD mode. Note that the affine merge mode and the affine inter mode are modes included in the affine mode. The affine mode is a mode for deriving an MV of each of a plurality of sub-blocks included in a current block as an MV of the current block in a case where an affine transformation is assumed. The FRUC mode is a mode for deriving an MV of a current block by performing estimation between coded regions, and in which neither motion information nor any MV difference is coded. Note that the respective modes will be described later in more detail.
[0399] Note that, Figure 38B and Figure 39 The classification of the modes illustrated in FIG. 1 is an example, and the classification is not limited thereto. For example, when the MV difference is coded in the CIIP mode, the CIIP mode is classified as the inter mode.
[0400] (MV derivation > normal inter mode)
[0401] The normal inter mode is an inter prediction mode for deriving an MV of a current block from a reference picture region specified by an MV candidate based on a block similar to an image of the current block. In this normal inter mode, the MV difference is coded.
[0402] Figure 40 is a flowchart illustrating an example of a procedure of inter prediction in the normal inter mode.
[0403] First, the inter predictor 126 obtains a plurality of MV candidates for the current block based on information (for example, MVs of a plurality of coded blocks surrounding the current block in time or space) (step Sg_1). In other words, the inter predictor 126 generates an MV candidate list.
[0404] Next, the inter predictor 126 extracts N (an integer of 2 or more) MV candidates as motion vector predictor candidates (also referred to as MV predictor candidates) from the plurality of MV candidates obtained in step Sg_1 according to the determined priority order (step Sg_2). Note that the priority order can be determined in advance for each of the N MV candidates.
[0405] Next, the inter predictor 126 selects one of the N prediction motion vector candidates as a motion vector predictor (also referred to as an MV predictor) of the current block (step Sg_3). At this time, the inter predictor 126 encodes motion vector predictor selection information for identifying the selected motion vector predictor in the stream. In other words, the inter predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0406] Next, the inter predictor 126 derives an MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter predictor 126 also encodes, in the stream, a difference between the derived MV and the motion vector predictor as an MV difference. In other words, the inter predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130. Note that the encoded reference picture is a picture including a plurality of blocks that have been reconstructed after encoding.
[0407] Finally, the inter predictor 126 generates a prediction image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a slice, the inter prediction of the slice using the normal inter mode ends. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a picture, the inter prediction of the picture using the normal inter mode ends. Note that not all blocks included in a slice can undergo the processes in steps Sg_1 to Sg_5, and the inter prediction of the slice using the normal inter mode can end when a part of the blocks undergoes the processes. The same applies to the processes in steps Sg_1 to Sg_5. The inter prediction of the picture using the normal inter mode can end when the processes are performed for a part of the blocks in the picture.
[0408] Note that the prediction image is an inter prediction signal as described above. Additionally, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the prediction image is, for example, encoded as a prediction parameter in the encoded signal.
[0409] Note that the MV candidate list can also be used as a list used in other modes. Additionally, a process related to the MV candidate list can be applied to a process related to a list used in another mode. Processes related to the MV candidate list include, for example, extracting or selecting an MV candidate from the MV candidate list, reordering the MV candidates, or deleting an MV candidate.
[0410] (MV derivation > normal merge mode)
[0411] The normal merge mode is an inter prediction mode for selecting an MV candidate as an MV of a current block from a MV candidate list, thereby deriving the MV. Note that the normal merge mode is a type of the merge mode, and can be simply referred to as the merge mode. In this embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a broader sense.
[0412] Figure 41 is a flowchart showing an example of inter prediction in the normal merge mode.
[0413] First, the inter predictor 126 obtains a plurality of MV candidates for the current block based on information (e.g., MVs of a plurality of encoded blocks around the current block in time or space) (step Sh_1). In other words, the inter predictor 126 generates a MV candidate list.
[0414] Next, the inter predictor 126 selects one MV candidate from the plurality of MV candidates obtained in step Sh_1, thereby deriving an MV of the current block (step Sh_2). At this time, the inter predictor 126 encodes MV selection information for identifying the selected MV candidate in a stream. In other words, the inter predictor 126 outputs the MV selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0415] Finally, the inter predictor 126 generates a prediction image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a slice, inter prediction of the slice ends using the normal merge mode. Additionally, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a picture, inter prediction of the picture ends using the normal merge mode. Note that not all blocks included in a slice can undergo the processes of steps Sh_1 to Sh_3, and when a part of the blocks undergo the processes, inter prediction of the slice using the normal merge mode can end. This also applies to the processes in steps Sh_1 to Sh_3. When the processes are performed for a part of the blocks in a picture, inter prediction of the picture using the normal merge mode can end.
[0416] Additionally, for example, information indicating an inter prediction mode for generating a prediction image (normal merge mode in the above example) and included in the encoded signal is encoded as a prediction parameter in the stream.
[0417] Figure 41 is a conceptual diagram for showing one example of a motion vector derivation process by normal merge mode on a current picture.
[0418] First, the inter predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of the MV candidates include: spatially neighboring MV candidates, which are MVs of a plurality of encoded blocks spatially located around a current block; temporally neighboring MV candidates, which are MVs of blocks around a position on which a current block in an encoded reference picture is projected; combined MV candidates, which are MVs generated by combining MV values of spatially neighboring MV predictors and MV values of temporally neighboring MV predictors; and zero MV candidates, which are MVs having zero values.
[0419] Next, the inter predictor 126 selects one MV candidate from among the plurality of MV candidates registered in the MV candidate list and determines the MV candidate as an MV of the current block.
[0420] Further, the entropy encoder 110 writes and encodes merge_idx, which is a signal indicating which MV candidate has been selected, in the stream.
[0421] Note that the MV candidates registered in the MV candidate list described in Figure 38B may be examples. The number of MV candidates can be different from that of the MV candidates in the drawing, and the MV candidate list can be configured in such a manner that some of the categories of the MV candidates in the drawing can not be included, or one or more MV candidates other than the categories of the MV candidates in the drawing are included.
[0422] The final MV can be determined by performing dynamic motion vector refresh (DMVR) described later using the MV of the current block derived by the normal merge mode. Note that, in the normal merge mode, motion information is encoded and MV difference is not encoded. In the MMVD mode, an MV candidate is selected from the MV candidate list as in the case of the normal merge mode, and the MV difference is encoded. As shown in Figure 42 MMVD can be classified as a merge mode together with the normal merge mode as shown in Note that the MV difference in the MMVD mode does not always need to be the same as the MV difference used in the inter mode. For example, MV difference derivation in the MMVD mode can be a process that requires a smaller amount of processing than that required for MV difference derivation in the inter mode.
[0423] Additionally, a combined inter merge / intra prediction (CIIP) mode can be performed. This mode is used to overlap a prediction image generated in inter prediction and a prediction image generated in intra prediction to generate a prediction image for a current block.
[0424] Note that the MV candidate list can be referred to as a candidate list. Additionally, merge_idx is MV selection information.
[0425] (MV derivation > HMVP mode)
[0426] Figure 42 is a conceptual diagram for showing one example of an MV derivation process for a current picture using the HMVP merge mode.
[0427] In the normal merge mode, an MV for a CU, for example, as a current block is determined by selecting one MV candidate from a list of MVs generated from a reference coded block (e.g., a CU). Here, another MV candidate can be registered in the MV candidate list. The mode in which such another MV candidate is registered is referred to as the HMVP mode.
[0428] In the HMVP mode, a first-in first-out (FIFO) server of the HMVP is used to manage MV candidates, separately from the MV candidate list for the normal merge mode.
[0429] In the FIFO buffer, first, motion information such as MVs of past processed blocks is newly stored. In managing the FIFO buffer, every time a block is processed, an MV for a newest block (i.e., a CU that was just processed before) is stored in the FIFO buffer, and an MV of an oldest CU (i.e., a CU that was processed earliest) is deleted from the FIFO buffer. In Figure 42 In the example shown in, HMVP1 is an MV for a newest block, and HMVP5 is an MV for an oldest MV.
[0430] Then, for example, the inter predictor 126 checks whether each MV managed in the FIFO buffer is a different MV from all MV candidates that have been registered in the MV candidate list for the normal merge mode, starting from HMVP1. When it is determined that the MV is different from all MV candidates, the inter predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for the normal merge mode as an MV candidate. At this time, one or more of the MV candidates in the FIFO buffer can be registered (added to the MV candidate list).
[0431] By using the HMVP mode in this way, not only MVs of blocks neighboring the current block in space or time can be added, but also MVs of blocks processed in the past can be added. As a result, the variation of MV candidates for the normal merge mode is enlarged, which increases the possibility that coding efficiency can be improved.
[0432] Note that the MV can be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer can include not only the MV value, but also reference picture information, reference direction, picture number, and the like. Additionally, the block can be, for example, a CU.
[0433] Note that, Figure 42 The MV candidate list and the FIFO buffer shown in FIG. 10 are examples. The size of the MV candidate list and the FIFO buffer can be different from Figure 42 in FIG. 10, or can be configured to register MV candidates in an order different from Figure 43 in FIG. 10. Additionally, the processes described here can be common between the encoder 100 and the decoder 200.
[0434] Note that the HMVP mode can be applied to modes other than the normal merge mode. For example, motion information such as MVs of blocks processed in the past in the affine mode can also be stored first and can be used as MV candidates, which can better promote efficiency. A mode obtained by applying the HMVP mode to the affine mode can be referred to as a history affine mode.
[0435] (MV derivation > FRUC mode)
[0436] Motion information can be derived at the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation at the decoder 200 side. In an embodiment, at the decoder side, motion estimation is performed without using any pixel value in the current block. Modes for performing motion estimation at the decoder 200 side without using any pixel value in the current block include a frame rate up conversion (FRUC) mode, a pattern matching motion vector derivation (PMMVD) mode, and the like.
[0437] Figure 44 One example of a FRUC process in the form of a flowchart is shown in FIG. 11. First, a list indicates MVs of encoded blocks (each of the encoded blocks is spatially or temporally adjacent to the current block) as MV candidates by referring to the MVs (the list can be the MV candidate list, or can be used as the MV candidate list for the normal merge mode) (step Si_1).
[0438] Next, a best MV candidate is selected from the plurality of MV candidates registered in the MV candidate list (step Si_2). For example, evaluation values of the respective MV candidates included in the MV candidate list are calculated, and one MV candidate is selected based on the evaluation values. Based on the selected motion vector candidate, a motion vector for the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector for the current block. Alternatively, for example, the motion vector for the current block can be derived using pattern matching in a surrounding area of a position in a reference picture, where the position in the reference picture corresponds to the selected motion vector candidate. In other words, estimation using pattern matching and evaluation values can be performed in a surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV for the current block. In some embodiments, the update of the motion vector that produces the better evaluation value can not be performed.
[0439] Finally, the inter predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed on each block. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in a slice, the inter prediction of the slice ends using the FRUC mode. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in a picture, the inter prediction of the picture ends using the FRUC mode. Note that not all blocks included in a slice can undergo the processes of steps Si_1 to Si_5, and the inter prediction of the slice using the FRUC mode can end when a portion of the blocks undergo the processes. The inter prediction of the picture using the FRUC mode can end when the processes in steps Si_1 to Si_5 are performed on a portion of the blocks included in the picture in a similar manner.
[0440] Similar processes can be performed in units of sub-blocks.
[0441] The evaluation values can be calculated according to various types of methods. For example, a comparison is made between a reconstructed image of a region in a reference picture corresponding to a motion vector and a reconstructed image in a determined region (which can be, for example, a region in another reference picture or a region in a neighboring block of the current picture, as indicated below). The determined region can be predetermined.
[0442] A difference between pixel values of the two reconstructed images can be used for the evaluation value of the motion vector. Note that information other than the value of the difference can be used to calculate the evaluation value.
[0443] Next, an example of pattern matching is described in detail. First, one MV candidate included in an MV candidate list (e.g., merge list) is selected as a starting point of estimation by pattern matching. For example, first pattern matching or second pattern matching can be used as pattern matching. The first pattern matching and the second pattern matching can be referred to as bilateral matching and template matching, respectively.
[0444] (MV derivation > FRUC > bilateral matching)
[0445] In the first pattern matching, pattern matching is performed between two blocks positioned along a motion trajectory of a current block and included in two different reference pictures. Accordingly, in the first pattern matching, a region along a motion trajectory of a current block in another reference picture is used as a determined region for calculating an evaluation value of a candidate described above. The determined region can be predetermined.
[0446] Figure 44 is a conceptual diagram for illustrating one example of the first pattern matching (bilateral matching) between two blocks along a motion trajectory in two reference pictures. As illustrated in Figure 45 In the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating a pair that best matches among pairs in two blocks included in two different reference pictures (Ref0, Ref1) and positioned along a motion trajectory of a current block (Cur block). More specifically, a difference between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by an MV candidate and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scanning MV candidates with a display time interval is derived for a current block, and a value of the obtained difference is used to calculate an evaluation value. An MV candidate that produces a best evaluation value and possibly a good result can be selected from among a plurality of MV candidates as a final MV.
[0447] Under the assumption of a continuous motion trajectory, motion vectors (MV0, MV1) of two reference blocks are proportional to temporal distances (TD0, TD1) between a current picture (CurPic) and two reference pictures (Ref0, Ref1). For example, when a current picture is located between two reference pictures in time and temporal distances from the current picture to the respective two reference pictures are equal to each other, bidirectional motion vectors that are mirror-symmetric are derived in the first pattern matching.
[0448] (MV derivation > FRUC > template matching)
[0449] In the second pattern matching (template matching), pattern matching is performed between a block in the reference picture and a template in the current picture (the template is a block in the current picture that is adjacent to the current block (the adjacent block is, for example, an upper and / or left adjacent block)). Thus, in the second pattern matching, the block in the current picture that is adjacent to the current block is used as a determination region for calculating the evaluation value of the MV candidate described above.
[0450] Figure 45 is a conceptual diagram for illustrating one example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As Figure 46A In the second pattern matching, as illustrated in
[0451] Such information indicating whether to apply the FRUC mode (for example, referred to as a FRUC flag) can be signaled at the CU level. Additionally, when the FRUC mode is applied (for example, when the FRUC flag is true), information indicating the applicable pattern matching method (for example, the first pattern matching or the second pattern matching) can be signaled at the CU level. Note that signaling such information does not necessarily need to be performed at the CU level, can be performed at another level (for example, a sequence level, a picture level, a slice level, a tile level, a CTU level, or a sub-block level).
[0452] (MV derivation > affine mode)
[0453] The affine mode is a mode for generating an MV using an affine transformation. For example, an MV can be derived in sub-block units based on motion vectors of a plurality of adjacent blocks. This mode is also referred to as an affine motion compensation prediction mode.
[0454] Figure 46A is a conceptual diagram for illustrating one example of MV derivation in sub-block units based on motion vectors of a plurality of adjacent blocks. In Figure 46BIn this context, the current block comprises, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 at the top-left control point in the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V1 at the top-right control point in the current block is derived based on the motion vectors of adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vector (v1) for the corresponding sub-blocks in the current block can be derived. x ,v y ).
[0455] [Mathematics 1]
[0456]
[0457] Here, x and y indicate the horizontal and vertical positions of the sub-block, respectively, and w indicates the determined weighting coefficients. The determined weighting coefficients can be predetermined.
[0458] This information indicating an affine mode (e.g., referred to as an affine flag) can be signaled at the CU level. It should be noted that signaling information indicating an affine mode does not necessarily need to be performed at the CU level, and can also be performed at another level (e.g., sequence level, picture level, slice level, fragment level, CTU level, or sub-block level).
[0459] Additionally, affine modes can include several modes for different methods of deriving motion vectors at the top-left and top-right control points. For example, affine modes include two modes: affine inter-frame mode (also known as affine normal inter-frame mode) and affine merge mode.
[0460] (MV derivation > Affine mode)
[0461] Figure 46B This is a conceptual diagram used to illustrate an example of MV derivation in a sub-block manner within an affine pattern that uses three control points. Figure 47A In this context, the current block comprises, for example, sixteen 4×4 blocks. Here, the motion vector V0 at the top-left control point of the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector V1 at the top-right control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V2 at the bottom-left control point of the current block is derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v1, v2, v2) for the corresponding sub-blocks in the current block can be derived. x ,v y ).
[0462] [Mathematics 2]
[0463]
[0464] Here, x and y indicate a horizontal position and a vertical position of a sub-block, respectively, and w and h can be weighting factors, which can be predetermined weighting factors. In an embodiment, w can indicate a width of the current block, and h can indicate a height of the current block.
[0465] The affine mode in which different numbers of control points (e.g., two and three control points) are used can be switched and signaled at a CU level. Note that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., a sequence level, a picture level, a slice level, a slice segment level, a CTU level, or a sub-block level).
[0466] Additionally, such an affine mode in which three control points are used can include different methods for deriving motion vectors at the top-left control point, the top-right control point, and the bottom-left control point. For example, as in the case of the affine mode in which two control points are used, the affine mode in which three control points are used can include both an affine inter mode and an affine merge mode.
[0467] Note that the size of each sub-block included in the current block can not be limited to 4x4 pixels in the affine mode, and can also be another size. For example, the size of each sub-block can be 8x8 pixels.
[0468] (MV derivation > affine mode > control point)
[0469] Figure 47B , Figure 47C and Figure 47A are conceptual diagrams for illustrating an example of MV derivation at a control point in the affine mode.
[0470] As illustrated in Figure 47B , in the affine mode, for example, a motion vector predictor at a corresponding control point of the current block is calculated based on a plurality of motion vectors corresponding to the encoded blocks A (left), B (above), C (top-right), D (bottom-left), and E (top-left) adjacent to the current block, in accordance with the affine mode being encoded. More specifically, the encoded blocks A (left), B (above), C (top-right), D (bottom-left), and E (top-left) are examined in the listed order, and a first valid block encoded in accordance with the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on the plurality of motion vectors corresponding to the identified block.
[0471] For example, as Figure 47CAs shown in FIG. 2, when a block A adjacent to the current block on the left has been coded according to the affine mode in which two control points are used, motion vectors v3 and v4 projected at the top-left position and the top-right position of the coded block including the block A are derived. Then, the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block are calculated according to the derived motion vectors v3 and v4.
[0472] For example, as Figure 47A to 47C As shown in FIG. 3, when a block A adjacent to the current block on the left has been coded according to the affine mode in which three control points are used, motion vectors v3, v4 and v5 projected at the top-left position, the top-right position and the bottom-left position of the coded block including the block A are derived. Then, the motion vector v0 at the top-left control point of the current block, the motion vector v1 at the top-right control point of the current block and the motion vector v2 at the bottom-left control point of the current block are calculated according to the derived motion vectors v3, v4 and v5.
[0473] Figure 50 The MV derivation method shown in FIG. 1 can be used for Figure 51 The MV derivation at each control point of the current block in step Sk_1 shown in FIG. 2, or can be used for the MV predictor derivation at each control point of the current block in step Sj_1 described later. Figure 48A The MV predictor derivation at each control point of the current block in step Sj_1 shown in FIG. 3.
[0474] Figure 48B and Figure 48A is a conceptual diagram for showing an example of the MV derivation at the control points in the affine mode.
[0475] Figure 48A is a conceptual diagram for showing an example affine mode in which two control points are used.
[0476] In the affine mode, as Figure 48B As shown in FIG. 2, a MV selected from MVs at coded blocks A, B and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Likewise, a MV selected from MVs at coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block.
[0477] Figure 48B is a conceptual diagram for showing an example affine mode in which three control points are used.
[0478] In the affine mode, as Figure 48AThe MV selected from the MVs of the encoded blocks A, B and C adjacent to the current block as shown in FIG. 10 is used as the motion vector v0 at the top-left corner control point of the current block. Likewise, the MV selected from the MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right corner control point of the current block. In addition, the MV selected from the MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v2 at the bottom-left corner control point of the current block.
[0479] Note that, Figure 48B and Figure 50 The MV derivation method shown in Figure 51 The MV predictor derivation at each control point of the current block in step Sk_1 shown in Figure 49A The MV predictor derivation at each control point of the current block in step Sk_1 shown in
[0480] Here, when the affine mode in which different numbers of control points (e.g., two and three control points) are used can be switched and signaled at the CU level, the number of control points of the encoded block and the number of control points of the current block can be different from each other.
[0481] Figure 49B and Figure 49A are conceptual diagrams for illustrating examples of the method for MV derivation at the control points when the number of control points of the encoded block and the number of control points of the current block are different from each other.
[0482] For example, as shown in Figure 49B The current block has three control points at the top-left corner, the top-right corner and the bottom-left corner, and the block A adjacent to the current block on the left has been encoded according to the affine mode in which two control points are used. In this case, the motion vectors v3 and v4 projected at the top-left corner position and the top-right corner position in the encoded block including the block A are derived. Then the motion vector v0 at the top-left corner control point and the motion vector v1 at the top-right corner control point of the current block are calculated according to the derived motion vectors v3 and v4. In addition, the motion vector v2 at the bottom-left corner control point is calculated according to the derived motion vectors v0 and v1.
[0483] For example, as shown in Figure 49AAs shown, the current block has two control points at the top left and top right corners, and the block A adjacent to the current block on the left has been encoded according to an affine pattern using three control points. In this case, motion vectors v3, v4, and v5 projected onto the top left, top right, and bottom left corners of the encoded block, including block A, are derived. Then, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated based on the derived motion vectors v3, v4, and v5.
[0484] Notice, Figure 49B and Figure 50 The MV derivation method shown can be used in the descriptions that follow. Figure 51 The MV derivation at each control point of the current block in step Sk_1 shown in the figure, or can be used for the description later. Figure 50 The MV predictor derivation at each control point of the current block in step Sj_1 is shown in the figure.
[0485] (MV Derivation > Affine Pattern > Affine Merging Pattern)
[0486] Figure 46A This is a flowchart illustrating an example of a process in affine merging mode.
[0487] In the affine merging mode as shown, firstly, the inter-frame predictor 126 derives the MV at the corresponding control point of the current block (step Sk_1). Figure 46B As shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 47A to 47C As shown, the control points are the top-left corner, top-right corner, and bottom-left corner of the current block. The inter-frame predictor 126 can encode MV selection information to identify two or three derived MVs in the stream.
[0488] For example, when using When the MV derivation method is shown in the figure, such as Figure 47A As shown, the inter-frame predictor 126 examines the encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in the listed order and identifies the first valid block encoded according to the affine pattern.
[0489] Inter-frame predictor 126 uses the first valid block identified, encoded according to the identified affine pattern, to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47BAs shown in FIG. 12, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block from the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block including the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block by projecting the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block onto the current block.
[0490] Alternatively, when the block A is identified and the block A has three control points, as shown in FIG. 13, the MVs at the three control points can be calculated, and as described above Figure 47C As shown in FIG. 12, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block from the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block including the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block by projecting the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block onto the current block.
[0491] Note that, as described above Figure 49A As shown in FIG. 12, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block from the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block including the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector vl at the top-right corner control point of the current block by projecting the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block onto the current block. Figure 49B As shown in FIG. 13, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block, the motion vector vl at the top-right corner control point of the current block, and the motion vector v2 at the bottom-left corner control point of the current block from the motion vectors v3, v4, and v5 at the top-left corner, the top-right corner, and the bottom-left corner of the coded block including the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block, the motion vector vl at the top-right corner control point of the current block, and the motion vector v2 at the bottom-left corner control point of the current block by projecting the motion vectors v3, v4, and v5 at the top-left corner, the top-right corner, and the bottom-left corner of the coded block onto the current block.
[0492] Next, the inter predictor 126 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter predictor 126 calculates the MVs for each of the plurality of sub-blocks as affine MVs using two motion vectors v0 and vl and the above expression (1A) or using three motion vectors v0, vl, and v2 and the above expression (IB) (step Sk_2). The inter predictor 126 then performs motion compensation on the sub-blocks using these affine MVs and the coded reference picture (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all the sub-blocks included in the current block, the process of generating a prediction image using the affine merge mode for the current block ends. In other words, motion compensation on the current block is performed to generate a prediction image of the current block.
[0493] Note that the MV candidate list described above can be generated in step Sk_1. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, any combination of Figures 47A-47C the MV derivation method illustrated in Figure 48A and Figure 48B the MV derivation method illustrated in Figure 49A and Figure 49B the MV derivation method illustrated in and other MV derivation methods.
[0494] Note that the MV candidate list can include MV candidates in modes in which prediction is performed in units of sub-blocks other than the affine mode.
[0495] Note that, for example, an MV candidate list including MV candidates in the affine merge mode in which two control points are used and the affine merge mode in which three control points are used can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode in which two control points are used and an MV candidate list including MV candidates in the affine merge mode in which three control points are used can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode in which two control points are used and the affine merge mode in which three control points are used can be generated. The MV candidate(s) can be, for example, MVs for the coded block A (left), block B (above), block C (upper right), block D (lower left), and block E (upper left), or MVs for the effective blocks in the blocks.
[0496] Note that an index indicating one of the MVs in the MV candidate list can be transmitted as the MV selection information.
[0497] (MV derivation > affine mode > affine inter mode)
[0498] Figure 51 is a flowchart illustrating one example of a procedure in the affine inter mode.
[0499] In the affine inter mode, first, the inter predictor 126 derives MV predictors (v0, v1) or (v0, v1, v2) of the respective two or three control points of the current block (step Sj_1). The control points can be, for example, the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block, as illustrated in Figure 46A or Figure 46B .
[0500] For example, when the MV derivation methods illustrated in Figure 48A and Figure 48B are used, the inter predictor 126 derives the MV predictor (v0, v1) or (v0, v1, v2) of the control point by selecting, for example, the MVs of the blocks A, B, C, D, and E in the MV candidate list.Figure 48A or Figure 48B the MVs of any of the encoded blocks in the vicinity of the respective control points of the current block, to derive the MV predictors (v0, v1) or (v0, v1, v2) at the respective two or three control points of the current block. At this time, the inter predictor 126 encodes in the stream the MV predictor selection information for identifying the selected two or three MV predictors.
[0501] For example, the inter predictor 126 can determine the blocks from which the MVs are selected as the MV predictors at the control points using cost evaluation or the like from the encoded blocks adjacent to the current block, and can write in the bitstream a flag indicating which MV predictor has been selected. In other words, the inter predictor 126 outputs the MV predictor selection information such as the flag as the prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0502] Next, the inter predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictors selected or derived in step Sj_1 (step Sj_2). In other words, the inter predictor 126 calculates the MVs of each sub-block corresponding to the updated MV predictors as affine MVs using the expression (1A) or the expression (1B) described above (step Sj_3). The inter predictor 126 then performs motion compensation on the sub-blocks using these affine MVs and the encoded reference picture (step Sj_4). The processes in steps Sj_3 and Sj_4 are performed on all the blocks in the current block when the MV predictors are updated in step Sj_2. As a result, for example, the inter predictor 126 determines the MV predictor that results in the minimum cost in the motion estimation loop as the MV at the control points (step Sj_5). At this time, the inter predictor 126 also encodes in the stream the difference between the determined MV and the MV predictor as the MV difference. In other words, the inter predictor 126 outputs the MV difference as the prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0503] Finally, the inter predictor 126 generates the prediction image for the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0504] Note that the MV candidate list described above can be generated in step Sj_1. The MV candidate list can be, for example, a list including the MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, Figures 47A-47C the MV derivation method shown in Figure 48A and Figure 48B the MV derivation method shown in Figure 49A and Figure 49Bany combination of the MV derivation methods illustrated in the above examples and other MV derivation methods.
[0505] Note that the MV candidate list can include MV candidates in modes in which prediction is performed in sub-block units other than the affine mode.
[0506] Note that, for example, an MV candidate list including MV candidates in the affine inter mode in which two control points are used and the affine inter mode in which three control points are used can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine inter mode in which two control points are used and an MV candidate list including MV candidates in the affine inter mode in which three control points are used can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine inter mode in which two control points are used and the affine inter mode in which three control points are used can be generated. The MV candidate(s) can be, for example, MVs for the coded blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), or MVs for the valid blocks in the blocks.
[0507] Note that an index indicating one of the MVs in the MV candidate list can be transmitted as the MV predictor selection information.
[0508] (MV derivation > triangle mode)
[0509] In the above example, the inter predictor 126 generates one rectangular prediction image for the current rectangular block. However, the inter predictor 126 can generate a plurality of prediction images each having a shape different from the rectangle of the current rectangular block, and can combine the plurality of prediction images to generate a final rectangular prediction image. The shape different from the rectangle can be, for example, a triangle.
[0510] Figure 52A is a conceptual diagram for illustrating generation of two triangular prediction images.
[0511] The inter predictor 126 generates a triangular prediction image by performing motion compensation on a first partition having a triangular shape in the current block using a first MV of the first partition to generate a triangular prediction image. Likewise, the inter predictor 126 generates a triangular prediction image by performing motion compensation on a second partition having a triangular shape in the current block using a second MV of the second partition to generate a triangular prediction image. Then, the inter predictor 126 generates a prediction image having a rectangular shape identical to the rectangular shape of the current block by combining these prediction images.
[0512] Note that a first prediction image having a rectangular shape corresponding to the current block can be generated using the first MV as the prediction image for the first partition. Additionally, a second prediction image having a rectangular shape corresponding to the current block can be generated using the second MV as the prediction image for the second partition. The prediction image for the current block can be generated by performing a weighted addition of the first prediction image and the second prediction image. Note that the portion where the weighted addition is performed can be a partial region across the boundary between the first partition and the second partition.
[0513] Figure 52B is a conceptual diagram for illustrating a first portion of a first partition that overlaps with a second partition, and a first set and a second set of samples that can be weighted as part of a correction process. The first portion can have, for example, one quarter of the width or height of the first partition. In another example, the first portion can have a width corresponding to N samples adjacent to an edge of the first partition, where N is an integer greater than zero, for example, N can be the integer 2. As shown, Figure 52B The left example of shows a rectangular partition having a rectangular portion that is one quarter of the width of the first partition, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion. Figure 52B The center example of shows a rectangular partition having a rectangular portion that is one quarter of the height of the first partition, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion. Figure 52B The right example of shows a triangular partition having a polygonal portion that is two samples in height, where the first set of samples includes samples outside the first portion and samples inside the first portion, and the second set of samples includes samples inside the first portion.
[0514] The first portion can be a portion of the first partition that overlaps with an adjacent partition. Figure 52C is a conceptual diagram for illustrating a first portion of a first partition that is a portion of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition is shown having an overlapping portion with a spatially adjacent rectangular partition. Partitions having other shapes (e.g., triangular partitions) can be employed, and the overlapping portion can overlap with a spatially or temporally adjacent partition.
[0515] Additionally, while examples are given in which inter prediction is used to generate a prediction image for each of the two partitions, intra prediction can be used to generate a prediction image for at least one of the partitions.
[0516] Figure 53is a flowchart showing one example of a process in the triangle mode.
[0517] In the triangle mode, first, the inter-frame predictor 126 splits the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode, as the prediction parameters, partition information that is information related to the splitting into the partitions in the stream. In other words, the inter-frame predictor 126 can output, as the prediction parameters, the partition information to the entropy encoder 110 through the prediction parameter generator 130.
[0518] First, the inter-frame predictor 126 obtains a plurality of MV candidates for the current block based on information (e.g., MVs of a plurality of encoded blocks around the current block in time or space) (step Sx_2). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0519] The inter-frame predictor 126 then selects, as a first MV and a second MV, an MV candidate for the first partition and an MV candidate for the second partition, respectively, from the plurality of MV candidates obtained in step Sx_1 (step Sx_3). At this time, the inter-frame predictor 126 encodes, as the prediction parameters, MV selection information for identifying the selected MV candidates in the stream. In other words, the inter-frame predictor 126 outputs, as the prediction parameters, the MV selection information to the entropy encoder 110 through the prediction parameter generator 130.
[0520] Next, the inter-frame predictor 126 generates a first prediction image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Likewise, the inter-frame predictor 126 generates a second prediction image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).
[0521] Finally, the inter-frame predictor 126 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_6).
[0522] Note that, although the first partition and the second partition are triangles in the example shown in Figure 52A , the first partition and the second partition can be trapezoids or other shapes different from each other. Further, although the current block includes two partitions in the examples shown in Figure 52A and Figure 52C , the current block can include three or more partitions.
[0523] Additionally, the first partition and the second partition can overlap each other. In other words, the first partition and the second partition can include the same pixel region. In this case, a prediction image in the first partition and a prediction image in the second partition can be used to generate a prediction image for the current block.
[0524] Additionally, while an example has been shown in which inter prediction is used to generate a prediction image for each of the two partitions, intra prediction can be used to generate a prediction image for at least one of the partitions.
[0525] Note that the MV candidate list used for selecting the first MV and the MV candidate list used for selecting the second MV can be different from each other, or the MV candidate list used for selecting the first MV can be used as the MV candidate list used for selecting the second MV.
[0526] Note that the partition information can include an index indicating a split direction in which at least the current block is split into a plurality of partitions. The MV selection information can include an index indicating the selected first MV and an index indicating the selected second MV. One index can indicate multiple pieces of information. For example, one index that collectively indicates a part or all of the partition information and a part or all of the MV selection information can be encoded.
[0527] (MV derivation > ATMVP mode)
[0528] Figure 54 is a conceptual diagram for showing one example of an advanced temporal motion vector prediction (ATMVP) mode in which an MV is derived in units of sub-blocks.
[0529] The ATMVP mode is a mode classified as the merge mode. For example, in the ATMVP mode, an MV candidate for each sub-block is registered in an MV candidate list for the normal merge mode.
[0530] More specifically, in the ATMVP mode, first, a temporal MV reference block associated with the current block is identified in an encoded reference picture specified by an MV (MV0) of a neighboring block located at a lower-left position with respect to the current block, as shown in Figure 54 Next, in each sub-block in the current block, an MV used for encoding a region corresponding to the sub-block in the temporal MV reference block is identified. The MV identified in this way is included in the MV candidate list for the sub-block in the current block as an MV candidate. When an MV candidate for each sub-block is selected from the MV candidate list, the sub-block undergoes motion compensation in which the MV candidate is used as an MV for the sub-block. In this way, a prediction image for each sub-block is generated.
[0531] While in the example shown in Figure 54 , a block located at a lower-left position with respect to the current block is used as a surrounding MV reference block, it should be noted that another block can be used. Additionally, the size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can be switched for units such as a slice, a tile, a picture, and the like.
[0532] (Motion Estimation > DMVR)
[0533] Figure 55 is a flowchart illustrating a relationship between the merge mode and decoder motion vector refinement DMVR.
[0534] The inter predictor 126 derives a motion vector for the current block according to the merge mode (step S1_1). Next, the inter predictor 126 determines whether to perform estimation of the motion vector, i.e., motion estimation (step S1_2). Here, when it is determined that the motion estimation is not performed (NO in step S1_2), the inter predictor 126 determines the motion vector derived in step S1_1 as the final motion vector for the current block (step S1_4). In other words, in this case, the motion vector for the current block is determined according to the merge mode.
[0535] When it is determined in step S1_1 that the motion estimation is performed (YES in step S1_2), the inter predictor 126 derives the final motion vector for the current block by estimating surrounding areas of a reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector for the current block is determined according to DMVR.
[0536] Figure 56 is a conceptual diagram for illustrating one example of a DMVR process for determining an MV.
[0537] First, MV candidates (L0 and L1) are selected for the current block, e.g., in the merge mode. Reference pixels are identified from a first reference picture (L0) that is a coded picture in the L0 list according to the MV candidate (L0). Likewise, reference pixels are identified from a second reference picture (L1) that is a coded picture in the L1 list according to the MV candidate (L1). A template is generated by computing an average of these reference pixels.
[0538] Next, the template is used to estimate each of the surrounding areas of the MV candidates of the first reference picture (L0) and the second reference picture (L1), and the MV that results in the least cost is determined as the final MV. Note that the cost can be computed, e.g., using a difference between each of the pixel values in the template and a corresponding one of the pixel values in the estimated area, a value of the MV candidate, etc.
[0539] It is not always necessary to perform the exact same process described here. Other processes can be used that achieve derivation of the final MV through estimation in the surrounding area of the MV candidate.
[0540] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining an MV. Unlike the example ofFigure 56 An example of DMVR illustrated in Figure 57 In the example illustrated in
[0541] First, the inter predictor 126 estimates a surrounding area of a reference block in each of the reference pictures included in the L0 list and the L1 list based on an initial MV that is an MV candidate obtained from each MV candidate list. For example, as illustrated in Figure 57 In the example illustrated in
[0542] Figure 58A is a conceptual diagram for illustrating one example of motion estimation in DMVR, and Figure 58B is a flowchart for illustrating one example of a process of motion estimation.
[0543] First, in step 1, the inter predictor 126 calculates a cost between a search position indicated by an initial MV (also referred to as a starting point) and eight surrounding search positions. The inter predictor 126 then determines whether the cost at each of the search positions other than the starting point is the minimum. Here, when it is determined that the cost at the search position other than the starting point is the minimum, the inter predictor 126 changes a target to the search position at which the minimum cost is obtained and performs a process in step 2. When the cost at the starting point is the minimum, the inter predictor 126 skips the process in step 2 and performs a process in step 3.
[0544] In step 2, the inter predictor 126 performs a search similar to the process in step 1, taking the search position after the target change as a new starting point from the result of the process in step 1. The inter predictor 126 then determines whether the cost at each of the search positions other than the starting point is minimum. Here, when the cost at the search position other than the starting point is determined to be minimum, the inter predictor 126 performs the process in step 4. When the cost at the starting point is minimum, the inter predictor 126 performs the process in step 3.
[0545] In step 4, the inter predictor 126 takes the search position at the starting point as a final search position, and determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0546] In step 3, the inter predictor 126 determines a pixel position in sub-pixel accuracy at which the cost is minimum, based on the costs at the four points of the positions above, below, left, and right with respect to the starting point in step 1 or step 2, and takes the pixel position as a final search position. The pixel position in sub-pixel accuracy is determined by performing a weighted addition of each of four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) using the cost at a corresponding one of the four search positions as a weight. The inter predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as a vector difference.
[0547] (Motion compensation > BIO / OBMC / LIC)
[0548] Motion compensation involves a mode for generating a prediction image and correcting the prediction image. The mode is, for example, bidirectional optical flow (BIO), overlapped block motion compensation (OBMC), local illumination compensation (LIC), and the like, which are described later.
[0549] Figure 59 is a flowchart showing one example of the process of generation of a prediction image.
[0550] The inter predictor 126 generates a prediction image (step Sm_1), and corrects the prediction image, for example, in accordance with any one of the modes described above (step Sm_2).
[0551] Figure 60 is a flowchart showing another example of the process of generation of a prediction image.
[0552] The inter predictor 126 determines a motion vector of the current block (step Sn_1). Next, the inter predictor 126 generates a prediction image using the motion vector (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter predictor 126 generates a final prediction image by correcting the prediction image (step Sn_4). Note that, in LIC described later, both the luminance and the chrominance can be corrected in step Sn_4. When it is determined not to perform the correction process (No in step Sn_3), the inter predictor 126 outputs the prediction image as the final prediction image without correcting the prediction image (step Sn_5).
[0553] (Motion compensation) OBMC
[0554] Note that, in addition to the motion information for the current block obtained by the motion estimation, the motion information for the neighboring blocks can be used to generate the inter prediction image. More specifically, by performing weighted addition of a prediction image (in the reference picture) based on the motion information obtained by the motion estimation and a prediction image (in the current picture) based on the motion information of the neighboring blocks, the inter prediction image can be generated for each sub-block in the current block. This inter prediction (motion compensation) is also referred to as overlapped block motion compensation (OBMC) or OBMC mode.
[0555] In the OBMC mode, information indicating a sub-block size for OBMC (e.g., referred to as OBMC block size) can be signaled at a sequence level. Further, information indicating whether to apply the OBMC mode (e.g., referred to as OBMC flag) can be signaled at a CU level. Note that, signaling such information does not necessarily need to be performed at the sequence level and the CU level, and can be performed at another level (e.g., a picture level, a slice level, a tile level, a CTU level, or a sub-block level).
[0556] The OBMC mode will be described in more detail. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for showing an outline of a prediction image correction process performed by the OBMC.
[0557] First, as shown in Figure 62 , a prediction image (Pred) by normal motion compensation is obtained using the MV assigned to the current block. In Figure 62 , the arrow “MV” points to the reference picture, and indicates what the current block of the current picture refers to in order to obtain the prediction image.
[0558] Next, a prediction picture (Pred_L) is obtained by applying a motion vector (MV_L) that has been derived for an encoded block neighboring to the left of the current block to the current block (reusing the motion vector for the current block). The motion vector (MV_L) is indicated by the arrow "MV_L" which indicates the reference picture from the current block. The first correction of the prediction picture is performed by overlapping the two prediction pictures Pred and Pred_L. This provides the effect of blending the boundary between the neighboring blocks.
[0559] Likewise, a prediction picture (Pred_U) is obtained by applying a MV (MV_U) that has been derived for an encoded block neighboring above the current block to the current block (reusing the MV for the current block). The MV (MV_U) is indicated by the arrow "MV_U" which indicates the reference picture from the current block. The second correction of the prediction picture is performed by overlapping the prediction picture Pred_U with the prediction picture (e.g. Pred and Pred_L) to which the first correction has been performed. This provides the effect of blending the boundary between the neighboring blocks. The prediction picture obtained by the second correction is the prediction picture in which the boundary between the neighboring blocks has been blended (smoothed), and is thus the final prediction picture for the current block.
[0560] While the above example is a two-path correction method using the left and above neighboring blocks, it should be noted that the correction method can be a three-path or more path correction method that also uses the right neighboring block and / or the below neighboring block.
[0561] Note that the area in which such overlapping is performed can just be a portion of the area in the region close to the block boundary, rather than the entire pixel area of the block.
[0562] Note that the prediction picture correction process for obtaining one prediction picture Pred from one reference picture by overlapping the additional prediction pictures Pred_L and Pred_U according to OBMC has been described above. However, when a prediction picture is corrected based on multiple reference pictures, a similar process can be applied to each of the multiple reference pictures. In this case, after the corrected prediction pictures are obtained from the respective reference pictures by performing the OBMC picture correction based on the multiple reference pictures, the obtained corrected prediction pictures are further overlapped to obtain the final prediction picture.
[0563] Note that in OBMC, the current block unit can be a PU, or a sub-block unit obtained by further splitting a PU.
[0564] One example of a method for determining whether to apply OBMC is a method for using an obmc_flag as a signal indicating whether to apply OBMC. As one specific example, the encoder 100 can determine whether the current block belongs to a region with complex motion. When the block belongs to a region with complex motion, the encoder 100 sets the obmc_flag to a value of "1" and applies OBMC at the time of encoding, and when the block does not belong to a region with complex motion, the encoder 100 sets the obmc_flag to a value of "0" and encodes the block without applying OBMC. The decoder 200 switches between applying and not applying OBMC by decoding the obmc_flag written in the stream.
[0565] (motion compensation > BIO)
[0566] Next, the MV derivation method is described. First, a mode for deriving MVs based on a model assuming uniform straight-line motion is described. This mode is also referred to as the bi-directional optical flow (BIO) mode. Additionally, this bi-directional optical flow can be written as BDOF instead of BIO.
[0567] Figure 63 is a conceptual diagram for showing the model assuming uniform straight-line motion. In Figure 63 , (v x , v y ) indicates a velocity vector, and τ0 and τ1 indicate the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) indicates the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) indicates the MV corresponding to the reference picture Ref1.
[0568] Here, assuming that uniform straight-line motion is exhibited by the velocity vector (v x , v y ), (MV x0 , MV y0 ) and (MV x1 , MV y1 ) are expressed as (v xτ0 , v yτ0 ) and (-v xτ1 , -v yτ1 ), respectively, and the following optical flow equation (2) is given.
[0569] [math 3]
[0570]
[0571] Here, I(k) indicates a motion-compensated luma value k (k = 0, 1) of the reference picture after the motion compensation. The optical flow equation represents that a sum of (i) a time derivative of the luma value, (ii) a product of the horizontal velocity and a horizontal component of the spatial gradient of the reference image, and (iii) a product of the vertical velocity and a vertical component of the spatial gradient of the reference image is equal to zero. Based on a combination of the optical flow equation and the Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected in pixel units.
[0572] Note that the motion vector can be derived at the decoder side 200 using a method other than the model based on the assumption of uniform straight-line motion to derive the motion vector. For example, the motion vector can be derived in sub-block units based on the motion vectors of a plurality of neighboring blocks.
[0573] Figure 64 is a flowchart showing one example of a procedure of inter prediction according to the BIO. Figure 65 is a functional block diagram showing one example of a functional configuration of the inter predictor 126 that can perform the inter prediction according to the BIO.
[0574] As shown in Figure 65 , the inter predictor 126 includes, for example, a memory 126a, an interpolated image deriver 126b, a gradient image deriver 126c, an optical flow deriver 126d, a correction value deriver 126e, and a predicted image corrector 126f. Note that the memory 126a can be the frame memory 122.
[0575] The inter predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) including the current block. Then, the inter predictor 126 derives a predicted image for the current block using the two motion vectors (M0, M1) (step Sy_1). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.
[0576] Next, the interpolated image deriver 126b derives an interpolated image I 0 for the current block using the motion vector M0 and the reference picture L0 by referring to the memory 126a. Next, the interpolated image deriver 126b derives an interpolated image I 1 for the current block using the motion vector M1 and the reference picture L1 by referring to the memory 126a (step Sy_2). Here, the interpolated image I 0is an interpolated image included in the reference picture Ref0 and derived for the current block, and the interpolated image I 1 is an interpolated image included in the reference picture Ref1 and derived for the current block. The interpolated image I 0 and the interpolated image I 1 Each of the interpolated image I 0 and the interpolated image I 1 may be an image larger than the current block. Further, the interpolated image I 0 and the interpolated image I 1 may include a predicted image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.
[0577] Additionally, the gradient image deriver 126c derives gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) of the current block from the interpolated image I 0 and the interpolated image I 1 (step Sy_3). Note that the gradient images in the horizontal direction are (Ix 0 , Ix 1 ), and the gradient images in the vertical direction are (Iy 0 , Iy 1 ). The gradient image deriver 126c can derive each gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image can indicate an amount of spatial change in pixel values in the horizontal direction, in the vertical direction, or in both directions.
[0578] Next, the optical flow deriver 126d derives, as velocity vectors, optical flows (vx, vy) for each sub-block of the current block using the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (step Sy_4). The optical flow indicates a coefficient for correcting an amount of spatial pixel movement, and can be referred to as a local motion estimation value, a corrected motion vector, or a corrected weighted vector. As one example, the sub-block can be a 4x4-pixel sub-CU. Note that the optical flow derivation can be performed for each pixel unit or the like, instead of for each sub-block.
[0579] Next, the inter predictor 126 uses the optical flow (vx, vy) to correct the prediction picture for the current block. For example, the correction value deriver 126e uses the optical flow (vx, vy) to derive correction values for values of pixels included in the current block (step Sy_5). The prediction picture correcter 126f can then use the correction values to correct the prediction picture for the current block (step Sy_6). Note that the correction values can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.
[0580] Note that the BIO process flow is not limited to the process disclosed in Figure 64 For example, only a part of the process disclosed in Figure 64 may be executed, or a different process or a different process can be added as an alternative, or the processes can be executed in a different processing order, etc.
[0581] (motion compensation > LIC)
[0582] Next, one example of a mode for generating a prediction picture (prediction) using a local illumination compensation (LIC) process is described.
[0583] Figure 66A is a conceptual diagram for illustrating one example of a process of a prediction picture generation method using an illumination correction process performed by LIC. Figure 66B is a flowchart illustrating one example of a process of a prediction picture generation method using LIC.
[0584] First, the inter predictor 126 derives MVs from the encoded reference pictures and obtains a reference picture corresponding to the current block (step Sz_1).
[0585] Next, the inter predictor 126 extracts, for the current block, information indicating how the luminance values change between the current block and the reference picture (step Sz_2). This extraction is performed based on luminance pixel values of an encoded left-side neighboring reference region (surrounding reference region) and an encoded upper neighboring reference region (surrounding reference region) in the current picture, and luminance pixel values at corresponding positions in the reference picture specified by the derived MVs. The inter predictor 126 uses the information indicating how the luminance values change to calculate illumination correction parameters (step Sz_3).
[0586] The inter predictor 126 generates the prediction image for the current block by performing an illumination correction process in which an illumination correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_4). In other words, the prediction image is corrected based on the illumination correction parameter, which is the reference image in the reference picture specified by the MV. In this correction, the illumination can be corrected, or the chrominance can be corrected, or both. In other words, the chrominance correction parameter can be calculated using information indicating how the chrominance changes, and the chrominance correction process can be performed.
[0587] Note that, Figure 66A The shape of the surrounding reference region illustrated in FIG. 12 is one example; another shape can be used.
[0588] Further, while the process in which the prediction image is generated from a single reference picture has been described here, the case in which the prediction image is generated from multiple reference pictures can be described in the same way. The prediction image can be generated in the same way as described above after performing the illumination correction process on the reference image obtained from the reference picture.
[0589] One example of a method for determining whether to apply LIC is a method for using lic_flag as a signal indicating whether to apply LIC. As one specific example, the encoder 100 determines whether the current block belongs to a region having an illumination change. When the block belongs to a region having an illumination change, the encoder 100 sets lic_flag to the value "1" and applies LIC at the time of encoding, and when the block does not belong to a region having an illumination change, the encoder 100 sets lic_flag to the value "0" and performs encoding without applying LIC. The decoder 200 can decode lic_flag written in the stream and decode the current block by switching between applying LIC and not applying LIC according to the flag value.
[0590] One example of a different method of determining whether to apply the LIC process is a determination method according to whether the LIC process has been applied to the surrounding block. As one specific example, when the current block has been processed in the merge mode, the inter predictor 126 determines whether the coded surrounding block selected in the MV derivation in the merge mode has been coded using LIC. The inter predictor 126 performs encoding by switching between applying LIC and not applying LIC according to the result. Note that the same process is applied in the process on the decoder 200 side in this example as well.
[0591] The illumination correction (LIC) process has been described with reference to Figure 66A and Figure 66B and is further described below.
[0592] First, the inter predictor 126 derives, from a reference picture that is an encoded picture, an MV for obtaining a reference image corresponding to a current block to be encoded.
[0593] Next, the inter predictor 126 extracts information indicating how the luminance values of the reference picture change to the luminance values of the current picture using the luminance pixel values of the encoded surrounding reference region neighboring the left and the top of the current block and the luminance values in the corresponding positions of the reference picture specified by the MV, and calculates the illumination correction parameters. For example, assume that the luminance pixel value of a given pixel in the surrounding reference region in the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region in the reference picture is p1. The inter predictor 126 calculates the coefficients A and B for optimizing Axp1+B=p0 as the illumination correction parameters for a plurality of pixels in the surrounding reference region.
[0594] Next, the inter predictor 126 performs illumination correction processing using the illumination correction parameters for the reference image in the reference picture specified by the MV to generate a prediction image for the current block. For example, assume that the luminance pixel value in the reference image is p2, and the luminance pixel value after the illumination correction of the prediction image is p3. The inter predictor 126 generates the prediction image after going through the illumination correction process by calculating Axp2+B=p3 for each of the pixels in the reference image.
[0595] For example, a region having a determined number of pixels extracted from each of the top neighboring pixel and the left neighboring pixel can be used as the surrounding reference region. Additionally, the surrounding reference region is not limited to a region neighboring the current block, and can be a region not neighboring the current block. In Figure 66A In the example shown in FIG. 10, the surrounding reference region in the reference picture can be a region specified by another MV in the current picture from the surrounding reference region in the current picture. For example, the other MV can be an MV in the surrounding reference region in the current picture.
[0596] Although the operations performed by the encoder 100 have been described here, it should be noted that the decoder 200 performs similar operations.
[0597] Note that LIC can be applied not only to luminance but also to chrominance. At this time, the correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of Y, Cb, and Cr.
[0598] Additionally, the LIC process can be applied in units of sub-blocks. For example, the correction parameters can be derived using the surrounding reference region in the current sub-block and the surrounding reference region in the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0599] (predictor)
[0600] The prediction controller 128 selects one of the intra prediction signal (the image or signal output from the intra predictor 124) and the inter prediction signal (the image or signal output from the inter predictor 126), and outputs the selected prediction image to the subtracter 104 and the adder 116 as a prediction signal.
[0601] (predictor parameter generator)
[0602] The predictor parameter generator 130 can output information related to the intra prediction, the inter prediction, the selection of the prediction image in the prediction controller 128, and the like, as a prediction parameter to the entropy encoder 110. The entropy encoder 110 can generate a stream based on the prediction parameter input from the predictor parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameter can be used in the decoder 200. The decoder 200 can receive and decode the stream, and perform the same process as the prediction processes performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameter can include, for example, (i) a selection of a prediction signal (for example, an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) a prediction process or an optional index, flag, or value indicating the prediction process based on the prediction process performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128.
[0603] (decoder)
[0604] Next, a decoder 200 capable of decoding a stream output from the above-described encoder 100 is described. Figure 67 is a block diagram illustrating a functional configuration of the decoder 200 according to the present embodiment. The decoder 200 is a device that decodes a stream that is an encoded image in units of blocks.
[0605] As Figure 67 illustrated in FIG. 27, the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformer 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a predictor parameter generator 222, and a split determiner 224. Note that the intra predictor 216 and the inter predictor 218 are configured as a part of a prediction performer.
[0606] (installation example of decoder)
[0607] Figure 68 is a functional block diagram illustrating an installation example of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, Figure 67The plurality of constituent elements of the decoder 200 shown in FIG. 2 are mounted on Figure 68 The processor b1 and the memory b2 shown in FIG. 2.
[0608] The processor b1 is a circuit that performs information processing and is coupled to the memory b2. For example, the processor b1 is a special-purpose or general-purpose electronic circuit that decodes a stream. The processor b1 can be a processor such as a CPU. Alternatively, the processor b1 can be a collection of a plurality of electronic circuits. Further, for example, the processor b1 can assume a role of Figure 67 Two or more of the plurality of constituent elements of the decoder 200 shown in FIG. 2 other than the constituent elements for storing information assume a role of, and the like.
[0609] The memory b2 is a special-purpose or general-purpose memory that stores information used by the processor b1 to decode a stream. The memory b2 can be an electronic circuit and can be connected to the processor b1. Alternatively, the memory b2 can be included in the processor b1. Further, the memory b2 can be a collection of a plurality of electronic circuits. Further, the memory b2 can be a magnetic disk, an optical disk, or the like, or can be denoted as a storage device, a recording medium, or the like. Further, the memory b2 can be a nonvolatile memory or a volatile memory.
[0610] For example, the memory b2 can store an image or a stream. Alternatively, the memory b2 can store a program for causing the processor b1 to decode a stream.
[0611] Further, for example, the memory b2 can assume a role of Figure 67 Two or more of the plurality of constituent elements of the decoder 200 shown in FIG. 2 for storing information assume a role of, and the like. More specifically, the memory b2 can assume a role of Figure 67 The block memory 210 and the frame memory 214 shown in FIG. 2. More specifically, the memory b2 can store a reconstructed image (specifically, a reconstructed block, a reconstructed picture, or the like).
[0612] Note that, in the decoder 200, all of the plurality of constituent elements indicated in FIG. 2, and the like can not be implemented, and all of the processes described herein can not be performed. Figure 67 Part of the constituent elements indicated in FIG. 2, and the like can be included in another device, or part of the processes described herein can be performed by another device. Figure 67
[0613] In the following, the overall flow of the process performed by the decoder 200 is described, and then each of the constituent elements included in the decoder 200 will be described. Note that some of the constituent elements included in the decoder 200 perform the same processes as some of those performed in the encoder 100, and thus the same processes are not repeatedly described in detail. For example, the inverse quantizer 204, the inverse transformer 206, the adder 208, the block memory 210, the frame memory 214, the intra predictor 216, the inter predictor 218, the prediction controller 220, and the loop filter 212 included in the decoder 200 perform similar processes to those performed by the inverse quantizer 112, the inverse transformer 114, the adder 116, the block memory 118, the frame memory 122, the intra predictor 124, the inter predictor 126, the prediction controller 128, and the loop filter 120 included in the encoder 100, respectively.
[0614] (Overall flow of the decoding process)
[0615] Figure 69 is a flowchart showing one example of the overall decoding process performed by the decoder 200.
[0616] First, the split determiner 224 in the decoder 200 determines a split pattern of each of a plurality of fixed-size blocks (e.g., 128 x 128 pixels) included in a picture based on parameters input from the entropy decoder 202 (step Sp_1). The split pattern is the split pattern selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks of the split pattern.
[0617] The entropy decoder 202 decodes (specifically, entropy-decodes) the encoded quantized coefficients and the prediction parameters of the current block (step Sp_2).
[0618] Next, the inverse quantizer 204 performs inverse quantization on the plurality of quantized coefficients, and the inverse transformer 206 performs inverse transformation on the result to recover the prediction residual (i.e., the difference block) (step Sp_3).
[0619] Next, a prediction performer including all or a part of the intra predictor 216, the inter predictor 218, and the prediction controller 220 generates a prediction signal of the current block (step Sp_4).
[0620] Next, the adder 208 adds the prediction image and the prediction residual to generate a reconstructed image (also referred to as a decoded image block) of the current block (step Sp_5).
[0621] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).
[0622] The decoder 200 then determines whether decoding of the entire picture has ended (step Sp_7). When it is determined that decoding has not ended ("No" in step Sp_7), the decoder 200 repeats the process starting from step Sp_1.
[0623] Note that the processes of these steps Sp_1 to Sp_7 can be sequentially executed by the decoder 200, or two or more of the processes can be executed in parallel. The processing order of two or more of the processes can be modified.
[0624] (Split determiner)
[0625] Figure 70 is a conceptual diagram for showing the relationship between the split determiner 224 and other constituent elements in the embodiment. As an example, the split determiner 224 can execute the following processes.
[0626] For example, the split determiner 224 collects block information from the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoder 202. The split determiner 224 can then determine a split pattern of a fixed-size block based on the block information and the parameters. The split determiner 224 can then output information indicating the determined split pattern to the inverse transformer 206, the intra predictor 216, and the inter predictor 218. The inverse transformer 206 can perform inverse transformation of the transform coefficients based on the split pattern indicated by the information from the split determiner 224. The intra predictor 216 and the inter predictor 218 can generate a predicted image based on the split pattern indicated by the information from the split determiner 224.
[0627] (Entropy decoder)
[0628] Figure 71 is a block diagram showing one example of a functional configuration of the entropy decoder 202.
[0629] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to a split mode by entropy-decoding the stream. For example, CABAC is used for the entropy-decoding. More specifically, the entropy-decoding 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a debinarizer 202c. The binary arithmetic decoder 202a arithmetically decodes the stream into a binary signal using a context value derived by the context controller 202b. The context controller 202b derives the context value from a feature or a surrounding state of a syntax element, i.e., a probability of occurrence of a binary signal, in the same manner as the context controller 110b of the encoder 100 performs. The debinarizer 202c performs debinarization to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal indicating the quantized coefficients, as described above. This binarization can be performed according to the binarization method described above.
[0630] In this way, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can output the prediction parameters included in the stream (see Figure 1 ) to the intra predictor 216, the inter predictor 218, and the prediction controller 220. The intra predictor 216, the inter predictor 218, and the prediction controller 220 are capable of performing the same prediction processes as those performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.
[0631] Figure 72 is a conceptual diagram for showing a flow of an example CABAC process in the entropy decoder 202.
[0632] First, initialization is performed in the CABAC in the entropy decoder 202. In the initialization, initialization in the binary arithmetic decoder 202a and setting of initial context values are performed. The binary arithmetic decoder 202a and the debinarizer 202c then perform arithmetic decoding and debinarization of the encoded data of, for example, a CTU. At this time, the context controller 202b updates the context values each time the line arithmetic decoding is performed. The context controller 202b then saves the context values as post-processing. For example, the saved context values are used to initialize the context values for the next CTU.
[0633] (inverse quantizer)
[0634] The inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block input from the entropy decoder 202. More specifically, the inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inverse-quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.
[0635] Figure 73 is a block diagram showing one example of a functional configuration of the inverse quantizer 204.
[0636] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage 204d, and an inverse quantization executor 204e.
[0637] Figure 74 is a flowchart showing one example of a process of inverse quantization performed by the inverse quantizer 204.
[0638] As one example, the inverse quantizer 204 can perform the inverse quantization process on each CU in the flow shown in Figure 74 The inverse quantization process is performed on each CU in the flow shown in
[0639] Next, the predicted quantization parameter generator 204b then obtains quantization parameters for processing units different from the current block from the quantization parameter storage 204d (step Sv_13). The predicted quantization parameter generator 204b generates a predicted quantization parameter for the current block based on the obtained quantization parameters (step Sv_14).
[0640] The quantization parameter generator 204a then generates a quantization parameter for the current block based on the delta quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the delta quantization parameter for the current block obtained from the entropy decoder 202 and the predicted quantization parameter for the current block generated by the predicted quantization parameter generator 204b can be added to generate the quantization parameter for the current block. Additionally, the quantization parameter generator 204a stores the quantization parameter for the current block in the quantization parameter storage 204d (step Sv_16).
[0641] Next, the inverse quantization executor 204e inverse quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).
[0642] Note that the delta quantization parameter can be decoded at a bit sequence level, a picture level, a slice level, a tile level, or a CTU level. Additionally, an initial value of the quantization parameter can be decoded at a sequence level, a picture level, a slice level, a tile level, or a CTU level. At this time, the initial value of the quantization parameter and the delta quantization parameter can be used to generate the quantization parameter.
[0643] Note that the inverse quantizer 204 can include a plurality of inverse quantizers, and can inverse quantize the quantized coefficients using an inverse quantization method selected from a plurality of inverse quantization methods.
[0644] (inverse transformer)
[0645] The inverse transformer 206 recovers the prediction residual by inverse transforming the transform coefficients as input from the inverse quantizer 204.
[0646] For example, when the information parsed from the stream indicates that EMT or AMT is to be applied (e.g., when the AMT flag is true), the inverse transformer 206 inverse transforms the transform coefficients of the current block based on the information indicating the parsed transform type.
[0647] Further, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a secondary inverse transform to the transform coefficients.
[0648] Figure 75 is a flowchart showing one example of the process performed by the inverse transformer 206.
[0649] For example, the inverse transformer 206 determines whether or not there is information in the stream indicating that the orthogonal transform is not to be performed (step St_11). Here, when it is determined that there is no such information (NO in step St_11) (e.g.: there is no any indication as to whether or not the orthogonal transform is to be performed; there is an indication that the orthogonal transform is to be performed), the inverse transformer 206 obtains information indicating the transform type decoded by the entropy decoder 202 (step St_12). Next, based on the information, the inverse transformer 206 determines the transform type for the orthogonal transform in the encoder 100 (step St_13). The inverse transformer 206 then performs the inverse orthogonal transform using the determined transform type (step St_14). As described in Figure 75 As described in the foregoing, when it is determined that there is information indicating that the orthogonal transform is not to be performed (YES in step St_11) (e.g.: an explicit indication that the orthogonal transform is not to be performed; there is no indication that the orthogonal transform is to be performed), the orthogonal transform is not performed.
[0650] Figure 76 is a flowchart showing one example of the process performed by the inverse transformer 206.
[0651] For example, the inverse transformer 206 determines whether the transform size is smaller than or equal to a determination value (step Su_11). The determination value can be predetermined. Here, when the transform size is determined to be smaller than or equal to the determination value (Yes in step Su_11), the inverse transformer 206 obtains, from the entropy decoder 202, information indicating which transform type among at least one transform type included in the first transform type group is used by the encoder 100 (step Su_12). Note that such information is decoded by the entropy decoder 202 and output to the inverse transformer 206.
[0652] Based on the information, the inverse transformer 206 determines a transform type used for the orthogonal transform in the encoder 100 (step Su_13). The inverse transformer 206 then performs inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su_14). When the transform size is determined not to be smaller than or equal to the determination value (No in step Su_11), the inverse transformer 206 performs inverse transform on the transform coefficients of the current block using the second transform type group (step Su_15).
[0653] Note that, as one example, the inverse orthogonal transform by the inverse transformer 206 can be performed according to the flow illustrated in Figure 75 or Figure 76 Additionally, the inverse orthogonal transform can be performed by using a defined transform type without decoding the information indicating the transform type used for the orthogonal transform. The defined transform type can be a pre-defined transform type or a default transform type. Additionally, the transform type can be specifically DST7, DCT8, or the like. In the inverse orthogonal transform, an inverse transform basis function corresponding to the transform type is used.
[0654] (adder)
[0655] The adder 208 reconstructs the current block by adding the prediction residual as an input from the inverse transformer 206 and the prediction image as an input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.
[0656] (block memory)
[0657] The block memory 210 is a storage device for storing blocks included in the current picture and which can be referred to in the intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.
[0658] (loop filter)
[0659] The loop filter 212 applies a loop filter to the reconstructed picture generated by the adder 208, and outputs the filtered reconstructed picture to the frame memory 214, and provides an output of the decoder 200, for example, and to a display device or the like.
[0660] When the information indicating the on or off of the ALF parsed from the stream indicates that the ALF is on, one filter is selected from a plurality of filters, for example, based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed picture.
[0661] Figure 77 is a block diagram showing one example of a functional configuration of the loop filter 212. Note that the loop filter 212 has a configuration similar to that of the loop filter 120 of the encoder 100.
[0662] For example, as shown in Figure 77 , the loop filter 212 includes a deblocking filter executor 212a, an SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a executes a deblocking filtering process on the reconstructed picture. The SAO executor 212b executes an SAO process on the reconstructed picture after the deblocking filtering process. The ALF executor 212c executes an ALF process on the reconstructed picture after the SAO process. Note that the loop filter 212 does not always need to include all the constituent elements disclosed in Figure 77 , and can include only a part of the constituent elements. Additionally, the loop filter 212 can be configured to execute the above-described processes in a processing order different from the processing order disclosed in Figure 77 , can not execute all the processes shown in Figure 77 , and the like.
[0663] (frame memory)
[0664] The frame memory 214 is a storage device for storing, for example, a reference picture used for inter prediction, and can also be referred to as a frame buffer. More specifically, the frame memory 214 stores the reconstructed picture filtered by the loop filter 212.
[0665] (predictor (intra predictor, inter predictor, prediction controller))
[0666] Figure 78 is a flowchart showing one example of a process performed by the predictor of the decoder 200. Note that the prediction executor can include all or a part of the following constituent elements: the intra predictor 216; the inter predictor 218; and the prediction controller 220. The prediction executor includes, for example, the intra predictor 216 and the inter predictor 218.
[0667] The predictor generates a prediction picture of the current block (step Sq_1). The prediction picture can also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra prediction picture or an inter prediction picture. More specifically, the predictor generates the prediction picture of the current block by prediction picture generation, prediction residual recovery, and prediction picture addition using a reconstructed picture that has been obtained for another block. The predictor of the decoder 200 generates the same prediction picture as the prediction picture generated by the predictor of the encoder 100. In other words, the prediction picture is generated according to a method common to or corresponding to each other between the predictors.
[0668] The reconstructed picture is, for example, a picture in a reference picture, or a picture of a decoded block in a current picture that includes the current picture (i.e., the other block described above). The decoded block in the current picture is, for example, a neighboring block of the current block.
[0669] Figure 79 is a flowchart showing another example of a process performed by the predictor of the decoder 200.
[0670] The predictor determines a method or mode for generating the prediction picture (step Sr_1). The method or mode can be determined, for example, based on, for example, a prediction parameter or the like.
[0671] When the first method is determined as the mode for generating the prediction picture, the predictor generates the prediction picture according to the first method (step Sr_2a). When the second method is determined as the mode for generating the prediction picture, the predictor generates the prediction picture according to the second method (step Sr_2b). When the third method is determined as the mode for generating the prediction picture, the predictor generates the prediction picture according to the third method (step Sr_2c).
[0672] The first method, the second method, and the third method can be mutually different methods for generating the prediction picture. Each of the first method to the third method can be an inter prediction method, an intra prediction method, or another prediction method. The reconstructed picture described above can be used in these prediction methods.
[0673] Figure 80 is a flowchart showing another example of a process performed by the predictor of the decoder 200.
[0674] As one example, the predictor can perform the prediction process according to the flow shown in Figure 80 Note that the in-block copy shown in Figure 80 The in-block copy is a mode belonging to inter prediction, and in which a block included in a current picture is referred to as a reference picture or a reference block. In other words, in the in-block copy, a picture different from the current picture is not referred to. Additionally, the Figure 80The PCM mode illustrated in FIG. 6 is a mode belonging to intra prediction, and in which no transform and quantization are performed.
[0675] (intra predictor)
[0676] The intra predictor 216 performs intra prediction based on an intra prediction mode parsed from the stream by referring to the block stored in the block memory 210 to generate a prediction image of the current block (i.e., an intra predicted block). More specifically, the intra predictor 216 performs intra prediction by referring to pixel values (e.g., luma and / or chroma values) of one or more blocks neighboring the current block to generate an intra predicted image, and then outputs the intra predicted image to the prediction controller 220.
[0677] Note that when an intra prediction mode in which a luma block is referred to in intra prediction of a chroma block is selected, the intra predictor 216 can predict a chroma component of the current block based on a luma component of the current block.
[0678] Further, when information parsed from the stream indicates that PDPC is to be applied, the intra predictor 216 corrects pixel values of intra prediction based on horizontal / vertical reference pixel gradients.
[0679] Figure 81 is a diagram illustrating one example of a process performed by the intra predictor 216 of the decoder 200.
[0680] The intra predictor 216 first determines whether or not MPM is employed. As Figure 81 As illustrated in FIG. 6, the intra predictor 216 determines whether or not an MPM flag indicating 1 is present in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 is present (Yes in step Sw_11), the intra predictor 216 obtains information indicating an intra prediction mode selected in the encoder 100 among the MPMs from the entropy decoder 202. Note that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPMs (step Sw_13). The MPMs include, for example, six kinds of intra prediction modes. The intra predictor 216 then determines an intra prediction mode included in the plurality of intra prediction modes included in the MPMs and indicated by the information obtained in step Sw_12 (step Sw_14).
[0681] When it is determined that the MPM flag indicating 1 is not present (NO in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains, from the entropy decoder 202, information indicating the intra prediction mode selected in the encoder 100 from among at least one intra prediction mode that has never been included in the MPM. Note that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. The intra predictor 216 then determines the intra prediction mode that is not included in the plurality of intra prediction modes included in the MPM and that is indicated by the information obtained in step Sw_15 (step Sw_17).
[0682] The intra predictor 216 generates a prediction image in accordance with the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).
[0683] (inter-predictor)
[0684] The inter-predictor 218 predicts the current block by referring to the reference picture stored in the frame memory 214. The prediction is performed in units of the current block or a current sub-block in the current block. Note that the sub-block is included in the block and is a unit smaller than the block. The size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can be switched for units such as slices, tiles, pictures, and the like.
[0685] For example, the inter-predictor 218 generates an inter-predicted image of the current block or the current sub-block by performing motion compensation using the motion information (e.g., MV) parsed from the stream (e.g., the prediction parameters output from the entropy decoder 202) and outputs the inter-predicted image to the prediction controller 220.
[0686] When the information parsed from the stream indicates that the OBMC mode is to be applied, the inter-predictor 218 generates the inter-predicted image using the motion information of the neighboring block in addition to the motion information of the current block obtained by the motion estimation.
[0687] Further, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter-predictor 218 derives the motion information by performing the motion estimation in accordance with the pattern matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter-predictor 218 then performs the motion compensation (prediction) using the derived motion information.
[0688] Further, when the BIO mode is to be applied, the inter-predictor 218 derives the MV based on a model assuming uniform rectilinear motion. Additionally, when the information parsed from the stream indicates that the affine mode is to be applied, the inter-predictor 218 derives the MV for each sub-block based on the MVs of a plurality of neighboring blocks.
[0689] (MV derivation process)
[0690] Figure 82 is a flowchart showing one example of the process of MV derivation in the decoder 200.
[0691] For example, the inter predictor 218 determines whether or not to decode motion information (e.g., MV). For example, the inter predictor 218 can make the determination in accordance with a prediction mode included in the stream, or can make the determination based on other information included in the stream. Here, when it is determined to decode motion information, the inter predictor 218 derives the MV for the current block in a mode in which motion information is decoded. When it is determined not to decode motion information, the inter predictor 218 derives the MV in a mode in which motion information is not decoded.
[0692] Here, the MV derivation mode includes a normal inter mode, a normal merge mode, an FRUC mode, an affine mode, and the like, which are described later. The mode in which motion information is decoded among the modes includes the normal inter mode, the normal merge mode, the affine mode (specifically, an affine inter mode and an affine merge mode), and the like. Note that the motion information can include not only the MV but also MV predictor selection information, which is described later. The mode in which motion information is not decoded includes the FRUC mode and the like. The inter predictor 218 selects a mode for deriving the MV for the current block from among the plurality of modes, and derives the MV for the current block using the selected mode.
[0693] Figure 83 is a flowchart showing one example of the process of MV derivation in the decoder 200.
[0694] For example, the inter predictor 218 can determine whether or not to decode the MV difference, i.e., for example, can make the determination in accordance with a prediction mode included in the stream, or can make the determination based on other information included in the stream. Here, when it is determined to decode the MV difference, the inter predictor 218 can derive the MV for the current block in a mode in which the MV difference is decoded. In this case, for example, the MV difference included in the stream is decoded as a prediction parameter.
[0695] When it is determined not to decode any MV difference, the inter predictor 218 derives the MV in a mode in which the MV difference is not decoded. In this case, an encoded MV difference is not included in the stream.
[0696] Here, as described above, the MV derivation mode includes normal inter mode, normal merge mode, FRUC mode, affine mode, and the like, which will be described later. Among the modes, the mode in which the MV difference is coded includes normal inter mode and affine mode (specifically, affine inter mode), and the like. Among the modes, the mode in which the MV difference is not coded includes FRUC mode, normal merge mode, affine mode (specifically, affine merge mode), and the like. The inter predictor 218 selects the mode for deriving the MV for the current block from among the modes, and derives the MV for the current block using the selected mode.
[0697] (MV derivation > normal inter mode)
[0698] For example, when the information parsed from the stream indicates that the normal inter mode is to be applied, the inter predictor 218 derives the MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.
[0699] Figure 84 is a flowchart illustrating an example of a procedure of inter prediction by the normal inter mode in the decoder 200.
[0700] The inter predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter predictor 218 obtains a plurality of MV candidates for the current block based on information (for example, MVs of a plurality of decoded blocks around the current block in time or space) (step Sg_11). In other words, the inter predictor 218 generates an MV candidate list.
[0701] Next, the inter predictor 218 extracts N (an integer of 2 or more) MV candidates from the plurality of MV candidates obtained in step Sg_11 as motion vector predictor candidates (also referred to as MV predictor candidates) in the order of the ranking in the priority order determined (step Sg_12). Note that the order of the ranking in the priority order can be determined in advance for the respective N MV predictor candidates, and the order can be determined in advance.
[0702] Next, the inter predictor 218 decodes the MV predictor selection information from the input stream, and selects one of the N MV predictor candidates as the MV predictor for the current block using the decoded MV predictor selection information (step Sg_13).
[0703] Next, the inter predictor 218 decodes the MV difference from the input stream, and derives the MV for the current block by adding the difference value as the decoded MV difference to the selected MV predictor (step Sg_14).
[0704] Finally, the inter predictor 218 generates a prediction image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sg_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed on each of all the blocks in a slice, the inter prediction of the slice using the normal inter mode ends. For example, when the processes in steps Sg_11 to Sg_15 are performed on each of all the blocks in a picture, the inter prediction of the picture using the normal inter mode ends. Note that not all the blocks included in a slice can undergo the processes in steps Sg_11 to Sg_15, and the inter prediction of the slice using the normal inter mode can end when a part of the blocks undergoes the processes. This also applies to the picture in steps Sg_11 to Sg_15. When the processes are performed on a part of the blocks in a picture, the inter prediction of the picture using the normal inter mode can end.
[0705] (MV derivation > normal merge mode)
[0706] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter predictor 218 derives an MV and performs motion compensation (prediction) using the MV.
[0707] Figure 85 is a flowchart showing an example of the processes of the inter prediction by the normal merge mode in the decoder 200.
[0708] First, the inter predictor 218 obtains a plurality of MV candidates for the current block based on information (e.g., MVs of a plurality of decoded blocks around the current block in time or space) (step Sh_11). In other words, the inter predictor 218 generates an MV candidate list.
[0709] Next, the inter predictor 218 selects one MV candidate from the plurality of MV candidates obtained in step Sh_11, thereby deriving an MV for the current block (step Sh_12). More specifically, the inter predictor 218 obtains MV selection information included in the stream as a prediction parameter, and selects the MV candidate identified by the MV selection information as the MV for the current block.
[0710] Finally, the inter predictor 218 generates a prediction picture for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Sh_13). The process in steps Sh_11 to Sh_13 is performed for each block, for example. The inter prediction of the slice using the normal merge mode ends when the process in steps Sh_11 to Sh_13 is performed on each of all the blocks in the slice, for example. Additionally, the inter prediction of the picture using the normal merge mode ends when the process in steps Sh_11 to Sh_13 is performed on each of all the blocks in the picture. Note that not all the blocks included in the slice undergo the process of steps Sh_11 to Sh_13, and the inter prediction of the slice using the normal merge mode can end when a part of the blocks undergo the process. This also applies to the picture in steps Sh_11 to Sh_13. The inter prediction of the picture using the normal merge mode can end when the process is performed on a part of the blocks in the picture.
[0711] (MV derivation > Merge mode)
[0712] When the information parsed from the stream indicates that the FRUC mode is to be applied, for example, the inter predictor 218 derives an MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is derived at the decoder 200 side without being signaled from the encoder 100 side. The decoder 200 can derive the motion information by performing motion estimation, for example. In this case, the decoder 200 performs motion estimation without using any pixel value in the current block.
[0713] Figure 86 is a flowchart showing an example of the process of inter prediction by the FRUC mode in the decoder 200.
[0714] First, the inter predictor 218 generates a list indicating MVs of decoded blocks spatially or temporally neighboring the current block by referring to the MVs as MV candidates (the list is an MV candidate list, and for example, can also be used as an MV candidate list for normal merge mode) (step Si_11). Next, a best MV candidate is selected from among the plurality of MV candidates registered in the MV candidate list (step Si_12). For example, the inter predictor 218 calculates an evaluation value of each of the MV candidates included in the MV candidate list, and selects one of the MV candidates as the best MV candidate based on the evaluation values. Based on the selected best MV candidate, the inter predictor 218 then derives an MV for the current block (step Si_14). More specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Additionally, for example, the MV for the current block can be derived using pattern matching in a surrounding area included in the reference picture and corresponding to a position of the selected best MV candidate. In other words, estimation using pattern matching in the reference picture and evaluation values can be performed in a surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV for the current block. In an embodiment, the update to the MV that produces the better evaluation value can not be performed.
[0715] Finally, the inter predictor 218 generates a prediction image for the current block by performing motion compensation on the current block using the derived MV and the decoded reference picture (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each of all blocks in a slice, the inter prediction of the slice ends using the FRUC mode. For example, when the processes in steps Si_11 to Si_15 are performed for each of all blocks in a picture, the inter prediction of the picture ends using the FRUC mode. Each sub-block can be processed similarly to the case of each block.
[0716] (MV derivation > FRUC mode)
[0717] For example, when information parsed from the stream indicates that the affine merge mode is to be applied, the inter predictor 218 derives an MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0718] Figure 87 is a flowchart showing an example of a process of inter prediction by the affine merge mode in the decoder 200.
[0719] In the affine merge mode, first, the inter predictor 218 derives MVs for the current block at respective control points (step Sk_11). As described above, for example, the inter predictor 218 derives the MVs for the current block at the respective control points by using the MVs of the decoded blocks neighboring the current block as the MV candidates.Figure 46A As shown, the control points are the top-left corner and the top-right corner of the current block, or as... Figure 46B As shown, the control points are the top-left corner, the top-right corner, and the bottom-left corner of the current block.
[0720] For example, when using Figures 47A-47C When the MV derivation method is shown in the figure, such as Figure 47A As shown, the inter-frame predictor 218 sequentially examines the decoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), and identifies the first valid block decoded according to the affine pattern. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47B As shown, the inter-frame predictor 218 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the decoded block including block A. In this way, the MV at each control point is derived.
[0721] Note that, as Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, and as follows: Figure 49B As shown, when block A is identified and when block A has three control points, the MV at two control points can be calculated.
[0722] Additionally, when MV selection information is included in the stream as a prediction parameter, the inter-frame predictor 218 can use the MV selection information to derive the MV for the current block at each control point.
[0723] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to compute an affine MV for each of the multiple sub-blocks (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and an encoded reference image to perform motion compensation for the sub-blocks (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of all sub-blocks included in the current block, the inter-frame prediction using the affine merging mode for the current block ends. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.
[0724] Note that the MV candidate list described above can be generated in step Sk_11. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, any combination of Figures 47A-47C the MV derivation method illustrated in Figure 48A and Figure 48B the MV derivation method illustrated in Figure 49A and Figure 49B the MV derivation method illustrated in and other MV derivation methods.
[0725] Note that the MV candidate list can include MV candidates in modes in which prediction is performed in units of sub-blocks other than the affine mode.
[0726] Note that, for example, an MV candidate list including MV candidates in the affine merge mode in which two control points are used and the affine merge mode in which three control points are used can be generated as the MV candidate list. Alternatively, an MV candidate list including MV candidates in the affine merge mode in which two control points are used and an MV candidate list including MV candidates in the affine merge mode in which three control points are used can be generated separately. Alternatively, an MV candidate list including MV candidates in one of the affine merge mode in which two control points are used and the affine merge mode in which three control points are used can be generated.
[0727] (MV derivation > affine inter mode)
[0728] For example, when information parsed from the stream indicates that the affine inter mode is to be applied, the inter predictor 218 derives an MV in the affine inter mode and performs motion compensation (prediction) using the MV.
[0729] Figure 88 is a flowchart illustrating an example of a procedure of inter prediction by the affine inter mode in the decoder 200.
[0730] In the affine inter mode, first, the inter predictor 218 derives MV predictors (v0, v1) or (v0, v1, v2) of the respective two or three control points of the current block (step Sj_11). The control points are the top-left corner point of the current block, the top-right corner point of the current block, and the bottom-left corner point of the current block, as illustrated in Figure 46A or Figure 46B .
[0731] The inter predictor 218 obtains MV predictor selection information included in the stream as a prediction parameter and derives MV predictors at each control point of the current block using the MV identified by the MV predictor selection information. For example, when using Figure 48A and Figure 48BIn the MV derivation method illustrated in FIG. 6, the inter predictor 218 derives the motion vector predictor (v0, v1) or (v0, v1, v2) at the control points of the current block by selecting Figure 48A or Figure 48B the MV of the block identified by the MV predictor selection information in the decoded block in the vicinity of the corresponding control point of the current block illustrated in FIG. 6.
[0732] Next, the inter predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block to the MV difference corresponding to the MV predictor (step Sj_12). In this way, the MV at each control point for the current block is derived.
[0733] Next, the inter predictor 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter predictor 218 calculates the MV for each of the plurality of sub-blocks as an affine MV using the two motion vectors v0 and v1 and the above expression (1A) or using the three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_13). The inter predictor 218 then performs motion compensation on the sub-blocks using these affine MVs and the encoded reference picture (step Sk_14). When the processes in steps Sk_13 and Sk_14 are performed for each of all the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block ends. In other words, motion compensation on the current block is performed to generate a predicted picture of the current block.
[0734] Note that the MV candidate list described above can be generated in step Sj_11 as in step Sk_11.
[0735] (MV derivation > triangle mode)
[0736] For example, when the information parsed from the stream indicates that the triangle mode is to be applied, the inter predictor 218 derives the MV in the triangle mode and performs motion compensation (prediction) using the MV.
[0737] Figure 89 is a flowchart illustrating an example of the procedure of inter prediction by the triangle mode in the decoder 200.
[0738] In the triangle mode, first, the inter predictor 218 splits the current block into a first partition and a second partition (step Sx_11). For example, the inter predictor 218 can obtain partition information as a prediction parameter from the stream, the partition information being information related to the splitting. The inter predictor 218 can then split the current block into the first partition and the second partition according to the partition information.
[0739] Next, the inter predictor 218 obtains a plurality of MV candidates for the current block based on the information (e.g., MVs of a plurality of decoded blocks around the current block in time or space) (step Sx_12). In other words, the inter predictor 218 generates a list of MV candidates.
[0740] The inter predictor 218 then selects an MV candidate for the first partition and an MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, the inter predictor 218 can obtain, as the prediction parameters, MV selection information for identifying each of the selected MV candidates from the stream. The inter predictor 218 can then select the first MV and the second MV in accordance with the MV selection information.
[0741] Next, the inter predictor 218 generates a first prediction image by performing motion compensation using the selected first MV and the decoded reference picture (step Sx_14). Likewise, the inter predictor 218 generates a second prediction image by performing motion compensation using the selected second MV and the decoded reference picture (step Sx_15).
[0742] Finally, the inter predictor 218 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_16).
[0743] (MV estimation > DMVR)
[0744] For example, the information parsed from the stream indicates that DMVR is to be applied, and the inter predictor 218 performs motion estimation using DMVR.
[0745] Figure 90 is a flowchart illustrating an example of a process of motion estimation by DMVR in the decoder 200.
[0746] The inter predictor 218 derives an MV for the current block in accordance with the merge mode (step Sl_11). Next, the inter predictor 218 derives a final MV for the current block by searching a region around a reference picture indicated by the MV derived in Sl_11 (step Sl_12). In other words, in this case, the MV of the current block is determined in accordance with DMVR.
[0747] Figure 91 is a flowchart illustrating an example of a process of motion estimation by DMVR in the decoder 200, and is the same as Figure 58B
[0748] First, in Figure 58A In step 2 shown in FIG. 6, the inter predictor 218 performs a search similar to the process in step 1, and changes the target to the search position where the minimum cost is obtained according to the result of the process in step 1. Then the inter predictor 218 determines whether the cost at each of the search positions other than the start point is minimum. Here, when it is determined that the cost at one of the search positions other than the start point is minimum, the inter predictor 218 performs the process in step 3. When the cost at the start point is minimum, the inter predictor 218 performs the process in step 4. Figure 58A In step 2 shown in FIG. 6, the inter predictor 218 performs a search similar to the process in step 1, and changes the target to the search position where the minimum cost is obtained according to the result of the process in step 1. Then the inter predictor 218 determines whether the cost at each of the search positions other than the start point is minimum. Here, when it is determined that the cost at one of the search positions other than the start point is minimum, the inter predictor 218 performs the process in step 3. When the cost at the start point is minimum, the inter predictor 218 performs the process in step 4. Figure 58A In step 2 shown in FIG. 6, the inter predictor 218 performs a search similar to the process in step 1, and changes the target to the search position where the minimum cost is obtained according to the result of the process in step 1. Then the inter predictor 218 determines whether the cost at each of the search positions other than the start point is minimum. Here, when it is determined that the cost at one of the search positions other than the start point is minimum, the inter predictor 218 performs the process in step 3. When the cost at the start point is minimum, the inter predictor 218 performs the process in step 4.
[0749] In step 2 shown in FIG. 6, the inter predictor 218 performs a search similar to the process in step 1, and changes the target to the search position where the minimum cost is obtained according to the result of the process in step 1. Then the inter predictor 218 determines whether the cost at each of the search positions other than the start point is minimum. Here, when it is determined that the cost at one of the search positions other than the start point is minimum, the inter predictor 218 performs the process in step 3. When the cost at the start point is minimum, the inter predictor 218 performs the process in step 4. Figure 58A In step 2 shown in FIG. 6, the inter predictor 218 performs a search similar to the process in step 1, and changes the target to the search position where the minimum cost is obtained according to the result of the process in step 1. Then the inter predictor 218 determines whether the cost at each of the search positions other than the start point is minimum. Here, when it is determined that the cost at one of the search positions other than the start point is minimum, the inter predictor 218 performs the process in step 3. When the cost at the start point is minimum, the inter predictor 218 performs the process in step 4.
[0750] In step 4, the inter predictor 218 regards the search position at the start point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0751] In step 3 shown in FIG. 6, the inter predictor 218 determines the pixel position in sub-pixel accuracy at which the minimum cost is obtained based on the costs at the four points located at the upper, lower, left, and right positions with respect to the start point in step 1 or step 2, and regards the pixel position as the final search position. Figure 58A The pixel position in sub-pixel accuracy is determined by performing a weighted addition using the costs at the corresponding one of the four search positions as weights, for each of the four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)). The inter predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0752] (Motion compensation > BIO / OBMC / LIC)
[0753]
[0754] For example, when information parsed from the stream indicates that correction of the prediction image is to be performed, the inter-frame predictor 218 corrects the prediction image based on a mode for the correction at the time of generation of the prediction image. The mode is, for example, one of the BIO, the OBMC, and the LIC described above.
[0755] Figure 92 is a flowchart showing one example of a process of generating a prediction image in the decoder 200.
[0756] The inter-frame predictor 218 generates a prediction image (step Sm_11), and corrects the prediction image according to any of the modes described above (step Sm_12).
[0757] Figure 93 is a flowchart showing another example of a process of generating a prediction image in the decoder 200.
[0758] The inter-frame predictor 218 derives an MV for the current block (step Sn_11). Next, the inter-frame predictor 218 generates a prediction image using the MV (step Sn_12), and determines whether or not to perform a correction process (step Sn_13). For example, the inter-frame predictor 218 obtains a prediction parameter included in the stream, and determines whether or not to perform the correction process based on the prediction parameter. For example, the prediction parameter is a flag indicating whether or not one or more of the modes described above is to be applied. Here, when it is determined that the correction process is to be performed (Yes in step Sn_13), the inter-frame predictor 218 generates a final prediction image by correcting the prediction image (step Sn_14). Note that, in the LIC, the luminance and the chrominance can be corrected in step Sn_14. When it is determined that the correction process is not to be performed (No in step Sn_13), the inter-frame predictor 218 outputs the final prediction image without correcting the prediction image (step Sn_15).
[0759] (Motion compensation > OBMC)
[0760] For example, when information parsed from the stream indicates that the OBMC is to be performed, the inter-frame predictor 218 corrects the prediction image according to the OBMC at the time of generation of the prediction image.
[0761] Figure 94 is a flowchart showing an example of a process of correction of a prediction image by the OBMC in the decoder 200. Note that, Figure 94 The flowchart in Figure 62 indicates a correction flow of a prediction image using the current picture and the reference picture shown in
[0762] First, as shown in Figure 62 , the inter-frame predictor 218 obtains a prediction image (Pred) by performing normal motion compensation using an MV assigned to the current block.
[0763] Next, the inter predictor 218 obtains a prediction image (Pred_L) by applying a motion vector (MV_L) that has been derived for an encoded block neighboring to the left of the current block to the current block (reusing the motion vector for the current block). Then, the inter predictor 218 performs a first correction of the prediction image by overlapping the two prediction images Pred and Pred_L. This provides an effect of blending the boundary between the neighboring blocks.
[0764] Likewise, the inter predictor 218 obtains a prediction image (Pred_U) by applying a motion vector (MV_U) that has been derived for an encoded block neighboring above the current block to the current block (reusing the motion vector for the current block). Then, the inter predictor 218 performs a second correction of the prediction image by overlapping the prediction image Pred_U with the prediction image that has been subjected to the first correction (e.g., Pred and Pred_L). This provides an effect of blending the boundary between the neighboring blocks. The prediction image obtained by the second correction is an image in which the boundary between the neighboring blocks has been blended (smoothed), and is thus the final prediction image of the current block.
[0765] (Motion Compensated > BIO)
[0766] For example, when information parsed from the stream indicates that BIO is to be performed, the inter predictor 218 corrects the prediction image according to the BIO when generating the prediction image.
[0767] Figure 95 is a flowchart showing an example of a process of correcting a prediction image by BIO in the decoder 200.
[0768] As Figure 63 shown in FIG. 11, the inter predictor 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) including the current block. Then, the inter predictor 218 derives a prediction image for the current block using the two motion vectors (M0, M1) (step Sy_11). Note that the motion vector M0 is a motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.
[0769] Next, the inter predictor 218 derives an interpolated image I 0Additionally, the inter predictor 218 derives an interpolated image I 1 (Step Sy_12). Here, the interpolated image I 0 is an image included in the reference picture Ref0 and derived for the current block, and the interpolated image I 1 is an image included in the reference picture Ref1 and derived for the current block. Each of the interpolated image I 0 and the interpolated image I 1 may be the same size as the current block. Alternatively, each of the interpolated image I 0 and the interpolated image I 1 may be an image larger than the current block. Further, the interpolated image I 0 and the interpolated image I 1 may include a predicted image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.
[0770] Additionally, the inter predictor 218 derives gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) of the current block from the interpolated image I 0 and the interpolated image I 1 (Step Sy_13). Note that the gradient images in the horizontal direction are (Ix 0 , Ix 1 ), and the gradient images in the vertical direction are (Iy 0 , Iy 1 ). The inter predictor 218 can derive the gradient images by, for example, applying a gradient filter to the interpolated images. The gradient images can be images indicating each of a spatial change amount of pixel values in the horizontal direction or a spatial change amount of pixel values in the vertical direction.
[0771] Next, the inter predictor 218 derives optical flows (vx, vy) as velocity vectors for each sub-block of the current block using the interpolated images (I 0 , I 1 ) and the gradient images (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (Step Sy_14). As one example, the sub-block can be a 4x4-pixel sub-CU.
[0772] Next, the inter predictor 218 corrects the prediction image for the current block using the optical flow (vx, vy). For example, the inter predictor 218 derives correction values of the values of the pixels included in the current block using the optical flow (vx, vy) (step Sy_15). The inter predictor 218 can then correct the prediction image for the current block using the correction values (step Sy_16). Note that the correction values can be derived in units of pixels, or can be derived in units of multiple pixels or in units of sub-blocks.
[0773] Note that the BIO process flow is not limited to the process disclosed in Figure 95 the present disclosure. Only a part of the process disclosed in the present disclosure can be executed, or a different process can be added or used instead, or the processes can be executed in a different processing order, and the like. Figure 95
[0774] (motion compensation > LIC)
[0775] For example, when information parsed from the stream indicates that LIC is to be executed, the inter predictor 218 corrects the prediction image according to the LIC at the time of generation of the prediction image.
[0776] Figure 96 is a flowchart showing an example of a process of correction of a prediction image by LIC in the decoder 200.
[0777] First, the inter predictor 218 obtains a reference image corresponding to the current block from a decoded reference picture using the MV (step Sz_11).
[0778] Next, the inter predictor 218 extracts information indicating how the illumination value changes between the current picture and the reference picture for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values of the encoded left-side neighboring reference region (surrounding reference region) and the encoded upper neighboring reference region (surrounding reference region), and the luminance pixel values at the corresponding positions in the reference picture specified by the derived MV. The inter predictor 218 calculates the illumination correction parameter using the information indicating how the luminance value changes (step Sz_13).
[0779] The inter predictor 218 generates the prediction image for the current block by performing an illumination correction process in which the illumination correction parameter is applied to the reference image in the reference picture specified by the MV (step Sz_14). In other words, the prediction image is corrected based on the illumination correction parameter, which is the reference image in the reference picture specified by the MV. In this correction, the illumination can be corrected, or the chrominance can be corrected.
[0780] (prediction controller)
[0781] The prediction controller 220 selects either an intra-predicted picture or an inter-predicted picture and outputs the selected picture to the summer 208. In general, the configuration, functions, and processes of the prediction controller 220, the intra-predictor 216, and the inter-predictor 218 on the decoder 200 side can correspond to the configuration, functions, and processes of the prediction controller 128, the intra-predictor 124, and the inter-predictor 126 on the encoder 100 side.
[0782] (Decoding using predicted chroma samples)
[0783] In a first aspect, it is determined whether a luma sample can be used to predict a chroma sample block of a current block, where the block is decoded using the predicted chroma sample. For example, embodiments can employ a process that uses a decoded result of an illumination signal in a decoding method or an encoding method to determine whether to enable a tool such as CCLM to predict a color difference signal.
[0784] Figure 97 is a flowchart illustrating one example of a process 1000 of decoding a block using predicted chroma samples, which may, for example, be performed by Figure 7 the encoder 100 of Figure 67 the decoder 200 of Figure 67 . For convenience, the process 1000 will be described with reference to Figure 97 the decoder 200 of
[0785] At S1001, the decoder 200 determines whether a current chroma block is inside an MxN non-overlapping region aligned with an MxN grid of chroma samples. Figure 99 and Figure 100 are conceptual diagrams for illustrating examples of determining whether a current chroma block is inside an MxN non-overlapping region aligned with an MxN grid of chroma samples. In some formats (e.g., YUV420 format), a 16x16-pixel region of chroma corresponds to a 32x32-pixel region of luma. As illustrated in Figure 99 and Figure 100 , a chroma block within a 32x32-luma region aligned with a 16x16-chroma grid is determined to be inside an MxN non-overlapping region aligned with an MxN grid of chroma samples. A chroma block not within the 32x32-luma region is not determined to be inside an MxN non-overlapping region aligned with an MxN grid of chroma samples. Even if the current chroma block crosses a boundary of a corresponding luma block, luma samples can be used to obtain chroma samples if it is included in the same VPDU. For example, Figure 99 chroma block A in uses corresponding samples in luma block B to predict samples in chroma block A-1 and uses corresponding samples in luma block C to predict samples in chroma block A-2.
[0786] As illustrated in Figure 100As shown in the middle, the chroma samples of the shown chroma block can be predicted using the luma samples of the chroma block, because the chroma block is included in a grid (as shown, a 16x16 grid) and the collocated luma block is also inside the collocated 32x32 region.
[0787] In some embodiments, for example, by default, when other conditions are met, for example, as discussed below with reference to S1002, etc., luma samples can not be used to predict chroma sample blocks that are not determined to be inside the MxN non-overlapping region that is aligned with the MxN grid of chroma samples, and luma samples can be used to predict chroma sample blocks that are determined to be inside the MxN non-overlapping region that is aligned with the MxN grid of chroma samples.
[0788] As Figure 97 As shown in the middle, the chroma samples of the shown chroma block can be predicted using the luma samples of the chroma block, because the chroma block is included in a grid (as shown, a 16x16 grid) and the collocated luma block is also inside the collocated 32x32 region.
[0789] At S1002, the decoder 200 determines whether to split the current luma VPDU into smaller blocks. A VPDU is a unit that is processed in parallel in the encoding or decoding process, for example, of size 64x64. The size of the VPDU can be determined by a standard, or can be encoded in the stream.
[0790] Whether to split the current luma VPDU into smaller blocks can be determined in various ways, and is discussed in more detail below with reference to Figure 102 and Figure 103 Some examples are discussed in more detail below.
[0791] When it is determined at S1002 that the current luma VPDU is not to be split into smaller blocks, process 1000 proceeds from S1002 to S1004, in which the decoder 200 predicts the block of chroma samples without using luma samples. Process 1000 proceeds from S1004 to S1005, in which the decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1002 that the current luma VPDU is to be split into smaller blocks, process 1000 proceeds from S1002 to S1003, in which the decoder 200 predicts the block of chroma samples using luma samples. Process 1000 proceeds from S1003 to S1005, in which the decoder 200 decodes the block using the predicted chroma samples. In some embodiments, additional considerations can be taken into account to determine whether to decode the block of chroma samples using luma samples, for example, as discussed below with reference to Figures 104-110 .
[0792] Figure 98 is a flowchart showing another example of a process 2000 of decoding a block using predicted chroma samples, which can be performed, for example, by an encoder 100 of Figure 7 or a decoder 200 of Figure 67 . For convenience, the process 2000 will be described with reference to the decoder 200 of Figure 67 . It will be understood that the process 2000 can be performed by the encoder 100 of Figure 98 .
[0793] At S2001, the decoder 200 determines whether to split the first VPDU and the second VPDU into smaller blocks. Whether to split the current luma VPDU into smaller blocks can be determined in various ways, and some examples are discussed in more detail below with reference to Figure 102 and Figure 103 .
[0794] When it is determined at S2001 that the first VPDU is not to be split into smaller blocks and that the second VPDU is to be split into smaller blocks, process 2000 proceeds from S2001 to S2002, in which the decoder 200 predicts the block of chroma samples without using luma samples. Process 2000 proceeds from S2002 to S2004, in which the decoder 200 decodes the block using the predicted chroma samples.
[0795] When it is not determined at S2001 that the first luma VPDU is not to be split into smaller blocks and that the second VPDU is to be split into smaller blocks, the process 2000 proceeds from S2001 to S2003, in which the decoder 200 uses luma samples to predict chroma sample blocks. The process 2000 proceeds from S2003 to S2004, in which the decoder 200 decodes the blocks using the predicted chroma samples. In some embodiments, additional considerations can be taken into account to determine whether to use luma samples to decode chroma sample blocks, for example, as discussed below with reference to Figures 104-110
[0796] Figure 101 is a conceptual diagram for illustrating VPDU. VPDU is non-overlapped, which represents the buffer size of the pipeline stage. Figure 101 The left side (labeled a) of shows an example of a 128x128 CTU with 4 64x64 VPDU. Figure 101 The right side (labeled b) of shows an example of a 128x128 CTU with 16 32x32 VPDU. For example, if the VPDU is 64x64, both M and N are set to 16. When the VPDU is to be further divided, the divided CU size becomes 2Mx2N (32x32) or smaller. In YUV420 format, 16x16 chroma region corresponds to 32x32 luma region, so the pixels in 16x16 grid of chroma can be predicted based on the pixels in the corresponding 32x32 grid of luma. Thus, when the decoding of 32x32 region of luma is completed, the prediction process of luma chroma in 16x16 region of chroma can be started. In the case of YUV444 format, luma MxN region corresponds to chroma MxN region. If 2Mx2N is half of the VPDU size in both horizontal and vertical directions, the prediction process of luma chroma in 2Mx2N region of chroma can be started when the decoding of 2Mx2N region of luma is completed. Figure 97 In step S1002 of or Figure 98 In step S2001 of or
[0797] Figure 102 is a conceptual diagram for illustrating an example of determining whether a current VPDU can use luma samples to predict chroma sample blocks based on whether the luma VPDU is split into blocks, in which the left side shows a luma CTU and the right side shows a corresponding chroma CTU. As shown, luma VPDU0 is to be split into blocks and luma VPDU1 is not to be split into blocks. Thus, with reference to Figure 97 In the process 1000 of FIG. 10, the luma samples can be used to predict the chroma samples of VPDU0, and the luma samples can not be used to predict the chroma samples of VPDU1.
[0798] Figure 103 is a conceptual diagram for illustrating two example ways to determine whether a luma VPDU is to be split into smaller blocks. In Figure 103 the first example (labeled a) shown on the left, whether a luma VPDU is to be split can be determined based on a split flag associated with the luma VPDU. As shown, when the split flag has a value of 1, the VPDU is to be split (and, referring to the process 1000 of FIG. 10, the luma samples are used to predict the chroma samples of the block). When the split flag has a value of 0, the VPDU is not split (and, referring to the process 1000 of FIG. 10, the luma samples are not used to predict the chroma samples of the block). Other split flag values can be employed to determine whether a luma VPDU is split. Figure 97 Figure 97 In the process 1000 of FIG. 10, the luma samples can be used to predict the chroma samples of VPDU0, and the luma samples can not be used to predict the chroma samples of VPDU1.
[0799] In the process 1000 of FIG. 10, the luma samples can be used to predict the chroma samples of VPDU0, and the luma samples can not be used to predict the chroma samples of VPDU1. Figure 103 the second example (labeled b) shown on the right, whether a luma VPDU is to be split can be determined based on a quad-tree split depth of a luma block of the VPDU. As shown, the quad-tree split depth of the luma block of VPDU0 is greater than 1, and thus, referring to the process 1000 of FIG. 10, the luma samples can be used to predict the chroma samples when decoding the block of VPDU0. In contrast, the quad-tree split depth of the block of VPDU1 is less than or equal to 1, and thus, referring to the process 1000 of FIG. 10, the luma samples can not be used to predict the chroma samples when decoding the block of VPDU0. Other split depth values can be employed to determine whether a luma VPDU is split. Figure 97 Figure 97
[0800] Figure 104 is a conceptual diagram for illustrating additional considerations that can be taken into account to determine whether to use luma samples to predict the chroma samples of a block. As shown, whether a current block size is equal to or smaller than a threshold block size can be employed as an additional consideration to determine whether to use luma samples to predict the chroma samples of a block.
[0801] The threshold block size can be a default block size, a signaled block size, or a determined block size, and can be a luma or chroma block size. For example, if the threshold block size is a 16x16 luma block size, the luma block size of VPDU0 is greater than 16x16, and thus it can be determined that luma samples are not used to decide the chroma samples of the block. The threshold block size can be employed at S1002 of the process 1000 of FIG. 10 or at S2001 of the process 2000 of FIG. 20 to determine whether to split a luma VPDU into smaller blocks. Figure 97 Figure 98
[0802] Figure 97 The process of 1000 aspects and Figure 98 The process 2000 can be modified in various ways. For example, process 1000 or 2000 can be modified to perform more actions than shown, can be modified to perform fewer actions than shown, can be modified to perform actions in various orders, can be modified to combine or split actions, etc. For example, before S1001 or S1002, process 1000 can be modified based on other considerations (e.g., reference). Figure 103 The size of the current block (discussed) determines whether to use luminance samples to predict chrominance samples for the block. In another example, process 1000 can be modified to omit S1001. In another example, Figure 98 An embodiment of process 2000 can be modified to perform step S1001 before performing step S2001. In another example, S2001 may determine whether the first VPDU and the second VPDU are split into smaller blocks.
[0803] Figure 105 This is a conceptual diagram illustrating an example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (e.g.) Figure 105 The example combination shown is based on whether both the luma VPDU and the corresponding chroma VPDU have a quadtree split depth greater than or equal to 2. Luma VPDU0 has a quadtree depth greater than or equal to 2, and chroma VPDU0 also has a quadtree split depth greater than or equal to 2, so a luma sample can be used to predict a chroma sample for chroma VPDU0. However, luma VPDU1 has a quadtree split depth not greater than or equal to 2, therefore one of the conditions is not met, and a chroma sample for chroma VPDU1 will be predicted without using a luma sample.
[0804] Figure 106 This is a conceptual diagram illustrating another example of a combination of conditions considered when determining whether to use luminance samples to predict chrominance samples for a block. (See diagram for example.) Figure 106As shown in FIG. 6, an example combination of conditions is: (i) whether the luma VPDUs quad tree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDUs quad tree split depth is equal to 1; and (iii) whether the 32x32 chroma split threshold condition is satisfied (e.g., when the chroma size is 32x32, the block is not split). Luma VPDUs 0 has a quad tree depth greater than or equal to 2, satisfying condition (i); chroma VPDUs 0 has a quad tree split depth equal to 1, satisfying condition (ii), and the chroma VPDUs are not split into blocks smaller than 32x32, so all three conditions are satisfied and the chroma samples for VPDUs 0 can be predicted using luma samples. However, the block for chroma VPDUs 1 is smaller than the 32x32 threshold, so condition (iii) is not satisfied and the chroma samples for VPDUs 1 will be predicted without using luma samples.
[0805] Figure 107 is a conceptual diagram for showing another example of considering a combination of conditions when determining whether to use luma samples to predict chroma samples of a block. As shown in FIG. 7, an example combination of conditions is: (i) whether the luma VPDUs quad tree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDUs quad tree split depth is equal to 1; and (iii) whether the 32x32 chroma split threshold condition is satisfied (e.g., when the chroma size is 32x32, the block is not split). Figure 107 As shown in FIG. 6, an example combination of conditions is: (i) whether the luma VPDUs quad tree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDUs quad tree split depth is equal to 1; and (iii) whether the 32x32 chroma split threshold condition is satisfied (e.g., when the chroma size is 32x32, the block is not split). Figure 107 qtDepthC in FIG. 6 indicates the chroma quad tree split depth, and Figure 107 mtDepthC in FIG. 7 indicates the chroma multi-type tree split depth. The quad tree split can be followed by another quad tree split or a multi-type tree split (binary or ternary split). To specify that the chroma quad tree split terminates at depth 1, the condition chromaSplit32x32 == CU_DONT_SPLIT (referring to Figure 106 condition iii) discussed, which means there is no further split at the chroma 32x32 level. Assuming the luma quad tree split depth qtDepthl is greater than or equal to 2, only chroma VPDUs 0 satisfies all three conditions and the chroma samples for VPDUs 0 can be predicted using luma samples. Chroma VPDUs 1 has a chroma quad tree split depth of 2 and the block is split into blocks smaller than 32x32, so the chroma samples for VPDUs 1 will be predicted without using luma samples. Chroma VPDUs 2 has a chroma quad tree split depth of 1 but the block is split into blocks smaller than 32x32, so the chroma samples for VPDUs 2 will be predicted without using luma samples. Chroma VPDUs 3 has a chroma quad tree split depth of 1 but the block is split into blocks smaller than 32x32, so the chroma samples for VPDUs 3 will be predicted without using luma samples.
[0806] Figure 108 is a conceptual diagram for illustrating another example of a combination of conditions considered when determining whether to use luma samples to predict chroma samples of a block. In the example, the combination of conditions is: (i) whether the luma VPDU quadtree split depth is greater than or equal to 2; (ii) whether the corresponding chroma VPDU quadtree split depth is equal to 1 ; and (iii) whether there is no vertical or horizontal ternary split after a horizontal chroma split of size 32x32. For VPDU0, the conditions are satisfied; VPDU0 is split horizontally into two 16x32 blocks, and these blocks are not further split using horizontal or vertical ternary splits. Thus, luma samples can be used to predict the chroma samples in all blocks of VPDU0. For VPDU1, the lower 16x32 block satisfies the conditions, is not further ternary split, and luma samples can be used to predict the chroma samples of the lower 16x32 block. The upper 16x32 block of VPDU1 does not satisfy the conditions because there is a further vertical ternary split, and the chroma samples of the upper 16x32 block of VPDU1 will be predicted without using luma samples.
[0807] Figure 109 is a c...
Claims
1. An encoder comprising: circuitry; and a memory coupled to the circuitry; wherein the circuitry, in operation, performs the following: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; in response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples using luma samples; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting a block of chroma samples using luma samples; and encoding the block using the predicted chroma samples, wherein the first virtual pipeline decoding unit is a luma virtual pipeline decoding unit and the second virtual pipeline decoding unit is a chroma virtual pipeline decoding unit, wherein determining whether to split a virtual pipeline decoding unit into smaller blocks is based on a split flag or a block split depth or a threshold block size.
2. A decoder comprising: circuitry; a memory coupled to the circuitry; wherein the circuitry, in operation, performs the following: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples without using luma samples; in response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting a block of chroma samples using luma samples; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting a block of chroma samples using luma samples; and decoding the block using the predicted chroma samples, wherein the first virtual pipeline decoding unit is a luma virtual pipeline decoding unit and the second virtual pipeline decoding unit is a chroma virtual pipeline decoding unit, wherein determining whether to split a virtual pipeline decoding unit into smaller blocks is based on a split flag or a block split depth or a threshold block size.
3. A non-transitory medium storing a bitstream and readable by a computer, the bitstream comprising syntax for causing the computer to perform a decoding process, the decoding process comprising: determining whether a first virtual pipeline decoding unit (VPDU) is split into smaller blocks and whether a second virtual pipeline decoding unit is split into smaller blocks; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples without using luma samples; in response to determining that the first virtual pipeline decoding unit is split into smaller blocks and determining that the second virtual pipeline decoding unit is split into smaller blocks, predicting the block of chroma samples using luma samples; in response to determining that the first virtual pipeline decoding unit is not split into smaller blocks and determining that the second virtual pipeline decoding unit is not split into smaller blocks, predicting the block of chroma samples using luma samples; and decoding the block using the predicted chroma samples, wherein the first virtual pipeline decoding unit is a luma virtual pipeline decoding unit and the second virtual pipeline decoding unit is a chroma virtual pipeline decoding unit, wherein determining whether to split a virtual pipeline decoding unit into smaller blocks is based on a split flag or a block split depth or a threshold block size.