Systems and methods for video coding
By splitting the luma VPDU as needed in the encoder and decoder and using luma samples to predict blocks of chroma samples, the difficulties in improving video coding efficiency and image quality in the existing technology are solved, and more efficient resource utilization and smaller circuit scale are achieved.
Patent Information
- Application Number
- CN202511206135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-17
- Filing Date
- 2020-05-15
- Publication Date
- 2025-10-10
AI Technical Summary
Existing video coding technologies have difficulty improving coding efficiency, enhancing image quality, and reducing circuit scale when processing the ever-increasing amount of digital video data.
Encoding and decoding of blocks is achieved by determining in the encoder and decoder whether to split the luma virtual pipeline decoding unit (VPDU) into smaller blocks and using luma samples to predict blocks of chroma samples when appropriate.
The coding efficiency is improved, the image quality is enhanced, the utilization of processing resources is reduced, and the circuit scale is reduced.
Smart Images

Figure CN120769040A_ABST
Abstract
Description
[0001] This application is a divisional application of the same named patent application filed on 15 May 2020 with the application number 202080028161.5. TECHNICAL FIELD
[0002] The present invention relates to video coding, and more specifically to video coding and decoding systems, components and methods in video coding and decoding, such as for performing encoding of a block using predicted chroma samples. BACKGROUND
[0003] With the advancement of video coding technology, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding), there is still a continuous need to provide improvements and optimizations to video coding technology to handle the ever-increasing amount of digital video data in various applications. The present disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in performing encoding of a block using predicted chroma samples. SUMMARY
[0004] In an aspect, an encoder includes circuitry and a memory coupled to the circuitry. The circuitry determines whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. When it is determined not to split the current luma VPDU into smaller blocks, the circuitry predicts blocks of chroma samples without using luma samples. When it is determined to split the luma VPDU into smaller blocks, the circuitry predicts blocks of chroma samples using luma samples. The circuitry encodes the blocks using the predicted chroma samples.
[0005] In an aspect, an encoder includes a block splitter that, in operation, splits a first picture into a plurality of blocks, an intra predictor that, in operation, predicts a block included in the first picture using a reference block included in the first picture, an inter predictor that, in operation, predicts a block included in the first picture using a reference block included in a second picture different from the first picture, a loop filter that, in operation, filters a block included in the first picture, a transformer that, in operation, transforms a prediction error between an original signal and a prediction signal generated by the intra predictor or the inter predictor to generate transform coefficients, a quantizer that, in operation, quantizes the transform coefficients to generate quantized coefficients, and an entropy encoder that, in operation, variable encodes the quantized coefficients to generate an encoded bitstream including encoded quantized coefficients and control information. Predicting a block includes determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining not to split the current luma VPDU into smaller blocks, predicting a block of chroma samples without using luma samples. In response to determining to split the luma VPDU into smaller blocks, predicting a block of chroma samples using luma samples.
[0006] In an aspect, a decoder includes circuitry and a memory coupled to the circuitry. The circuitry determines whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. When determining not to split the current luma VPDU into smaller blocks, the circuitry predicts a block of chroma samples without using luma samples. When determining to split the luma VPDU into smaller blocks, the circuitry predicts a block of chroma samples using luma samples. The circuitry decodes the block using the predicted chroma samples.
[0007] In an aspect, a decoding device includes a decoder that, in operation, decodes an encoded bitstream to output quantized coefficients, an inverse quantizer that, in operation, inverse quantizes the quantized coefficients to output transform coefficients, an inverse transformer that, in operation, inverse transforms the transform coefficients to output a prediction error, an intra predictor that, in operation, predicts a block included in a first picture using a reference block included in the first picture, an inter predictor that, in operation, predicts a block included in the first picture using a reference block included in a second picture different from the first picture, a loop filter that, in operation, filters a block included in the first picture, and an output that, in operation, outputs a picture including the first picture. Predicting a block includes determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining not to split the current luma VPDU into smaller blocks, predicting a block of chroma samples without using luma samples. In response to determining to split the luma VPDU into smaller blocks, predicting a block of chroma samples using luma samples.
[0008] In an aspect, an encoding method includes determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining not to split the current luma VPDU into smaller blocks, predicting blocks of chroma samples without using luma samples. In response to determining to split the luma VPDU into smaller blocks, predicting blocks of chroma samples using luma samples. Encoding the blocks using the predicted chroma samples.
[0009] In an aspect, a decoding method includes determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining not to split the current luma VPDU into smaller blocks, predicting blocks of chroma samples without using luma samples. In response to determining to split the luma VPDU into smaller blocks, predicting blocks of chroma samples using luma samples. Decoding the blocks using the predicted chroma samples.
[0010] In video encoding techniques, it is desirable to propose new methods to improve encoding efficiency, enhance image quality, and reduce circuit size. Some implementations of embodiments of the present disclosure, including constituent elements of embodiments of the present disclosure considered individually or in various combinations, can facilitate one or more of the following: improving encoding efficiency, enhancing image quality, reducing utilization of processing resources associated with encoding / decoding, reducing circuit size, improving processing speed of encoding / decoding, and the like.
[0011] Furthermore, some implementations of embodiments of the present disclosure, including constituent elements of embodiments of the present disclosure considered individually or in various combinations, can facilitate encoding and decoding, suitable selection of one or more elements (e.g., filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations). It should be noted that the present disclosure includes disclosure of configurations and methods that can provide advantages beyond those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing increases in processing resource usage.
[0012] Additional benefits and advantages of the disclosed embodiments will become apparent to those of ordinary skill in the art upon reading and understanding the following detailed description and accompanying drawings. Benefits and / or advantages can be realized independently and / or in any combination with one another.
[0013] It should be noted that general or specific embodiments can be implemented as a system, a method, an integrated circuit, a computer program, a storage medium, or any selective combination thereof. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1is a schematic diagram showing one example of a functional configuration of a transmission system according to an embodiment.
[0015] Figure 2 is a conceptual diagram for showing one example of a hierarchical structure of data in a stream.
[0016] Figure 3 is a conceptual diagram for showing one example of a slice configuration.
[0017] Figure 4 is a conceptual diagram for showing one example of a tile configuration.
[0018] Figure 5 is a conceptual diagram for showing one example of an encoding structure in scalable coding.
[0019] Figure 6 is a conceptual diagram for showing one example of an encoding structure in scalable coding.
[0020] Figure 7 is a block diagram showing a functional configuration of an encoder according to an embodiment.
[0021] Figure 8 is a functional block diagram showing one example of an installation of an encoder.
[0022] Figure 9 is a flowchart showing one example of an overall encoding process performed by an encoder.
[0023] Figure 10 is a conceptual diagram for showing one example of block splitting.
[0024] Figure 11 is a block diagram showing one example of a functional configuration of a splitter according to an embodiment.
[0025] Figure 12 is a conceptual diagram for showing one example of a splitting mode.
[0026] Figure 13A is a conceptual diagram for showing one example of a syntax tree of a splitting mode.
[0027] Figure 13B is a conceptual diagram for showing another example of a syntax tree of a splitting mode.
[0028] Figure 14 is a chart showing example transform basis functions for various transform types.
[0029] Figure 15 is a conceptual diagram for showing an example spatial-variation transform (SVT).
[0030] Figure 16is a flowchart showing one example of a process performed by a transformer.
[0031] Figure 17 is a flowchart showing another example of a process performed by a transformer.
[0032] Figure 18 is a block diagram showing one example of a functional configuration of a quantizer according to an embodiment.
[0033] Figure 19 is a flowchart showing one example of a quantization process performed by a quantizer.
[0034] Figure 20 is a block diagram showing one example of a functional configuration of an entropy encoder according to an embodiment.
[0035] Figure 21 is a conceptual diagram for showing an example flow of a context-based adaptive binary arithmetic coding (CABAC) process in an entropy encoder.
[0036] Figure 22 is a block diagram showing one example of a functional configuration of a loop filter according to an embodiment.
[0037] Figure 23A is a conceptual diagram for showing one example of a filter shape used in an adaptive loop filter (ALF).
[0038] Figure 23B is a conceptual diagram for showing another example of a filter shape used in an ALF.
[0039] Figure 23C is a conceptual diagram for showing another example of a filter shape used in an ALF.
[0040] Figure 23D is a conceptual diagram for showing an example flow of a cross-component ALF (CC-ALF).
[0041] Figure 23E is a conceptual diagram for showing an example of a filter shape used in a CC-ALF.
[0042] Figure 23F is a conceptual diagram for showing an example flow of a joint chroma CCALF (JC-CCALF).
[0043] Figure 23G is a table showing example weight exponent candidates that can be employed in a JC-CCALF.
[0044] Figure 24 is a block diagram showing one example of a specific configuration of a loop filter used as a deblocking filter (DBF).
[0045] Figure 25 FIG. 1 is a conceptual diagram for illustrating an example of a concept of a deblocking filter having symmetric filter characteristics with respect to a block boundary.
[0046] Figure 26 FIG. 2 is a conceptual diagram for illustrating a block boundary on which a deblocking filtering process is performed.
[0047] Figure 27 FIG. 3 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0048] Figure 28 FIG. 4 is a flowchart illustrating one example of a process performed by a predictor of an encoder.
[0049] Figure 29 FIG. 5 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0050] Figure 30 FIG. 6 is a flowchart illustrating another example of a process performed by a predictor of an encoder.
[0051] Figure 31 FIG. 7 is a conceptual diagram for illustrating sixty-seven intra prediction modes used in intra prediction in an embodiment.
[0052] Figure 32 FIG. 8 is a flowchart illustrating one example of a process performed by an intra predictor.
[0053] Figure 33 FIG. 9 is a conceptual diagram for illustrating an example of a reference picture.
[0054] Figure 34 FIG. 10 is a conceptual diagram for illustrating an example of a reference picture list.
[0055] Figure 35 FIG. 11 is a flowchart illustrating an example basic processing flow of inter prediction.
[0056] Figure 36 FIG. 12 is a flowchart illustrating one example of a process of derivation of a motion vector.
[0057] Figure 37 FIG. 13 is a flowchart illustrating another example of a process of derivation of a motion vector.
[0058] Figure 38A FIG. 14 is a conceptual diagram for illustrating example representations of MV derivation modes.
[0059] Figure 38B FIG. 15 is a conceptual diagram for illustrating example representations of MV derivation modes.
[0060] Figure 39 FIG. 16 is a flowchart illustrating an example of an inter prediction process in a normal inter mode.
[0061] Figure 40 is a flowchart illustrating an example of inter prediction process in normal merge mode.
[0062] Figure 41 is a conceptual diagram for illustrating one example of motion vector derivation process in merge mode.
[0063] Figure 42 is a conceptual diagram for illustrating one example of MV derivation process of HMVP merge mode for current picture.
[0064] Figure 43 is a flowchart illustrating one example of frame rate up conversion (FRUC) process.
[0065] Figure 44 is a conceptual diagram for illustrating one example of pattern matching between two blocks along a motion trajectory (bi-lateral matching).
[0066] Figure 45 is a conceptual diagram for illustrating one example of pattern matching between a template in current picture and a block in reference picture (template matching).
[0067] Figure 46A is a conceptual diagram for illustrating one example of deriving motion vector of each sub-block based on motion vectors of multiple neighboring blocks.
[0068] Figure 46B is a conceptual diagram for illustrating one example of deriving motion vector of each sub-block in affine mode where three control points are used.
[0069] Figure 47A is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0070] Figure 47B is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0071] Figure 47C is a conceptual diagram for illustrating example MV derivation at control points in affine mode.
[0072] Figure 48A is a conceptual diagram for illustrating affine mode where two control points are used.
[0073] Figure 48B is a conceptual diagram for illustrating affine mode where three control points are used.
[0074] Figure 49Ais a conceptual diagram for illustrating one example of a method of MV derivation at a control point when the number of control points used for an encoded block and the number of control points used for a current block are different from each other.
[0075] Figure 49B is a conceptual diagram for illustrating another example of a method of MV derivation at a control point when the number of control points used for an encoded block and the number of control points used for a current block are different from each other.
[0076] Figure 50 is a flowchart illustrating one example of a process in an affine merge mode.
[0077] Figure 51 is a flowchart illustrating one example of a process in an affine inter mode.
[0078] Figure 52A is a conceptual diagram for illustrating generation of two triangular prediction images.
[0079] Figure 52B is a conceptual diagram for illustrating a first portion of a first partition that overlaps a second partition and a first and second set of samples that can be weighted as part of a correction process.
[0080] Figure 52C is a conceptual diagram for illustrating a first portion of a first partition that is a portion of the first partition that overlaps a portion of a neighboring partition.
[0081] Figure 53 is a flowchart illustrating one example of a process in a triangular mode.
[0082] Figure 54 is a conceptual diagram for illustrating one example of an advanced temporal motion vector prediction (ATMVP) mode in which MVs are derived in sub-block units.
[0083] Figure 55 is a flowchart illustrating a relationship between a merge mode and a dynamic motion vector refresh (DMVR).
[0084] Figure 56 is a conceptual diagram for illustrating one example of DMVR.
[0085] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining an MV.
[0086] Figure 58A is a conceptual diagram for illustrating one example of motion estimation in DMVR.
[0087] Figure 58B is a flowchart illustrating one example of a process of motion estimation in DMVR.
[0088] Figure 59 is a flowchart showing one example of a process of generating a prediction image.
[0089] Figure 60 is a flowchart showing another example of a process of generating a prediction image.
[0090] Figure 61 is a flowchart showing one example of a process of correction of a prediction image by overlapped block motion compensation (OBMC).
[0091] Figure 62 is a conceptual diagram for showing one example of a prediction image correction process by OBMC.
[0092] Figure 63 is a conceptual diagram for showing a model assuming uniform straight-line motion.
[0093] Figure 64 is a flowchart showing one example of an inter prediction process according to BIO.
[0094] Figure 65 is a functional block diagram showing one example of a functional configuration of an inter predictor that can perform inter prediction according to BIO.
[0095] Figure 66A is a conceptual diagram for showing one example of a process of a prediction image generation method using a brightness correction process performed by LIC.
[0096] Figure 66B is a flowchart showing one example of a process of a prediction image generation method using LIC.
[0097] Figure 67 is a block diagram showing a functional configuration of a decoder according to an embodiment.
[0098] Figure 68 is a functional block diagram showing an installation example of a decoder.
[0099] Figure 69 is a flowchart showing one example of an overall decoding process performed by a decoder.
[0100] Figure 70 is a conceptual diagram for showing a relationship between a split determiner and other constituent elements.
[0101] Figure 71 is a block diagram showing one example of a functional configuration of an entropy decoder.
[0102] Figure 72 is a conceptual diagram for showing an example flow of a CABAC process in an entropy decoder.
[0103] Figure 73 is a block diagram showing one example of a functional configuration of an inverse quantizer.
[0104] Figure 74 is a flowchart showing one example of an inverse quantization process performed by the inverse quantizer.
[0105] Figure 75 is a flowchart showing one example of a process performed by the inverse transformer.
[0106] Figure 76 is a flowchart showing another example of a process performed by the inverse transformer.
[0107] Figure 77 is a block diagram showing one example of a functional configuration of a loop filter.
[0108] Figure 78 is a flowchart showing one example of a process performed by the predictor of the decoder.
[0109] Figure 79 is a flowchart showing another example of a process performed by the predictor of the decoder.
[0110] Figure 80 is a flowchart showing another example of a process performed by the predictor of the decoder.
[0111] Figure 81 is a diagram showing one example of a process performed by the intra predictor of the decoder.
[0112] Figure 82 is a flowchart showing one example of an MV derivation process in the decoder.
[0113] Figure 83 is a flowchart showing another example of an MV derivation process in the decoder.
[0114] Figure 84 is a flowchart showing an example of an inter prediction process by normal inter mode in the decoder.
[0115] Figure 85 is a flowchart showing an example of an inter prediction process by normal merge mode in the decoder.
[0116] Figure 86 is a flowchart showing an example of an inter prediction process by FRUC mode in the decoder.
[0117] Figure 87 is a flowchart showing an example of an inter prediction process by affine merge mode in the decoder.
[0118] Figure 88is a flowchart illustrating an example of an inter-prediction process by affine inter mode in a decoder.
[0119] Figure 89 is a flowchart illustrating an example of an inter-prediction process by triangle mode in a decoder.
[0120] Figure 90 is a flowchart illustrating an example of a process of motion estimation by DMVR in a decoder. Figure 91 is a flowchart illustrating an example process of motion estimation by DMVR in a decoder.
[0121] Figure 92 is a flowchart illustrating an example of a process of generating a prediction image in a decoder.
[0122] Figure 93 is a flowchart illustrating another example of a process of generating a prediction image in a decoder.
[0123] Figure 94 is a flowchart illustrating an example of a process of correction of a prediction image by OBMC in a decoder.
[0124] Figure 95 is a flowchart illustrating an example of a process of correction of a prediction image by BIO in a decoder.
[0125] Figure 96 is a flowchart illustrating an example of a process of correction of a prediction image by LIC in a decoder.
[0126] Figure 97 is a flowchart illustrating an example of a process of decoding a block using predicted chroma samples.
[0127] Figure 98 is a conceptual diagram for illustrating an example of determining whether a current chroma block is inside an MxN non-overlapping region aligned with chroma samples of an MxN grid.
[0128] Figure 99 is a conceptual diagram for illustrating an example of determining whether a current chroma block is inside an MxN non-overlapping region aligned with chroma samples of an MxN grid.
[0129] Figure 100 is a conceptual diagram for illustrating a virtual pipeline decoding unit (VPDU).
[0130] Figure 101 is a conceptual diagram for illustrating an example of determining whether a current VPDU can be employed to predict a block of chroma samples.
[0131] Figure 102is a conceptual diagram illustrating an example manner of determining whether a luma VPDU is to be split into smaller blocks.
[0132] Figure 103 is a conceptual diagram illustrating the use of a threshold size to determine whether to use luma samples to predict chroma samples of a block.
[0133] Figure 104 is a conceptual diagram for illustrating an example of non-rectangular partitions.
[0134] Figure 105 is a diagram showing an example overall configuration of a content providing system for realizing a content distribution service. Figure 106 is a conceptual diagram for illustrating an example of a display screen of a web page.
[0135] Figure 107 is a conceptual diagram for illustrating an example of a display screen of a web page.
[0136] Figure 108 is a block diagram illustrating one example of a smartphone.
[0137] Figure 109 is a block diagram illustrating an example of a functional configuration of a smartphone. DETAILED DESCRIPTION
[0138] In the drawings, like reference numerals denote like elements unless context dictates otherwise. The sizes and relative positions of elements in the drawings are not necessarily drawn to scale.
[0139] Hereinafter, embodiments will be described with reference to the accompanying drawings. Note that the embodiments described below each illustrate general or specific examples. The numerical values, shapes, materials, components, arrangement and connection of components, steps, relationships between steps, and order of steps, etc., referred to in the following embodiments are merely examples and are not intended to limit the scope of the claims.
[0140] Embodiments of encoders and decoders are described below. The embodiments are examples of encoders and decoders to which the processes and / or configurations presented in the description of aspects of the present disclosure may be applied. The processes and / or configurations may also be implemented in encoders and decoders that are different from the encoders and decoders according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented:
[0141] (1) Any one of the components of the encoder or decoder according to the embodiment presented in the description of the aspects of the present disclosure may be replaced by or combined with another component presented anywhere in the description of the aspects of the present disclosure.
[0142] (2) In an encoder or decoder according to an embodiment, any of the functions or processes performed by one or more components of the encoder or decoder can be changed arbitrarily, e.g., added, replaced, removed, etc. For example, any function or process can be replaced by or combined with another function or process presented anywhere in the description of aspects of the disclosure.
[0143] (3) In a method implemented by an encoder or decoder according to an embodiment, changes can be made as appropriate, e.g., adding, replacing, and removing one or more processes included in the method. For example, any process in the method can be replaced by or combined with another process presented anywhere in the description of aspects of the disclosure.
[0144] One or more components included in an encoder or decoder according to an embodiment can be combined with components presented anywhere in the description of aspects of the disclosure, can be combined with components including one or more functions presented anywhere in the description of aspects of the disclosure, and can be combined with components implementing one or more processes implemented by components presented in the description of aspects of the disclosure.
[0145] Components including one or more functions of an encoder or decoder according to an embodiment, or components implementing one or more processes of an encoder or decoder according to an embodiment, can be combined with or replaced by components presented anywhere in the description of aspects of the disclosure, can be combined with or replaced by components including one or more functions presented anywhere in the description of aspects of the disclosure, or can be combined with or replaced by components implementing one or more processes presented anywhere in the description of aspects of the disclosure.
[0146] In a method implemented by an encoder or decoder according to an embodiment, any of the processes included in the method can be replaced by or combined with a process presented anywhere in the description of aspects of the disclosure, or by or combined with any corresponding or equivalent process.
[0147] One or more processes included in a method implemented by an encoder or decoder according to an embodiment can be combined with a process presented anywhere in the description of aspects of the disclosure.
[0148] Implementation of the processes and / or configurations presented in the description of aspects of the disclosure are not limited to an encoder or decoder according to an embodiment. For example, the processes and / or configurations can be implemented in a device for a different purpose than a motion picture encoder or motion picture decoder disclosed in an embodiment.
[0149] (Term Definitions)
[0150] Various terms can be defined as follows as examples.
[0151] An image is a unit of data composed of a set of pixels, is a picture, or includes a block smaller than a pixel. In addition to video, images also include still images.
[0152] A picture is a unit of image processing configured with a set of pixels, and can also be referred to as a frame or a field. For example, a picture can take the form of an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0153] A block is a unit of processing that is a set of a determined number of pixels. A block can have any number of different shapes. For example, a block can have a rectangular shape of MxN (M columns x N rows) pixels, a square shape of MxM pixels, a triangle, a circle, etc. Examples of blocks include slices, tiles, blocks, CTUs, superblocks, elementary splitting units, VPDU, hardware-oriented processing splitting units, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and sub-blocks. A block can take the form of an MxN array of samples or an MxN array of transform coefficients. For example, a block can be a square or rectangular area of pixels including one luma matrix and two chroma matrices.
[0154] A pixel or a sample is the smallest point of an image. A pixel or a sample includes pixels at integer positions, and pixels at sub-pixel positions, for example, pixels generated based on pixels at integer positions.
[0155] A pixel value or a sample value is a characteristic value of a pixel. A pixel value or a sample value can include one or more of the following: a luma value, a chroma value, an RGB gray level, a depth value, a binary value 0 or 1, etc.
[0156] Chroma or chrominance is the intensity of a color, usually denoted by the symbols Cb and Cr, which specify: a value of a sample array or a single sample value represents the value of one of two color difference signals related to primary colors.
[0157] Luma or luminance is the lightness of an image, usually denoted by the symbols or subscripts Y or L, which specify: a value of a sample array or a single sample value represents the value of a monochrome signal related to primary colors.
[0158] A flag includes one or more bits indicating, for example, a value of a parameter or an index. A flag can be a binary flag, which represents a binary value of the flag, which can also represent a non-binary value of a parameter.
[0159] A signal conveys information, which is signalized or encoded into the signal. A signal includes a discrete digital signal and a continuous analog signal.
[0160] A stream or bitstream is a digital data string of a digital data stream. The stream or bitstream can be one stream or can be configured with multiple streams having multiple hierarchical layers. The stream or bitstream can be transmitted in a serial communication manner using a single transmission path, or can be transmitted in a packet communication manner using multiple transmission paths.
[0161] Difference refers to various mathematical differences such as a simple difference (x-y), an absolute value of a difference (|x-y|), a squared difference (x^2-y^2), a square root of a difference (√(x-y)), a weighted difference (ax-by: a and b are constants), an offset difference (x-y+a: a is an offset), and the like. In the case of a scalar, a simple difference can be sufficient and includes difference calculation.
[0162] Sum refers to various mathematical sums such as a simple sum (x+y), an absolute value of a sum (|x+y|), a squared sum (x^2+y^2), a square root of a sum (√(x+y)), a weighted sum (ax+by: a and b are constants), an offset sum (x+y+a: a is an offset), and the like. In the case of a scalar, a simple sum can be sufficient and includes sum calculation.
[0163] A frame is a combination of a top field and a bottom field, in which sample rows 0, 2, 4,... originate from the top field and sample rows 1, 3, 5,... originate from the bottom field.
[0164] A slice is an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) up to, but not including, the next independent slice segment (if any) within the same access unit.
[0165] A tile is a rectangular region of coding tree blocks within a particular tile column and a particular tile row in a picture. A tile can be a rectangular region of a frame intended to be able to be decoded and coded independently, although circular filtering across tile edges can still be applied.
[0166] A coding tree unit (CTU) can be a coding tree block of luma samples of a picture having three sample arrays, or two corresponding coding tree blocks of chroma samples. Alternatively, a CTU can be a coding tree block of samples of a monochrome picture and one of the pictures encoded using three separate color planes and syntax structures for encoding samples. A superblock can be a square block of 64x64 pixels composed of 1 or 2 mode information blocks, or recursively divided into four 32x32 blocks, which themselves can be further divided.
[0167] (System configuration)
[0168] First, a transmission system according to an embodiment will be described. Figure 1 is a schematic diagram showing one example of a configuration of a transmission system 400 according to an embodiment.
[0169] The transmission system 400 is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. As illustrated, the transmission system 400 includes Figure 1 the encoder 100, the network 300, and the decoder 200.
[0170] An image is input to the encoder 100. The encoder 100 generates a stream by encoding the input image and outputs the stream to the network 300. The stream includes, for example, an encoded image and control information for decoding the encoded image. The image is compressed by encoding.
[0171] It should be noted that the image before being encoded by the encoder 100 is also referred to as an original image, an original signal, or an original sample. The image can be a video or a still image. The image is a general concept of a sequence, a picture, and a block, and thus is not limited to a spatial region having a specific size and a temporal region having a specific size unless otherwise specified. The image is an array of pixels or pixel values, and a signal representing the image or the pixel values is also referred to as a sample. The stream can be referred to as a bitstream, an encoded bitstream, a compressed bitstream, or an encoded signal. Furthermore, the encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by the encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.
[0172] The network 300 transmits the stream generated by the encoder 100 to the decoder 200. The network 300 can be the Internet, a wide area network (WAN), a local area network (LAN), or any combination of networks. The network 300 is not limited to a bidirectional communication network, and can be a unidirectional communication network that transmits a broadcast wave of digital terrestrial broadcasting, satellite broadcasting, or the like. Alternatively, the network 300 can be replaced by a recording medium such as a digital versatile disc (DVD) and a Blu-ray disc (BD) or the like on which a stream is recorded.
[0173] The decoder 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted by the network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method employed by the encoder 100.
[0174] It should be noted that the decoder 200 can also be referred to as an image decoder or a video decoder, and the decoding method performed by the decoder 200 can also be referred to as a decoding method, an image decoding method, or a video decoding method.
[0175] (Data structure)
[0176] Figure 2 is a conceptual diagram for illustrating one example of a hierarchical structure of data in a stream. For convenience, the transmission system 400 of Figure 1 will be described. Figure 2A stream includes, for example, a video sequence. As shown in (a) of FIG. 1, a video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplemental enhancement information (SEI), and a plurality of pictures. Figure 2
[0177] In a video having a plurality of layers, a VPS can include coding parameters common among some of the plurality of layers, and coding parameters related to some of the plurality of layers included in the video or to a single layer.
[0178] An SPS includes parameters for a sequence, i.e., coding parameters referenced by the decoder 200 for decoding the sequence. For example, the coding parameters can indicate a width or height of a picture. It should be noted that there can be a plurality of SPSs.
[0179] A PPS includes parameters for a picture, i.e., coding parameters referenced by the decoder 200 for decoding each picture in the sequence. For example, the coding parameters can include a reference value for a quantization width for decoding the picture and a flag indicating application of weighted prediction. It should be noted that there can be a plurality of SPSs. Each of the SPS and the PPS can be referred to simply as a parameter set.
[0180] As shown in (b) of FIG. 1, a picture can include a picture header and one or more slices. The picture header includes coding parameters referenced by the decoder 200 for decoding the one or more slices. Figure 2 As shown in (c) of FIG. 1, a slice includes a slice header and one or more blocks. The slice header includes coding parameters referenced by the decoder 200 for decoding the one or more blocks.
[0181] Figure 2 As shown in (d) of FIG. 1, a block includes one or more coding tree units (CTUs).
[0182] As shown in (e) of FIG. 1, a CTU includes a CTU header and at least one coding unit (CU). As shown, the CTU includes four coding units CU(10), CU(11), CU(12), and CU(13). The CTU header includes coding parameters referenced by the decoder 200 for decoding the at least one CU. Figure 2 It should be noted that a picture can not include any slice and can include a tile group instead of a slice. In this case, the tile group includes at least one tile. Further, a block can include a slice.
[0183] A CTU can also be referred to as a superblock or a base split unit. As shown in (e) of FIG. 1, a CTU includes a CTU header and at least one coding unit (CU). As shown, the CTU includes four coding units CU(10), CU(11), CU(12), and CU(13). The CTU header includes coding parameters referenced by the decoder 200 for decoding the at least one CU.
[0184] Figure 2 It should be noted that a picture can not include any slice and can include a tile group instead of a slice. In this case, the tile group includes at least one tile. Further, a block can include a slice.
[0185] One CU can be split into multiple smaller CUs. As shown, CU (10) is not split into smaller coding units; CU (11) is split into four smaller coding units CU (110), CU (111), CU (112), and CU (113); CU (12) is not split into smaller coding units; and CU (13) is split into seven smaller coding units CU (1310), CU (1311), CU (1312), CU (1313), CU (132), CU (133), and CU (134). As Figure 2 As shown in (f), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual described later. Although a CU is basically the same as a prediction unit (PU) and a transform unit (TU), it should be noted that, for example, a sub-block transform (SBT) to be described later can include multiple TUs smaller than a CU. Further, a CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. A VPDU is, for example, a fixed unit that can be processed in one stage when pipeline processing is performed in hardware.
[0186] It should be noted that a stream can not include all the hierarchical layers shown in Figure 2 Here, a picture that is a target of a process to be performed by an apparatus such as the encoder 100 or the decoder 200 is referred to as a current picture. When the process is an encoding process, the current picture means a current picture to be encoded; and when the process is a decoding process, the current picture means a current picture to be decoded. Likewise, a CU or a CU block that is a target of a process to be performed by an apparatus such as the encoder 100 or the decoder 200 is referred to as a current block. When the process is an encoding process, the current block means a current block to be encoded; and when the process is a decoding process, the current block means a current block to be decoded.
[0187] (Picture structure: slice / tile)
[0188] A picture can be configured with one or more slice units or one or more tile units to facilitate parallel encoding / decoding of the picture.
[0189] A slice is a basic coding unit included in a picture. A picture can include, for example, one or more slices. Further, a slice includes one or more coding tree units (CTUs).
[0190] Figure 3 is a conceptual diagram for illustrating one example of a slice configuration. For example, in Figure 3In the example of FIG. 6, a picture includes 11x8 CTUs and is split into four slices (slices 1 to 4). Slice 1 includes 16 CTUs, slice 2 includes 21 CTUs, slice 3 includes 29 CTUs, and slice 4 includes 22 CTUs. Here, each CTU in the picture belongs to one of the slices. The shape of each slice is a shape obtained by horizontally splitting the picture. The boundary of each slice does not need to coincide with the picture end and can coincide with any boundary between CTUs in the picture. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, a raster scan order. A slice includes a slice header and coded data. Characteristics of a slice can be written in the slice header. The characteristics can include the CTU address of the top CTU in the slice, the slice type, and so on.
[0191] A tile is a rectangular region unit included in a picture. Tiles of a picture can be assigned numbers called Tileld in a raster scan order.
[0192] Figure 4 is a conceptual diagram for showing one example of a tile configuration. For example, in Figure 4 , a picture includes 11x8 CTUs and is split into four tiles (tiles 1 to 4) of rectangular regions. When tiles are used, the processing order of CTUs can be different from that in a case where tiles are not used. When tiles are not used, a plurality of CTUs in a picture are generally processed in a raster scan order. When a plurality of tiles are used, at least one CTU in each of the plurality of tiles is processed in a raster scan order. For example, as shown in Figure 4 , the processing order of CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then continues from the left end of the second column of tile 1 to the right end of the second column of tile 1.
[0193] It should be noted that one tile can include one or more slices, and one slice can include one or more tiles.
[0194] It should be noted that a picture can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more tiles. A picture can be configured with one of a tile set, a tile group, and a tile. For example, assume that the order in which a plurality of tiles are scanned in a raster scan order for each tile set is the basic encoding order of the tiles. Assume that the set of one or more tiles that are consecutive in the basic encoding order in each tile set is a tile group. Such a picture can be configured by a splitter 102 (see Figure 7 ) described later.
[0195] (Scalable coding)
[0196] Figure 5 and Figure 6is a conceptual diagram showing an example of a scalable stream structure, and for convenience, reference will be made to Figure 1 Provide a description.
[0197] like Figure 5 As shown, encoder 100 can generate a temporally and spatially scalable stream by dividing each of multiple pictures into any of multiple layers and encoding the pictures in the layers. For example, encoder 100 encodes pictures of each layer, thereby achieving scalability when an enhancement layer exists above a base layer. This encoding of each picture is also called scalable coding. In this way, decoder 200 can switch the image quality of the image displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as decoder 200's processing power and external factors such as the status of the communication bandwidth. Therefore, decoder 200 can decode content while freely switching between low and high resolutions. For example, a user of a stream may watch a streamed video on a smartphone on their way home, finishing watching the video and continuing watching the video at home on an internet-connected device (e.g., a TV). It should be noted that each of the smartphones and devices described above includes a decoder 200 with the same or different capabilities. In this case, when the device decodes a layer into a higher layer in the stream, the user can watch the video at high quality at home. In this way, the encoder 100 does not need to generate a plurality of streams having different image qualities for the same content, and thus can reduce the processing load.
[0198] In addition, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 may generate a video whose image quality has been enhanced by performing super-resolution imaging on the pictures in the base layer based on the metadata. Super-resolution imaging may include, for example, an improvement in the signal-to-noise ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information for identifying linear or nonlinear filter coefficients (such as used in a super-resolution process), or information for identifying parameter values in a filtering process, or machine learning, least squares method, etc. used in a super-resolution process.
[0199] In an embodiment, a configuration may be provided in which a picture is divided into, for example, tiles according to the meaning of, for example, an object in the picture. In this case, the decoder 200 can decode only a portion of the area in the picture by selecting the tile to be decoded. In addition, the attributes of the object (person, car, ball, etc.) and the position of the object in the picture (coordinates in the same image) can be stored as metadata. In this case, the decoder 200 is able to identify the position of the desired object based on the metadata and determine the tile that includes the object. For example, Figure 6As shown, metadata can be stored using a data storage structure different from that of image data (e.g., SEI (Supplementary Enhancement Information) messages in HEVC). The metadata indicates, for example, the position, size, or color of the main object.
[0200] The metadata can be stored in units of multiple pictures (e.g., streams, sequences, or random access units). In this way, the decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the picture unit information, it can identify the picture in which the object (person) exists and determine the position of the object in the picture.
[0201] (Encoder)
[0202] An encoder according to an embodiment will be described. Figure 7 1 is a block diagram illustrating a functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes a video in units of blocks.
[0203] like Figure 7 As shown, the encoder 100 is a device for encoding an image in units of blocks, and includes a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a loop filter 120, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown in the figure, the intra-frame predictor 124 and the inter-frame predictor 126 are part of the prediction controller.
[0204] The encoder 100 is implemented as, for example, a general-purpose processor and a memory. In this case, when a software program stored in the memory is executed by the processor, the processor functions as the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, the subtractor 104, the transformer 106, the quantizer 108, the entropy encoder 110, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the intra-frame predictor 124, the inter-frame predictor 126, and the prediction controller 128.
[0205] (Encoder installation example)
[0206] Figure 8 1 is a functional block diagram showing an example of an installation of the encoder 100. The encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The plurality of constituent elements of the encoder 100 shown are mounted on Figure 8 The processor a1 and the memory a2 shown are coupled to each other.
[0207] The processor a1 is a circuit that performs information processing and is coupled to the memory a2. For example, the processor a1 is a dedicated or general electronic circuit that encodes an image. The processor a1 can be a processor, such as a CPU. Further, the processor a1 can be an aggregation of a plurality of electronic circuits. Further, for example, the processor a1 can function as Figure 7 two or more of the plurality of constituent elements of the encoder 100 shown, and the like.
[0208] The memory a2 is a dedicated or general memory for storing information used by the processor a1 to encode an image. The memory a2 can be an electronic circuit and can be connected to the processor a1. Further, the memory a2 can be included in the processor a1. Further, the memory a2 can be an aggregation of a plurality of electronic circuits. In addition, the memory a2 can be a magnetic disk, an optical disk, or the like, or can be denoted as a storage, a recording medium, or the like. Further, the memory a2 can be a non-volatile memory or a volatile memory.
[0209] For example, the memory a2 can store an image to be encoded or a bitstream corresponding to an encoded image. Further, the memory a2 can store a program for causing the processor a1 to encode an image.
[0210] Further, for example, the memory a2 can function as Figure 7 two or more of the plurality of constituent elements of the encoder 100 shown, and the like. For example, the memory a2 can function as Figure 7 the block memory 118 and the frame memory 122 shown in More specifically, the memory a2 can store a reconstructed block, a reconstructed picture, or the like.
[0211] It should be noted that, in the encoder 100, all of the plurality of constituent elements and the like indicated in Figure 7 the processes described herein can not be performed. Figure 7 Part of the constituent elements and the like shown in may be included in another device, or part of the processes described herein can be performed by another device.
[0212] Hereinafter, the overall flow of the processes performed by the encoder 100 is described, and then each constituent element included in the encoder 100 will be described.
[0213] (Overall flow of the encoding processes)
[0214] Figure 9is a flowchart representing one example of an overall encoding process performed by the encoder 100, and will be described with reference to Figure 7 is described.
[0215] First, the splitter 102 of the encoder 100 splits each picture included in an input image into a plurality of blocks having a fixed size (e.g., 128 x 128 pixels) (step Sa_1). The splitter 102 then selects a split mode for the fixed-size block (also referred to as a block shape) (step Sa_2). In other words, the splitter 102 further splits the fixed-size block into a plurality of blocks forming the selected split mode. The encoder 100 performs steps Sa_3 to Sa_9 for each of the plurality of blocks, for the block (i.e., the current block to be encoded).
[0216] The prediction controller 128 and the prediction executor (which includes the intra predictor 124 and the inter predictor 126) generate a prediction image of the current block (step Sa-3). The prediction image can also be referred to as a prediction signal, a prediction block, or a prediction sample.
[0217] Next, the subtracter 104 generates a difference between the current block and the prediction image as a prediction residual (step Sa_4). The prediction residual can also be referred to as a prediction error.
[0218] Next, the transformer 106 transforms the prediction image, and the quantizer 108 quantizes the result to generate a plurality of quantized coefficients (step Sa_5). The plurality of quantized coefficients can sometimes be referred to as a coefficient block.
[0219] Next, the entropy encoder 110 encodes (specifically, entropy-encodes) the plurality of quantized coefficients and prediction parameters related to the generation of the prediction image to generate a stream (step Sa_6). The stream can sometimes be referred to as an encoded bitstream or a compressed bitstream.
[0220] Next, the inverse quantizer 112 inverse-quantizes the plurality of quantized coefficients, and the inverse transformer 114 inverse-transforms the result to restore the prediction residual (step Sa_7).
[0221] Next, the adder 116 adds the prediction image and the restored prediction residual to reconstruct the current block (step Sa_8). In this way, a reconstructed image is generated. The reconstructed image can also be referred to as a reconstructed block or a decoded image block.
[0222] When the reconstructed image is generated, the loop filter 120 performs filtering of the reconstructed image as necessary (step Sa_9).
[0223] The encoder 100 then determines whether the encoding of the entire picture has been completed (step Sa_10). When it is determined that the encoding has not been completed (NO in step Sa_10), the process starting with step Sa_2 is repeatedly performed for the next block of the picture.
[0224] Although the encoder 100 selects one split mode for the fixed-size block and encodes each block according to the split mode in the above-described example, it should be noted that each block can be encoded according to a respective one of a plurality of split modes. In this case, the encoder 100 can evaluate the cost of each of the plurality of split modes, and can select, for example, the stream obtained by encoding according to the split mode that yields the smallest cost as the output stream.
[0225] As illustrated, the processes in steps Sa_1 to Sa_10 are sequentially performed by the encoder 100. Alternatively, two or more processes can be performed in parallel, the processes can be reordered, and so on.
[0226] The encoding process employed by the encoder 100 is hybrid encoding using predictive encoding and transform encoding. Further, the predictive encoding is performed by an encoding loop configured with the subtracter 104, the transformer 106, the quantizer 108, the inverse quantizer 112, the inverse transformer 114, the adder 116, the loop filter 120, the block memory 118, the frame memory 122, the intra predictor 124, the inter predictor 126, and the prediction controller 128. In other words, the prediction executor configured with the intra predictor 124 and the inter predictor 126 is part of the encoding loop.
[0227] (Splitter)
[0228] The splitter 102 splits each picture included in the original image into a plurality of blocks, and outputs each block to the subtracter 104. For example, the splitter 102 first splits a picture into blocks of a fixed size (e.g., 128 x 128 pixels). Other fixed block sizes can be employed. The fixed-size blocks are also referred to as coding tree units (CTUs). The splitter 102 then splits each fixed-size block into blocks of variable sizes (e.g., 64 x 64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, the splitter 102 selects a split mode. The variable-size blocks can also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that, in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in a picture can be processed in units of CUs, PUs, or TUs.
[0229] Figure 10 is a conceptual diagram for illustrating one example of block splitting according to an embodiment. In Figure 10In this case, the solid lines indicate block boundaries of blocks split by quadtree block splitting, and the dashed lines indicate block boundaries of blocks split by binary tree block splitting.
[0230] Here, the block 10 is a square block having 128 x 128 pixels (128 x 128 block). The 128 x 128 block 10 is first split into four square 64 x 64 pixel blocks (quadtree block splitting).
[0231] The 64 x 64 pixel block at the upper left is further split vertically into two rectangular 32 x 64 pixel blocks, and the 32 x 64 pixel block at the left is further split vertically into two rectangular 16 x 64 pixel blocks (binary tree block splitting). As a result, the 64 x 64 pixel block at the upper left is split into two 16 x 64 pixel blocks 11 and 12, and one 32 x 64 pixel block 13.
[0232] The 64 x 64 pixel block at the upper right is split horizontally into two rectangular 64 x 32 pixel blocks 14 and 15 (binary tree block splitting).
[0233] The 64 x 64 pixel block at the lower left is first split into four square 32 x 32 pixel blocks (quadtree block splitting). The upper left block and the lower right block of the four square 32 x 32 pixel blocks are further split. The 32 x 32 pixel block at the upper left is split vertically into two rectangular 16 x 32 pixel blocks, and the 16 x 32 pixel block at the right is further split horizontally into two 16 x 16 pixel blocks (binary tree block splitting). The square 32 x 32 pixel block at the upper right is split horizontally into two rectangular 32 x 16 pixel blocks (binary tree block splitting). The square 32 x 32 pixel block at the upper right is split horizontally into two rectangular 32 x 16 pixel blocks (binary tree block splitting). As a result, the 64 x 64 pixel block at the lower left is split into a rectangular 16 x 32 pixel block 16, two square 16 x 16 pixel blocks 17 and 18, two square 32 x 32 pixel blocks 19 and 20, and two rectangular 32 x 16 pixel blocks 21 and 22.
[0234] The 64 x 64 pixel block 23 at the lower right is not split.
[0235] As described above, in Figure 10 In this case, the block 10 is split into 13 variable size blocks 11 to 23 based on recursive quadtree and binary tree block splitting. This type of splitting is also referred to as quadtree plus binary tree (QTBT) splitting.
[0236] It should be noted that, in Figure 10In the HEVC, one block is split into four or two blocks (quad-tree or binary-tree block split), but the split is not limited to these examples. For example, one block can be split into three blocks (ternary block split). The split including such ternary block split is also referred to as multi-type tree (MBT) split.
[0237] Figure 11 is a block diagram showing one example of a functional configuration of a splitter according to one embodiment. As shown in Figure 11 the splitter 102 can include a block split determiner 102a. As an example, the block split determiner 102a can perform the following process.
[0238] For example, the block split determiner 102a can obtain or retrieve block information from the block memory 118 and / or the frame memory 122, and determine a split pattern (e.g., the above-described split patterns) based on the block information. The splitter 102 splits the original image according to the split pattern, and outputs at least one block resulting from the split to the subtracter 104.
[0239] Further, for example, the block split determiner 102a outputs one or more parameters indicating the determined split pattern (e.g., the above-described split patterns) to the transformer 106, the inverse transformer 114, the intra predictor 124, the inter predictor 126, and the entropy encoder 110. The transformer 106 can transform the predicted residual based on the one or more parameters. The intra predictor 124 and the inter predictor 126 can generate a predicted image based on the one or more parameters. Further, the entropy encoder 110 can entropy-encode the one or more parameters.
[0240] The parameters related to the split pattern can be written in a stream as shown below, to name one example.
[0241] Figure 12 is a conceptual diagram for showing examples of split patterns. The examples of split patterns include: split into four regions (QT), in which one block is divided into two regions horizontally and vertically; split into three regions (HT or VT), in which one block is split in the same direction at a ratio of 1:2:1; split into two regions (HB or VB), in which one block is split in the same direction at a ratio of 1:1; and no split (NS).
[0242] It should be noted that the split pattern does not have a block split direction in the case of split into four regions and no split, whereas the split pattern has split direction information in the case of split into two regions or three regions.
[0243] Figure 13A is a conceptual diagram for showing one example of a syntax tree of a split pattern.
[0244] Figure 13Bis a conceptual diagram for showing another example of a syntax tree of a split mode.
[0245] Figure 13A and Figure 13B is a conceptual diagram for showing an example of a syntax tree of a split mode. In Figure 13A the example, first is information indicating whether or not to split (S: split flag), next is information indicating whether or not to split into four regions (QT: QT flag). Next is information indicating whether to perform splitting into two regions or splitting into three regions (TT: TT flag, or BT: BT flag), and then is information indicating a direction of splitting (Ver: vertical flag, or Hor: horizontal flag). Note that each of the at least one block divided in such a division manner can be further divided repeatedly in a similar procedure. In other words, as an example, whether or not to split, whether or not to split into four regions, which of the horizontal direction and the vertical direction is a direction of a splitting method performed, and which of splitting into three regions or splitting into two regions is to be performed can be determined recursively, and the determination result can be encoded in a stream in the order of coding disclosed by the syntax tree shown in Figure 13A
[0246] In addition, although the information items indicating S, QT, TT, and Ver, respectively, are arranged in the order listed in the syntax tree shown in Figure 13A , the information items indicating S, QT, Ver, and BT, respectively, can be arranged in the order listed. That is, in the example of Figure 13B , first is information indicating whether or not to split (S: split flag), next is information indicating whether or not to split into four regions (QT: QT flag). Next is information indicating a direction of splitting (Ver: vertical flag, or Hor: horizontal flag), and then is information indicating whether to perform splitting into two regions or splitting into three regions (BT: BT flag, or TT: TT flag).
[0247] Note that the above-described split manner is an example, and a split manner other than the above-described split manner can be used, or a part of the above-described split manner can be used.
[0248] (subtracter)
[0249] The subtracter 104 subtracts a prediction image (a prediction sample input from the prediction controller 128 indicated below) from an original image in units of a block input from the splitter 102 and split by the splitter 102. In other words, the subtracter 104 calculates a prediction residual (also referred to as an error) of a current block. The subtracter 104 then outputs the calculated prediction residual to the transformer 106.
[0250] The original picture can be a picture that has been input to the encoder 100 as a signal (e.g., a luma signal and two chroma signals) representing each picture included in a video. The signal representing the picture can also be referred to as a sample.
[0251] (transformer)
[0252] The transformer 106 transforms the prediction residual in the spatial domain into a transform coefficient in the frequency domain, and outputs the transform coefficient to the quantizer 108. More specifically, the transformer 106 applies, for example, a discrete cosine transform (DCT) or a discrete sine transform (DST) defined to the prediction residual in the spatial domain. The defined DCT or DST can be predefined.
[0253] It should be noted that the transformer 106 can adaptively select a transform type from among a plurality of transform types, and transform the prediction residual into a transform coefficient by using a transform basis function corresponding to the selected transform type. Such a transform is also referred to as an explicit multi-kernel transform (EMT) or an adaptive multi-kernel transform (AMT). The transform basis function can also be referred to as a kernel.
[0254] The transform types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transform types can be denoted as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 is a graph indicating an example transform basis function of an example transform type. In Figure 14 In, N denotes the number of input pixels. For example, the selection of the transform type from among a plurality of transform types can depend on a prediction type (one of intra prediction and inter prediction), and can depend on an intra prediction mode.
[0255] Information indicating whether to apply such an EMT or AMT (e.g., referred to as an EMT flag or an AMT flag) and information indicating the selected transform type are generally signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level, but can also be performed at other levels (e.g., at the sequence level, the picture level, the slice level, the tile level, or the CTU level).
[0256] Furthermore, the transformer 106 can perform a retransformation on the transform coefficients, which are the result of the transform. This retransformation is also referred to as adaptive secondary transform (AST) or non-separable secondary transform (NSST). For example, the transformer 106 performs the retransformation in units of sub-blocks (e.g., 4x4 pixel sub-blocks) included in a transform coefficient block corresponding to an intra prediction residual. Information indicating whether or not to apply the NSST and information related to a transform matrix used for the NSST are typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level, but can also be performed at another level (e.g., at the sequence level, picture level, slice level, tile level, or CTU level).
[0257] The transformer 106 can employ separable transforms and non-separable transforms. A separable transform is a method in which a transform is performed multiple times by performing a transform for each of a plurality of directions separately according to the dimension of the input. A non-separable transform is a method of performing a collective transform in which two or more dimensions in a multi-dimensional input are collectively regarded as one dimension.
[0258] In one example of a non-separable transform, when the input is a 4x4 pixel block, the 4x4 pixel block is regarded as an array including 16 elements, and the transform applies a 16x16 transform matrix to the array.
[0259] In another example of a non-separable transform, an input block of 4x4 pixels is regarded as a single array including 16 elements, and then a transform that performs a given rotation multiple times on the array (hypercube given transform) can be performed.
[0260] In the transform in the transformer 106, the transform type to be transformed into a transform basis function in the frequency domain according to the region in the CU can be switched. Examples include a spatially varying transform (SVT).
[0261] Figure 15 is a conceptual diagram for illustrating one example of an SVT.
[0262] In the SVT, as shown in Figure 15 , a CU is split horizontally or vertically into two equal regions, and only one of the regions is transformed into the frequency domain. A transform basis type can be set for each region. For example, DST7 and DST8 are used. For example, in two regions obtained by vertically splitting one CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Or, in the two regions, DST7 is used for the region at position 1. Likewise, in two regions obtained by horizontally splitting one CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Or, in the two regions, DST7 is used for the region at position 1. Although in the above examples, DST7 and DCT8 are used, other transform bases can also be used. Figure 15 In the illustrated example, only one of the two regions in the CU is transformed and the other is not transformed, but each of the two regions can be transformed. In addition, the splitting method can include not only splitting into two regions, but also splitting into four regions. Furthermore, the splitting method can be more flexible. For example, information indicating the splitting method can be encoded and can be signaled in the same manner as the CU split. It should be noted that the SVT can also be referred to as sub-block transform (SBT).
[0263] The AMT and the EMT described above can be referred to as MTS (multiple transform selection). When the MTS is applied, a transform type of DST7, DCT8, etc. can be selected, and information indicating the selected transform type can be encoded as index information for each CU. There is another process referred to as IMTS (implicit MTS) as a process for selecting a transform type for an orthogonal transform performed without encoding the index information. When the IMTS is applied, for example, when a CU has a rectangular shape, an orthogonal transform of the rectangular shape can be performed by using DST7 for a short side and DST2 for a long side. In addition, for example, when a CU has a square shape, an orthogonal transform of the rectangular shape can be performed by using DCT2 when the MTS is active in a sequence and using DST7 when the MTS is inactive in the sequence. DCT2 and DST7 are only examples. Other transform types can be used, and the combination of the transform types used can also be changed to a different transform type combination. The IMTS can be used only for an intra prediction block, or can be used for both an intra prediction block and an inter prediction block.
[0264] The three processes of MTS, SBT, and IMTS have been described above as selection processes for selectively switching the transform type used for the orthogonal transform. However, all three selection processes can be employed, or only some of the selection processes can be selectively employed. For example, whether one or more of these selection processes is employed can be identified based on flag information in a header such as an SPS, or the like. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching the transform type can be a different selection process from the above three selection processes, or each of the three selection processes can be replaced by another process. In general, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing the orthogonal transform of the entire CU and determining the transform type based on a determined rule without encoding information indicating the transform type. Function [3] is a function for performing the orthogonal transform of a partial region of the CU and encoding information indicating the transform type used in the transform. Function [4] is a function for performing the orthogonal transform of a partial region of the CU and determining the transform type based on a determined rule without encoding information indicating the transform type used in the transform. The determined rule can be predetermined.
[0265] It should be noted that whether MTS, IMTS, and / or SBT is applied can be determined for each processing unit. For example, whether MTS, IMTS, and / or SBT is applied can be determined for each sequence, picture, block, slice, CTU, or CU.
[0266] It should be noted that the tool for selectively switching the transform type in the present disclosure can be described as a method, selection process, or process for selecting a basis used in a transform process for selectively selecting a basis. In addition, the tool for selectively switching the transform type can be described as a mode for adaptively selecting a transform type.
[0267] Figure 16 is a flowchart showing one example of a process performed by the converter 106, and will be described with reference to Figure 7 for convenience.
[0268] For example, the transformer 106 determines whether to perform the orthogonal transform (step St l). Here, when it is determined to perform the orthogonal transform (Yes in step St l), the transformer 106 selects a transform type for the orthogonal transform from among a plurality of transform types (step St_2). Next, the transformer 106 performs the orthogonal transform by applying the selected transform type to the prediction residual of the current block (step St_3). The transformer 106 then outputs information indicating the selected transform type to the entropy encoder 110 to allow the entropy encoder 110 to encode the information (step St_4). On the other hand, when it is determined not to perform the orthogonal transform (No in step St l), the transformer 106 outputs information indicating that the orthogonal transform is not performed to allow the entropy encoder 110 to encode the information (step St_5). It should be noted that whether to perform the orthogonal transform in step St l can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, and the like. Alternatively, the orthogonal transform can be performed using a defined transform type without encoding information indicating the transform type for the orthogonal transform. The defined transform type can be predefined.
[0269] Figure 17 is a flowchart showing one example of a process performed by the transformer 106, and will be described with reference to Figure 7 for convenience. It should be noted that Figure 17 the example shown is in the case where Figure 16 the example shown in FIG. 10. The orthogonal transform in the case where the transform type for the orthogonal transform is selectively switched in the example shown in FIG. 10 will be described.
[0270] As one example, the first transform type group can include DCT2, DST7, and DCT8. As another example, the second transform type group can include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group can partially overlap, or can be completely different from each other.
[0271] The transformer 106 determines whether the transform size is smaller than or equal to a determination value (step Su l). Here, when it is determined that the transform size is smaller than or equal to the determination value (Yes in step Su l), the transformer 106 orthogonally transforms the prediction residual of the current block using a transform type included in the first transform type group (step Su_2). Next, the transformer 106 outputs information indicating the transform type to be used among at least one transform type included in the first transform type group to the entropy encoder 110 to allow the entropy encoder 110 to encode the information (step Su_3). On the other hand, when it is determined that the transform size is not smaller than or equal to the predetermined value (No in step Su l), the transformer 106 orthogonally transforms the prediction residual of the current block using the second transform type group (step Su_4). The determination value can be a threshold value, or a predetermined value.
[0272] In step Su_3, the information indicating the transform type used in the orthogonal transform can be information indicating a combination of a transform type applied vertically in the current block and a transform type applied horizontally in the current block. A first type group can include only one transform type, and information indicating a transform type used for the orthogonal transform can not be encoded. A second transform type group can include a plurality of transform types, and information indicating a transform type used for the orthogonal transform among one or more transform types included in the second transform type group can be encoded.
[0273] Alternatively, the transform type can be indicated based on the transform size without encoding information indicating the transform type. It should be noted that such a determination is not limited to a determination of whether the transform size is smaller than or equal to a determined value, and other processes can be used to determine the transform type used for the orthogonal transform based on the transform size.
[0274] (Quantizer)
[0275] The quantizer 108 quantizes the transform coefficients output from the transformer 106. More specifically, the quantizer 108 scans the transform coefficients of the current block in a determined scan order, and quantizes the scanned transform coefficients based on a quantization parameter (QP) corresponding to the transform coefficients. The quantizer 108 then outputs the quantized transform coefficients (hereinafter also referred to as quantized coefficients) of the current block to the entropy encoder 110 and the inverse quantizer 112. The determined scan order can be predetermined.
[0276] The determined scan order is an order for quantizing / inverse quantizing the transform coefficients. For example, the determined scan order can be defined as an ascending order of frequency (from low frequency to high frequency) or a descending order of frequency (from high frequency to low frequency).
[0277] The quantization parameter (QP) is a parameter defining a quantization step (quantization width). For example, when the value of the quantization parameter increases, the quantization step also increases. In other words, when the value of the quantization parameter increases, the error (quantization error) of the quantized coefficients increases.
[0278] In addition, a quantization matrix can be used for quantization. For example, several quantization matrices can be used for frequency transform sizes such as 4x4, 8x8, prediction modes such as intra prediction and inter prediction, and pixel components such as luminance and chrominance pixel components, respectively. It should be noted that quantization means digitizing values sampled at a determined interval corresponding to a determined level. In this technical field, quantization can refer to using other expressions such as rounding and scaling, and rounding and scaling can be employed. The determined interval and the determined level can be predetermined.
[0279] Methods using a quantization matrix can include a method of using a quantization matrix directly set at the encoder 100 side, and a method of using a quantization matrix set as a default (default matrix). At the encoder 100 side, a quantization matrix suitable for the characteristics of an image can be set by directly setting the quantization matrix. However, this case can have a disadvantage of increasing the amount of coding for coding the quantization matrix. It should be noted that a quantization matrix for quantizing a current block can be generated based on a default quantization matrix or a coded quantization matrix, rather than directly using the default quantization matrix or the coded quantization matrix.
[0280] There is a method of quantizing high frequency coefficients and low frequency coefficients without using a quantization matrix. It should be noted that this method can be regarded as a method equivalent to using a quantization matrix whose coefficients have the same value (flat matrix).
[0281] A quantization matrix can be coded, for example, at a sequence level, a picture level, a slice level, a block level, or a CTU level. A quantization matrix can be specified using, for example, a sequence parameter set (SPS) or a picture parameter set (PPS). The SPS includes parameters for a sequence, and the PPS includes parameters for a picture. Each of the SPS and the PPS can be simply referred to as a parameter set.
[0282] When a quantization matrix is used, the quantizer 108 scales a quantization width, which can be calculated based on a quantization parameter or the like, for each transform coefficient using the values of the quantization matrix. A quantization process performed without using a quantization matrix can be a process of quantizing transform coefficients based on a quantization width calculated from a quantization parameter or the like. It should be noted that in a quantization process performed without using any quantization matrix, a quantization width can be multiplied by a determined value common to all transform coefficients in a block. The determined value can be predetermined.
[0283] Figure 18 is a block diagram illustrating one example of a functional configuration of a quantizer according to an embodiment. For example, the quantizer 108 includes a differential quantization parameter generator 108a, a predicted quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter memory 108d, and a quantization executor 108e.
[0284] Figure 19 is a flowchart illustrating one example of a quantization process performed by the quantizer 108, and will be described with reference to Figure 7 and 18 for convenience.
[0285] As one example, the quantizer 108 can generate a quantization matrix for quantizing a current block based on a default quantization matrix or a coded quantization matrix, rather than directly using the default quantization matrix or the coded quantization matrix. Figure 19The flowchart shown performs quantization for each CU. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_l). Here, when it is determined to perform quantization (Yes in step Sv_l), the quantization parameter generator 108c generates a quantization parameter of the current block (step Sv_2), and stores the quantization parameter to the quantization parameter memory 108d (step Sv_3).
[0286] Next, the quantization performer 108e quantizes the transform coefficients of the current block using the quantization parameter generated in step Sv_2 (step Sv_4). The prediction quantization parameter generator 108b then obtains the quantization parameter of the processing unit different from the current block from the quantization parameter memory 108d (step Sv_5). The prediction quantization parameter generator 108b generates a prediction quantization parameter of the current block based on the obtained quantization parameter (step Sv_6). The difference quantization parameter generator 108a calculates a difference between the quantization parameter of the current block generated by the quantization parameter generator 108c and the prediction quantization parameter of the current block generated by the prediction quantization parameter generator 108b (step Sv_7). The difference quantization parameter can be generated by calculating the difference. The difference quantization parameter generator 108a outputs the difference quantization parameter to the entropy encoder 110 to allow the entropy encoder 110 to encode the difference quantization parameter (step Sv_8).
[0287] It should be noted that different quantization parameters can be encoded, for example, at a sequence level, a picture level, a slice level, a block level, or a CTU level. Further, an initial value of the quantization parameter can be encoded at a sequence level, a picture level, a slice level, a block level, or a CTU level. In the initialization stage, the initial value of the quantization parameter and the difference quantization parameter can be used to generate the quantization parameter.
[0288] It should be noted that the quantizer 108 can include a plurality of quantizers, and a dependent quantization in which a transform coefficient is quantized using a quantization method selected from a plurality of quantization methods can be applied.
[0289] (Entropy encoder)
[0290] Figure 20 is a block diagram showing one example of a functional configuration of the entropy encoder 110 according to the embodiment, and will be described with reference to Figure 7is described. The entropy encoder 110 generates a stream by entropy encoding the quantized coefficients input from the quantizer 108 and the prediction parameters input from the prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy encoding. More specifically, the entropy encoder 110 as illustrated includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization in which a multi-level signal such as a quantized coefficient and a prediction parameter is transformed into a binary signal. Examples of the binarization method include truncated Rice binarization, exponential Golomb coding, and fixed length binarization. The context controller 110b derives a context value according to a feature or a surrounding state of a syntax element, that is, an occurrence probability of a binary signal. Examples of the method for deriving a context value include bypass, reference to a syntax element, reference to an upper neighbor and a left neighbor block, reference to level information, and the like. The binary arithmetic encoder 110c performs arithmetic encoding of a binary signal using the derived context.
[0291] Figure 21 is a conceptual diagram for illustrating an example flow of a CABAC process in the entropy decoder 110. First, initialization is performed in the CABAC in the entropy encoder 110. In the initialization, initialization in the binary arithmetic encoder 110c and setting of an initial context value are performed. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic encoding of a plurality of quantized coefficients in a CTU. The context controller 110b can update a context value each time the arithmetic encoding is performed. The context controller 110b can then save the context value as a post-process. For example, the saved context value can be used to initialize a context value for a next CTU.
[0292] (inverse quantizer)
[0293] The inverse quantizer 112 inverse quantizes the quantized coefficients input from the quantizer 108. More specifically, the inverse quantizer 112 inverse quantizes the quantized coefficients of the current block in a determined scan order. The inverse quantizer 112 then outputs the inverse quantized transform coefficients of the current block to the inverse transformer 114. The determined scan order can be predetermined.
[0294] (inverse transformer)
[0295] The inverse transformer 114 recovers the prediction residual by inverse transforming the transform coefficients input from the inverse quantizer 112. More specifically, the inverse transformer 114 recovers the prediction residual of the current block by performing inverse transformation corresponding to the transformation applied to the transform coefficients by the transformer 106. The inverse transformer 114 then outputs the recovered prediction residual to the adder 116.
[0296] It should be noted that the recovered prediction residual does not match the prediction residual calculated by the subtractor 104 because information is generally lost in quantization. In other words, the recovered prediction residual generally includes quantization error.
[0297] (adder)
[0298] The adder 116 reconstructs the current block by adding the prediction residual input from the inverse transformer 114 and the prediction image input from the prediction controller 128. Subsequently, a reconstructed image is generated. The adder 116 then outputs the reconstructed image to the block memory 118 and the loop filter 120. The reconstructed block can also be referred to as a locally decoded block.
[0299] (block memory)
[0300] The block memory 118 is a storage for storing blocks in a current picture, for example, for intra prediction. More specifically, the block memory 118 stores the reconstructed image output from the adder 116.
[0301] (frame memory)
[0302] The frame memory 122 is a memory, for example, for storing reference pictures used in inter prediction, and is also referred to as a frame buffer. More specifically, the frame memory 122 stores the reconstructed image filtered by the loop filter 120.
[0303] (loop filter)
[0304] The loop filter 120 applies a loop filter to the reconstructed image output from the adder 116 and outputs the filtered reconstructed image to the frame memory 122. The loop filter is a filter used in an encoding loop (loop filter). Examples of the loop filter include, for example, an adaptive loop filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, and the like.
[0305] Figure 22 is a block diagram showing one example of a functional configuration of the loop filter 120 according to the embodiment. For example, as Figure 22As shown, the loop filter 120 includes a deblocking filter enforcer 120a, an SAO enforcer 120b, and an ALF enforcer 120c. The deblocking filter enforcer 120a performs a deblocking filter process on the reconstructed picture. The SAO enforcer 120b performs an SAO process on the reconstructed picture after the deblocking filter process. The ALF enforcer 120c performs an ALF process on the reconstructed picture after the SAO process. The ALF and the deblocking filter will be described in detail later. The SAO process is a process to improve the picture quality by reducing ringing (a phenomenon in which pixel values around an edge are distorted like a wave) and correcting bias of pixel values. Examples of the SAO process include an edge offset process and a band offset process. It should be noted that, in some embodiments, the loop filter 120 can not include all of the above-described processes, can not perform all of the above-described processes, etc. Figure 22 All of the constituent elements disclosed in the above-described embodiments can be included, some of the constituent elements can be included, and additional elements can be included. Further, the loop filter 120 can be configured to perform the above-described processes in a processing order different from the processing order disclosed in the above-described embodiments, can not perform all of the above-described processes, etc. Figure 22 All of the constituent elements disclosed in the above-described embodiments can be included, some of the constituent elements can be included, and additional elements can be included. Further, the loop filter 120 can be configured to perform the above-described processes in a processing order different from the processing order disclosed in the above-described embodiments, can not perform all of the above-described processes, etc.
[0306] (Loop filter > Adaptive loop filter)
[0307] In the ALF, a least square error filter for removing compression artifacts is applied. For example, one filter selected from a plurality of filters based on a direction and activity of local gradient is applied for each 2x2 pixel sub-block in the current block.
[0308] More specifically, first, each sub-block (e.g., each 2x2 pixel sub-block) is classified into one of a plurality of classes (e.g., fifteen or twenty-five classes). The classification of the sub-block can be based on, for example, gradient directionality and activity. In one example, a class index C (e.g., C=5D+A) is calculated or determined from a gradient directionality D (e.g., 0 to 2 or 0 to 4) and a gradient activity A (e.g., 0 to 4). Then, based on the class index C, each sub-block is classified into one of the plurality of classes.
[0309] The gradient directionality D is calculated, for example, by comparing gradients of a plurality of directions (e.g., horizontal, vertical, and two diagonal directions). Further, the gradient activity A is calculated, for example, by adding the gradients of the plurality of directions and quantizing the added result.
[0310] A filter to be used for each sub-block can be determined from a plurality of filters based on such a classification result.
[0311] The filter shape used in the ALF is, for example, a circularly symmetric filter shape. Figures 23A to 23C is a conceptual diagram for showing an example of a filter shape used in the ALF. Figure 23A shows a 5x5 diamond filter, Figure 23BA 7x7 diamond filter is shown, Figure 23C A 9x9 diamond filter is shown. The information indicating the filter shape is usually signaled at the picture level. It should be noted that the signaling of the information indicating the filter shape does not necessarily need to be performed at the picture level, but can also be performed at other levels (e.g., at the sequence level, slice level, tile level, or CTU level).
[0312] For example, the turning on or off of ALF can be determined at the picture level or CU level. For example, the decision whether to apply ALF for luma can be made at the CU level, and the decision whether to apply ALF for chroma can be made at the picture level. The information indicating the turning on or off of ALF is usually signaled at the picture level or CU level. It should be noted that the signaling of the information indicating the turning on or off of ALF does not necessarily need to be performed at the picture level or CU level, but can also be performed at other levels (e.g., at the sequence level, slice level, tile level, or CTU level).
[0313] Further, as mentioned above, one filter is selected from a plurality of filters, and the ALF processing of a subblock is performed. The coefficient set for the coefficients of each filter (e.g., up to the fifteenth or twenty-fifth filter) of the plurality of filters is usually signaled at the picture level. It should be noted that the signaling of the coefficient set does not necessarily need to be performed at the picture level, but can also be performed at other levels (e.g., at the sequence level, slice level, tile level, CTU level, CU level, or subblock level).
[0314] (circular filter > cross-component adaptive loop filter)
[0315] Figure 23D is a conceptual diagram for illustrating an example procedure of cross-component ALF (CC-ALF). Figure 23E is a conceptual diagram for illustrating an example filter shape used in CC-ALF (e.g. Figure 23D of CC-ALF. Figure 23D and Figure 23E The example CC-ALF of operates by applying a linear diamond filter to the luma channel of each chroma component. For example, the filter coefficients can be sent in the APS, scaled by a factor of 2A10, and rounded for fixed-point representation. For example, in Figure 23D In, Y samples (first component) are used for CCALF of Cb and CCALF of Cr (a component different from the first component).
[0316] The application of the filter can be controlled on a variable block size and signaled by a context coded flag received for each sample block. The block size can be received at slice level for each chroma component along with the CC-ALF enable flag. CC-ALF can support various block sizes, for example, 16x16 pixels, 32x32 pixels, 64x64 pixels, 128x128 pixels in chroma samples.
[0317] (Cycle filter) Joint chroma cross-component adaptive loop filter
[0318] One example of joint chroma-CCALF is illustrated in Figure 23F and 23G . Figure 23F is a conceptual diagram showing an example process of joint chroma CCALF. Figure 23G is a table showing example weight exponent candidates. As shown, one CCALF filter is used to generate one CCALF filter output as a chroma refinement signal for one color component, while applying a weighted version of the same chroma refinement signal to another color component. In this way, the complexity of the existing CCALF is reduced by about half. The weight value can be coded as a sign flag and a weight index. The weight index, denoted as weight_index, can be coded as 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight, which is of non-zero size. For example, the magnitude of JcCcWeight can be determined as follows:
[0319] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index » 2;
[0320] Otherwise, JcCcWeight is equal to 4 / (weight_index - 4).
[0321] The block-level on / off control of ALF filtering for Cb and Cr can be separate. This is the same as in CCALF and two separate sets of block-level on / off control flags can be coded. Unlike in CCALF, the Cb, Cr on / off control block sizes are the same here, so only one block size variable can be coded.
[0322] (Cycle filter) Deblocking filter
[0323] In the deblocking filter process, the cycle filter 120 performs a filtering process on block boundaries in the reconstructed image to reduce distortion that occurs at the block boundaries.
[0324] Figure 24 is a conceptual diagram showing the cycle filter 120 used as a deblocking filter (see Figure 7 and Figure 22a block diagram of one example of a specific configuration of the deblocking filter enforcer 120a of FIG. 1.
[0325] The deblocking filter enforcer 120a includes a boundary determiner 1201, a filter determiner 1203, a filter enforcer 1205, a process determiner 1208, a filter property determiner 1207, and switches 1202, 1204, and 1206.
[0326] The boundary determiner 1201 determines whether a pixel to be deblocking filtered (i.e., a current pixel) exists around a block boundary. The boundary determiner 1201 then outputs the determination result to the switch 1202 and the process determiner 1208.
[0327] In a case where the boundary determiner 1201 has determined that the current pixel exists around the block boundary, the switch 1202 outputs an unfiltered image to the switch 1204. In the opposite case where the boundary determiner 1201 has determined that the current pixel does not exist around the block boundary, the switch 1202 outputs the unfiltered image to the switch 1206. It should be noted that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.
[0328] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on pixel values of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.
[0329] In a case where the filter determiner 1203 has determined to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the filter enforcer 1205. In the opposite case where the filter determiner 1203 has determined not to perform deblocking filtering on the current pixel, the switch 1204 outputs the unfiltered image obtained through the switch 1202 to the switch 1206.
[0330] When the unfiltered image is obtained through the switches 1202 and 1204, the filter enforcer 1205 performs deblocking filtering on the current pixel with a filter property determined by the filter property determiner 1207. The filter enforcer 1205 then outputs the filtered pixel to the switch 1206.
[0331] The switch 1206 selectively outputs one of a pixel that has not been deblocking filtered and a pixel that has been deblocking filtered by the filter enforcer 1205 under control of the process determiner 1208.
[0332] The process determiner 1208 controls the switch 1206 based on the results of the determinations made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 determines that the current pixel exists in the vicinity of the block boundary and when the filter determiner 1203 determines that deblocking filtering of the current pixel is performed, the process determiner 1208 causes the switch 1206 to output the pixel that has been deblocking filtered. Further, the process determiner 1208 causes the switch 1206 to output the pixel that has not been deblocking filtered except for the above-described case. By repeating the output of the pixel in this way, the filtered image is output from the switch 1206. It should be noted that Figure 24 The configuration shown in FIG. 12B is one example of the configuration in the deblocking filter performer 120a. The deblocking filter performer 120a can have various configurations.
[0333] Figure 25 is a conceptual diagram for illustrating an example of a deblocking filter having a symmetric filtering characteristic with respect to a block boundary.
[0334] In the deblocking filtering process, one of two deblocking filters having different characteristics (i.e., a strong filter and a weak filter) can be selected using a pixel value and a quantization parameter. In the case of the strong filter, when the pixels p0to p2and the pixels q0to q2exist across the block boundary, as shown in Figure 25 by performing a calculation according to, for example, the following expression, the pixel values of the respective pixels q0to q2are changed to the pixel values q'0to q'2.
[0335] q'0= (p1+ 2 x p0+ 2 x q0+ 2 x q1+ q2+ 4) / 8
[0336] q'1= (p0+ q0+ q1+ q2+ 2) / 4
[0337] q'2= (p0+ q0+ q1+ 3 x q2+ 2 x q3+ 4) / 8
[0338] It should be noted that in the above expressions, p0to p2and q0to q2are the pixel values of the pixels p0to p2and the pixels q0to q2, respectively. In addition, q3is the pixel value of the neighboring pixel q3that is located on the opposite side of the block boundary from the pixel q2. Further, on the right side of each expression, the coefficients multiplied by the respective pixel values of the pixels to be used for deblocking filtering are filter coefficients.
[0339] Further, in the deblocking filtering, clipping can be performed so that the change in the calculated pixel value does not exceed a threshold value. For example, in the clipping process, the pixel value calculated according to the above expression can be clipped to a value obtained according to "calculated pixel value ± 2 x threshold value" using a threshold value determined based on the quantization parameter. In this way, excessive smoothing can be prevented.
[0340] Figure 26is a conceptual diagram for illustrating a block boundary against which a deblocking filtering process is performed. Figure 27 is a conceptual diagram for illustrating an example of a boundary strength (Bs) value.
[0341] A block boundary against which a deblocking filtering process is performed is, for example, a boundary between CUs, Pus, or TUs having 8x8 pixels, as shown in Figure 26 The deblocking filtering process can be performed in units of four rows or four columns, for example. First, for a block P and a block Q as shown in Figure 26 a boundary strength (Bs) value is determined as shown in Figure 27
[0342] According to the Bs value in Figure 27 , it can be determined whether to perform a deblocking filtering process against a block boundary belonging to the same picture using different strengths. When the Bs value is 2, a deblocking filtering process for a chroma signal is performed. When the Bs value is 1 or more and a determination condition is satisfied, a deblocking filtering process for a luma signal is performed. The determined condition can be predetermined. It should be noted that the conditions for determining the Bs value are not limited to those shown in Figure 27 and the Bs value can be determined based on another parameter.
[0343] (predictor (intra predictor, inter predictor, prediction controller))
[0344] Figure 28 is a flowchart showing one example of a process performed by a predictor of the encoder 100. It should be noted that the predictor includes all or a part of the following constituent elements: the intra predictor 124; the predictor 126; and the prediction controller 128. The prediction performer includes, for example, the intra predictor 124 and the inter predictor 126.
[0345] The predictor generates a prediction image of the current block (step Sb_1). The prediction image can also be referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra prediction image (intra prediction signal) or an inter prediction image (inter prediction signal). The predictor generates the prediction image of the current block using a reconstructed image that has been obtained through generation of a prediction image, generation of a prediction residual, generation of quantized coefficients, restoration of the prediction residual, and addition of the prediction image, of another block.
[0346] The reconstructed image can be, for example, an image in a reference picture, or an image of an encoded block in a current picture (which is a picture including the current block), that is, the above-described other block. The encoded block in the current picture is, for example, a neighboring block of the current block.
[0347] Figure 29 is a flowchart showing another example of a process performed by a predictor of the encoder 100.
[0348] The predictor generates a prediction image using a first method (step Sc la), generates a prediction image using a second method (step Sc lb), and generates a prediction image using a third method (step Sc lc). The first, second, and third methods can be different methods for generating a prediction image from each other. Each of the first to third methods can be an inter prediction method, an intra prediction method, or another prediction method. The above-described reconstructed image can be used in these prediction methods.
[0349] Next, the prediction processor evaluates the prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 2). For example, the predictor calculates the cost C with respect to the prediction images generated in steps Sc la, Sc lb, and Sc lc, and evaluates the prediction images by comparing the costs C of the prediction images. It should be noted that the cost C can be calculated, for example, according to an expression of an RD optimization model (e.g., C = D + λ x R). In this expression, D denotes a compression artifact of a prediction image, and is expressed as, for example, a sum of absolute differences between pixel values of a current block and pixel values of a prediction image. Further, R denotes a bit rate of a stream. Further, λ denotes a multiplier according to, for example, a Lagrange multiplier method.
[0350] The predictor then selects one of the prediction images generated in steps Sc la, Sc lb, and Sc lc (step Sc 3). In other words, the predictor selects a method or mode for obtaining a final prediction image. For example, the predictor selects a prediction image having a minimum cost C based on the costs C calculated for the prediction images. Alternatively, the evaluation in step Sc 2 and the selection of the prediction image in step Sc 3 can be made based on parameters used in the encoding process. The encoder 100 can convert information for identifying the selected prediction image, method, or mode into a stream. The information can be, for example, a flag or the like. In this way, the decoder 200 is able to generate a prediction image according to the method or mode selected by the encoder 100 based on the information. It should be noted that, in the example shown, the predictor selects any of the prediction images after generating the prediction images by using the respective methods. However, the predictor can select a method or mode based on parameters used in the above-described encoding process before generating the prediction images, and can generate the prediction images according to the selected method or mode. Figure 29
[0351] For example, the first and second methods can be intra prediction and inter prediction, respectively, and the predictor can select a final prediction image of a current block from the prediction images generated according to the prediction methods.
[0352] Figure 30 is a flowchart showing another example of a process performed by the prediction of the encoder 100.
[0353] First, the predictor generates a prediction image using intra prediction (step Sd la), and generates a prediction image using inter prediction (step Sd lb). Note that the prediction image generated by the intra prediction is also referred to as an intra prediction image, and the prediction image generated by the inter prediction is also referred to as an inter prediction image.
[0354] Next, the predictor evaluates each of the intra prediction image and the inter prediction image (step Sd 2). The above-described cost C can be used in the evaluation. The predictor can then select, as the final prediction image of the current block, the prediction image for which the minimum cost C has been calculated, among the intra prediction image and the inter prediction image (step Sd 3). In other words, the prediction method or mode used to generate the prediction image of the current block is selected.
[0355] The prediction processor then selects, as the final prediction image of the current block, the prediction image for which the minimum cost C has been calculated, among the intra prediction image and the inter prediction image (step Sd 3). In other words, the prediction method or mode used to generate the prediction image of the current block is selected.
[0356] (Intra predictor)
[0357] The intra predictor 124 generates a prediction signal (i.e., an intra prediction image) by performing intra prediction of the current block by referring to one or more blocks in the current picture and storing it in the block memory 118. More specifically, the intra predictor 124 generates an intra prediction image by performing intra prediction by referring to pixel values (e.g., luma and / or chroma values) of one or more blocks adjacent to the current block, and then outputs the intra prediction image to the prediction controller 128.
[0358] For example, the intra predictor 124 performs intra prediction by using one of a plurality of defined intra prediction modes. The intra prediction modes typically include one or more non-directional prediction modes and a plurality of directional prediction modes. The defined modes can be predefined.
[0359] The one or more non-directional prediction modes include, for example, a planar prediction mode and a DC prediction mode defined in the H.265 / high efficiency video coding (HEVC) standard.
[0360] The plurality of directional prediction modes include, for example, thirty-three directional prediction modes defined in the H.265 / HEVC standard. Note that the plurality of directional prediction modes can include, in addition to the thirty-three directional prediction modes, thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31is a conceptual diagram showing a total of 67 intra prediction modes (two non-directional prediction modes and 65 directional prediction modes) that can be used in intra prediction. The solid arrows represent thirty-three directions defined in the H.265 / HEVC standard, while the dashed arrows represent the additional thirty-two directions (two non-directional prediction modes are not shown in Figure 31
[0361] In various processing examples, a luma block can be referenced in intra prediction of a chroma block. In other words, a chroma component of a current block can be predicted based on a luma component of the current block. Such intra prediction is also referred to as cross-component linear model (CCLM) prediction. An intra prediction mode for a chroma block in which such a luma block is referenced (also referred to as, for example, CCLM mode) can be added as one of the intra prediction modes of the chroma block.
[0362] The intra predictor 124 can correct the intra prediction pixel values based on horizontal / vertical reference pixel gradients. The intra prediction accompanied by such correction is also referred to as position-dependent intra prediction combination (PDPC). Information indicating whether to apply PDPC (for example, referred to as a PDPC flag) is typically signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level, but can also be performed at other levels (for example, at the sequence level, picture level, slice level, tile level, or CTU level).
[0363] Figure 32 is a flowchart showing one example of a process performed by the intra predictor 124.
[0364] The intra predictor 124 selects one of the intra prediction modes from the plurality of intra prediction modes (step Sw_1). The intra predictor 124 then generates a prediction image according to the selected intra prediction mode (step Sw_2). Next, the intra predictor 124 determines most probable modes (MPMs) (step Sw_3). The MPMs include, for example, six intra prediction modes. For example, two of the six intra prediction modes can be a planar mode and a DC prediction mode, and the other four modes can be directional prediction modes. The intra predictor 124 determines whether the intra prediction mode selected in step Sw_1 is included in the MPMs (step Sw_4).
[0365] Here, when it is determined that the intra prediction mode selected in step Sw_1 is included in the MPMs (Yes in step Sw_4), the intra predictor 124 sets an MPM flag to 1 (step Sw_5), and generates information indicating the intra prediction mode selected among the MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra prediction mode can be encoded as a prediction parameter by the entropy encoder 110.
[0366] When it is determined that the selected intra prediction mode is not included in the MPM (NO in step Sw_4), the intra predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra predictor 124 does not set any MPM flag. The intra predictor 124 then generates information indicating the intra prediction mode selected from the at least one intra prediction mode that is not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra prediction mode can be encoded as the prediction parameter by the entropy encoder 110. The information indicating the intra prediction mode indicates, for example, any one of 0 to 60.
[0367] (Intra predictor)
[0368] The inter predictor 126 generates a predicted image (inter predicted image) by performing inter prediction (also referred to as inter prediction) of a current block by referring to one or more blocks in a reference picture (which is different from the current picture) and storing it in the frame memory 122. Inter prediction is performed in units of a current block or a current sub-block (e.g., a 4x4 block) in the current block. A sub-block is included in a block and is a smaller unit than the block. The size of the sub-block can be in the form of a slice, a block, a picture, and the like.
[0369] For example, the inter predictor 126 performs motion estimation in the reference picture of the current block or the current sub-block and finds a reference block or a reference sub-block that best matches the current block or the current sub-block. The inter predictor 126 then obtains motion information (e.g., a motion vector) that compensates for motion or changes from the reference block or the reference sub-block to the current block or the sub-block. The inter predictor 126 generates an inter predicted image of the current block or the sub-block by performing motion compensation (or motion prediction) based on the motion information. The inter predictor 126 outputs the generated inter predicted image to the prediction controller 128.
[0370] The motion information used in the motion compensation can be signaled as an inter prediction signal in various forms. For example, a motion vector can be signaled. As another example, a difference between a motion vector and a motion vector predictor can be signaled.
[0371] (Reference picture list)
[0372] Figure 33 is a conceptual diagram for illustrating an example of a reference picture. Figure 34 is a conceptual diagram for illustrating an example of a reference picture list. The reference picture list is a list indicating at least one reference picture stored in the frame memory 122. It should be noted that, in the example of the reference picture list illustrated in FIG. 10, the reference picture list is a list indicating two reference pictures. However, the reference picture list can be a list indicating one or more reference pictures. Figure 33In each of the diagrams, each rectangle represents one picture, each arrow represents one picture reference relationship, the horizontal axis represents time, I, P, and B in the rectangle respectively represent an intra prediction picture, a single prediction picture, and a bi-directional prediction picture, and the number in the rectangle represents the decoding order. As shown in FIG. 10A, the decoding order of the pictures is the order of I0, P1, B2, B3, and B4, and the display order of the pictures is the order of I0, B3, B2, B4, and P1. As shown in FIG. 10B, the display order of the pictures is the order of I0, B3, B2, B4, and P1, and the decoding order of the pictures is the order of I0, P1, B2, B3, and B4. Figure 33 Figure 34 As shown in FIG. 11, a reference picture list is a list indicating reference picture candidates. For example, one picture (or slice) can include at least one reference picture list. For example, one reference picture list is used when the current picture is a single prediction picture, and two reference picture lists are used when the current picture is a bi-directional prediction picture. In the example of FIG. 11, the reference picture list is a L0 list and a L1 list. The reference picture candidates of the current picture currPic are I0, P1, and B2, and the reference picture lists (which are the L0 list and the L1 list) indicate these pictures. Figure 33 34 In the example of FIG. 11, the picture B3, which is the current picture currPic, has two reference picture lists, i.e., a L0 list and a L1 list. When the current picture currPic is the picture B3, the reference picture candidates of the current picture currPic are I0, P1, and B2, and the reference picture lists (which are the L0 list and the L1 list) indicate these pictures. The inter predictor 126 or the prediction controller 128 specifies which picture in each reference picture list is actually to be referred to in the form of a reference picture index refidxLx. In the example of FIG. 11, the reference pictures P1 and B2 are specified by the reference picture indexes refIdxL0 and refIdxL1. Figure 34
[0373] Such a reference picture list can be generated for each unit such as a sequence, a picture, a slice, a block, a CTU, or a CU. In addition, among the reference pictures indicated in the reference picture list, a reference picture index indicating a reference picture to be referred to in inter prediction can be signaled at a sequence level, a picture level, a slice level, a block level, a CTU level, or a CU level. Furthermore, a common reference picture list can be used in a plurality of inter prediction modes.
[0374] (Basic flow of inter prediction)
[0375] Figure 35 FIG. 12 is a flowchart showing an example basic processing flow of an inter prediction process.
[0376] First, the inter predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtracter 104 generates a difference between the current block and the prediction image as a prediction residual (step Se_4).
[0377] Here, in the generation of the prediction image, the inter prediction unit 126 generates the prediction image by determination of a motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Further, in the determination of the MV, the inter prediction unit 126 determines the MV by selection of a motion vector candidate (MV candidate) (step Se_1) and derivation of the MV (step Se_2). The selection of the MV candidate is performed by, for example, the inter prediction unit 126 generating a list of MV candidates and selecting at least one MV candidate from the list of MV candidates. Note that a past-derived MV can be added to the list of MV candidates. Alternatively, in the derivation of the MV, the inter prediction unit 126 can select at least one MV candidate from the at least one MV candidate and determine the selected at least one MV candidate as the MV of the current block. Alternatively, the inter prediction unit 126 can determine the MV of the current block by performing estimation in a reference picture region specified by each of the selected at least one MV candidate. Note that the estimation in the reference picture region can be referred to as motion estimation.
[0378] Further, although the steps Se_1 to Se_3 are performed by the inter prediction unit 126 in the above-described example, the processes of, for example, the step Se_1, the step Se_2, and the like can be performed by another constituent element included in the encoder 100.
[0379] Note that a list of MV candidates can be generated for each process in the inter prediction mode, or a common list of MV candidates can be used in a plurality of inter prediction modes. The processes in the steps Se_3 and Se_4 correspond to the processes in the steps Sa_3 and Sa_4, respectively. Figure 9 The process in the step Se_3 corresponds to the process in the step Sd_1b. Figure 30
[0380] (Motion vector derivation flow)
[0381] Figure 36 is a flowchart showing one example of a process of derivation of a motion vector.
[0382] The inter prediction unit 126 can derive the MV of the current block in a mode in which motion information (e.g., MV) is encoded. In this case, for example, the motion information can be encoded as a prediction parameter and can be signaled. In other words, the encoded motion information is included in a stream.
[0383] Alternatively, the inter prediction unit 126 can derive the MV in a mode in which the motion information is not encoded. In this case, the motion information is not included in the stream.
[0384] Here, the MV derivation mode can include a normal inter mode, a normal merge mode, a FRUC mode, an affine mode, and the like described later. The mode in which motion information is coded includes the normal inter mode, the normal merge mode, the affine mode (specifically, the affine inter mode and the affine merge mode), and the like. Note that the motion information can include not only the MV but also the motion vector predictor selection information described later. The mode in which no motion information is coded includes the FRUC mode and the like. The inter predictor 126 selects the mode for deriving the MV of the current block from among the plurality of modes and derives the MV of the current block using the selected mode.
[0385] Figure 37 is a flowchart illustrating another example of the derivation of a motion vector.
[0386] The inter predictor 126 can derive the MV of the current block in a mode in which the MV difference is coded. In this case, for example, the MV difference can be coded as a prediction parameter, and can be signaled. In other words, the coded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. Note that the MV predictor is the motion vector predictor.
[0387] Alternatively, the inter predictor 126 can derive the MV in a mode in which the MV difference is not coded. In this case, the coded MV difference is not included in the stream.
[0388] Here, as described above, the MV derivation mode includes the normal inter mode, the normal merge mode, the FRUC mode, the affine mode, and the like described later. The mode in which the MV difference is coded includes the normal inter mode, the affine mode (specifically, the affine inter mode), and the like. The mode in which the MV difference is not coded includes the FRUC mode, the normal merge mode, the affine mode (specifically, the affine merge mode), and the like. The inter predictor 126 selects the mode for deriving the MV of the current block from among the plurality of modes and derives the MV of the current block using the selected mode.
[0389] (Motion vector derivation mode)
[0390] Figure 38A and Figure 38B is a conceptual diagram for illustrating an example classification of the modes for MV derivation. For example, as Figure 38A indicated, the MV derivation mode is roughly classified into three modes according to whether or not the motion information is coded and whether or not the MV difference is coded. The three modes are an inter mode, a merge mode, and a frame rate up conversion (FRUC) mode. The inter mode is a mode in which motion estimation is performed and in which the motion information and the MV difference are coded. For example, as Figure 38BAs shown, the inter modes include an affine inter mode and a normal inter mode. The merge mode is a mode in which motion estimation is not performed, and in which an MV is selected from coded surrounding blocks, and the MV of the current block is derived using the MV. The merge mode is a mode in which motion information is substantially encoded without encoding an MV difference. For example, as shown in Figure 38B As shown, the merge mode includes a normal merge mode (also referred to as a normal merge mode or a regular merge mode), a motion vector difference merge (MMVD) mode, a combined inter / intra prediction (CIIP) mode, a triangle mode, an ATMVP mode, and an affine merge mode. Here, in the MMVD mode among the modes included in the merge mode, an MV difference is exceptionally encoded. It should be noted that the affine merge mode and the affine inter mode are modes included in the affine mode. The affine mode is a mode in which an MV of each of a plurality of sub-blocks included in a current block is derived as an MV of the current block assuming an affine transformation. The FRUC mode is a mode in which an MV of a current block is derived by performing estimation between coded regions, and in which neither motion information nor any MV difference is encoded. It should be noted that the respective modes will be described in more detail later.
[0391] It should be noted that Figure 38A and Figure 38B The classification of the modes shown in FIGS. 1, 2, 3, 4, 5, and 6 is an example, and the classification is not limited thereto. For example, when an MV difference is encoded in the CIIP mode, the CIIP mode is classified as an inter mode.
[0392] (MV derivation > normal inter mode)
[0393] The normal inter mode is an inter prediction mode in which an MV of a current block is derived from a reference picture region specified by an MV candidate based on a block similar to an image of the current block. In this normal inter mode, an MV difference is encoded.
[0394] Figure 39 is a flowchart showing an example of an inter prediction process in the normal inter mode.
[0395] First, the inter predictor 126 obtains a plurality of MV candidates of the current block based on information such as MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sg_1). In other words, the inter predictor 126 generates an MV candidate list.
[0396] Next, the inter predictor 126 extracts N (N is an integer of 2 or more) MV candidates as motion vector prediction candidates (also referred to as MV prediction candidates) from the plurality of MV candidates obtained in step Sg_1 according to the determined priority order (step Sg_2). Note that the priority order can be determined in advance for each of the N MV candidates.
[0397] Next, the inter predictor 126 selects one of the N prediction motion vector candidates as a motion vector predictor (also referred to as an MV predictor) for the current block (step Sg_3). At this time, the inter predictor 126 encodes, in the stream, prediction motion vector selection information for identifying the selected motion vector predictor. In other words, the inter predictor 126 outputs the MV predictor selection information as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0398] Next, the inter predictor 126 derives the MV of the current block by referring to the encoded reference picture (step Sg_4). At this time, the inter predictor 126 also encodes, in the stream, a difference between the derived MV and the motion vector predictor as an MV difference. In other words, the inter predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130. Note that the encoded reference picture is a picture including a plurality of blocks that are reconstructed after encoding.
[0399] Finally, the inter predictor 126 generates a prediction image of the current block by performing motion compensation of the current block using the derived MV and the encoded reference picture (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a slice, the inter prediction of the slice using the normal inter mode is completed. For example, when the processes in steps Sg_1 to Sg_5 are performed for all blocks in a picture, the inter prediction of the picture using the normal inter mode is completed. Note that it can not be the case that all blocks included in a slice undergo these processes in steps Sg_1 to Sg_5, and the inter prediction of the slice using the normal inter mode can be completed when some blocks undergo these processes. The same applies to the processes in steps Sg_1 to Sg_5. The inter prediction of the picture using the normal inter mode can be completed when some blocks in the picture undergo these processes.
[0400] Note that the prediction image is an inter prediction signal as described above. In addition, information indicating the inter prediction mode (normal inter mode in the above example) used to generate the prediction image is, for example, encoded as a prediction parameter in the encoded signal.
[0401] It should be noted that the MV candidate list can also be used as a list used in another mode. Furthermore, a process related to the MV candidate list can be applied to a process related to the list to be used in another mode. The process related to the MV candidate list includes, for example, extraction or selection of an MV candidate from the MV candidate list, reordering of the MV candidates, or deletion of the MV candidates.
[0402] (MV derivation > normal merge mode)
[0403] The normal merge mode is an inter prediction mode in which an MV candidate is selected from the MV candidate list as the MV of the current block to derive the MV. It should be noted that the normal merge mode is a kind of merge mode, which can be simply referred to as the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode is used in a wider range.
[0404] Figure 40 is a flowchart showing an example of inter prediction in the normal merge mode.
[0405] First, the inter predictor 126 obtains a plurality of MV candidates of the current block based on information such as MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sh_1). In other words, the inter predictor 126 generates the MV candidate list.
[0406] Next, the inter predictor 126 selects one MV candidate from the plurality of MV candidates obtained in step Sh_1 to derive the MV of the current block (step Sh_2). At this time, the inter predictor 126 encodes MV selection information for identifying the selected MV candidate in the stream. In other words, the inter predictor 126 outputs the MV selection information as the prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.
[0407] Finally, the inter predictor 126 generates a prediction image of the current block by performing motion compensation of the current block using the derived MV and the coded reference picture (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a slice, inter prediction of the slice using the normal merge mode is completed. Furthermore, when the processes in steps Sh_1 to Sh_3 are performed for all blocks in a picture, inter prediction of the picture using the normal merge mode is completed. It should be noted that not all blocks included in a slice can undergo these processes in steps Sh_1 to Sh_3, and inter prediction of the slice using the normal merge mode can be completed when part of the blocks undergoes these processes. This also applies to the processes in steps Sh_1 to Sh_3. Inter prediction of the picture using the normal merge mode can be completed when part of the blocks in the picture undergoes these processes.
[0408] In addition, information indicating the inter prediction mode used to generate the prediction image and included in the encoded signal (normal merge mode in the above example) is, for example, encoded as a prediction parameter in the stream.
[0409] Figure 41 is a conceptual diagram for showing one example of a motion vector derivation process by the normal merge mode.
[0410] First, the inter predictor 126 generates an MV candidate list in which MV candidates are registered. Examples of the MV candidates include: spatially neighboring MV candidates, which are MVs of a plurality of coded blocks that spatially surround the current block; temporally neighboring MV candidates, which are MVs of surrounding blocks in a coded reference picture onto which a position of the current block is projected; combined MV candidates, which are MVs generated by combining MV values of the spatially neighboring MV predictor and MV values of the temporally neighboring MV predictor; and a zero MV candidate, which is an MV having a zero value.
[0411] Next, the inter predictor 126 selects one of the MV candidates registered in the MV candidate list and determines the MV candidate as the MV of the current block.
[0412] Further, the entropy encoder 110 writes and encodes merge_idx, which is a signal indicating which MV candidate has been selected, in the stream.
[0413] It should be noted that the MV candidates registered in the MV candidate list described in Figure 41 The number of MV candidates can be different from the number of MV candidates in the figure, and the MV candidate list can be configured in such a manner that some of the kinds of candidate MVs in the figure can not be included, or one or more candidate MVs other than the kinds of MV candidates in the figure can be included.
[0414] The final MV can be determined by performing dynamic motion vector refresh (DMVR) described later using the MV of the current block derived by the normal merge mode. It should be noted that, in the normal merge mode, the motion information is encoded and the MV difference is not encoded. In the MMVD mode, one of the MV candidates is selected from the MV candidate list as in the case of the normal merge mode, and the MV difference is encoded. As Figure 38B indicated in the figure, MMVD can be classified as a merge mode together with the normal merge mode. It should be noted that the MV difference in the MMVD mode does not always need to be the same as the MV difference used for the inter mode. For example, the MV difference derivation in the MMVD mode can be a process that requires a smaller amount of processing than the amount of processing required for the MV difference derivation in the inter mode.
[0415] In addition, a combined inter / intra prediction (CIIP) mode can be performed. This mode is used to overlap a prediction image generated in inter prediction and a prediction image generated in intra prediction to generate a prediction image of a current block.
[0416] Note that the MV candidate list can be referred to as a candidate list. In addition, merge_idx is MV selection information.
[0417] (MV derivation > HMVP mode)
[0418] Figure 42 is a conceptual diagram for showing one example of an MV derivation process for a current picture using the HMVP merge mode.
[0419] In the normal merge mode, the MV of, for example, a CU that is a current block is determined by selecting one MV candidate from a list of MVs generated from a reference coded block (e.g., a CU). Here, another MV candidate can be registered in the MV candidate list. The mode of registering such another MV candidate is referred to as the HMVP mode.
[0420] In the HMVP mode, a first-in first-out (FIFO) server of the HMVP is used to manage MV candidates, separately from the MV candidate list used for the normal merge mode.
[0421] In the FIFO buffer, the latest motion information (e.g., MV) of a past processed block is stored first. In the management of the FIFO buffer, every time a block is processed, the MV of the latest block (i.e., the previous processed CU) is deposited in the FIFO buffer, and the MV of the oldest CU (i.e., the earliest processed CU) is deleted from the FIFO buffer. In Figure 42 In the example shown, HMVP1 is the MV of the latest block, and HMVP5 is the MV of the oldest MV.
[0422] Then, for example, the inter predictor 126 checks whether each MV managed in the FIFO buffer is a different MV from all the MV candidates that have been registered in the MV candidate list of the normal merge mode starting from HMVP1. When it is determined that the MV is different from all the MV candidates, the inter predictor 126 can add the MV managed in the FIFO buffer to the MV candidate list for the normal merge mode as an MV candidate. At this time, one or more candidate MVs in the FIFO buffer can be registered (added to the MV candidate list).
[0423] By using the HMVP mode in this way, not only the MV of a block that is spatially or temporally adjacent to the current block, but also the MV of a past processed block can be added. As a result, the variation of the MV candidates of the normal merge mode is expanded, which increases the possibility that the coding efficiency can be improved.
[0424] It should be noted that the MVs can be motion information. In other words, the information stored in the MV candidate list and the FIFO buffer can include not only MV values, but also reference picture information, reference directions, picture numbers, and the like. In addition, the block can be, for example, a CU.
[0425] It should be noted that, Figure 42 The MV candidate list and the FIFO buffer shown are examples. The size of the MV candidate list and the FIFO buffer can be different from Figure 42 in the MV candidate list and the FIFO buffer can be configured to register the MV candidates in an order different from Figure 42 In addition, the processes described herein can be common between the encoder 100 and the decoder 200.
[0426] It should be noted that the HMVP mode can be applied to modes other than the normal merge mode. For example, the latest motion information (e.g., MV) of a block that was processed in the affine mode in the past can be stored first and can be used as an MV candidate, which can improve efficiency. The mode obtained by applying the HMVP mode to the affine mode can be referred to as a history affine mode.
[0427] (MV derivation > FRUC mode)
[0428] Motion information can be derived at the decoder side without being signaled from the encoder side. For example, the motion information can be derived by performing motion estimation at the decoder 200 side. In one embodiment, at the decoder side, motion estimation is performed without using any pixel values in the current block. Modes for performing motion estimation at the decoder 200 side without using any pixel values in the current block include a frame rate up conversion (FRUC) mode, a pattern matching motion vector derivation (PMMVD) mode, and the like.
[0429] Figure 43 One example of a FRUC process in the form of a flowchart is explained in FIG. 6. First, a list is created, which indicates the MV of each encoded block that is spatially or temporally adjacent to the current block by referring to the MV as an MV candidate (this list can be the MV candidate list, or can be used as the MV candidate list for the normal merge mode (step Si_1).
[0430] Next, a best MV candidate is selected from the plurality of MV candidates registered in the MV candidate list (step Si_2). For example, evaluation values of the respective MV candidates included in the candidate MV list are calculated, and one MV candidate is selected based on the evaluation values. Based on the selected motion vector candidate, a motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector of the current block. Alternatively, for example, the motion vector of the current block can be derived using pattern matching in a surrounding area of a certain position in a reference picture, where the position in the reference picture corresponds to the selected motion vector candidate. In other words, estimation using pattern matching and evaluation values can be performed in a surrounding area of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, the update of the motion vector that produces the better evaluation value can not be performed.
[0431] Finally, the inter predictor 126 generates a predicted image of the current block by performing motion compensation of the current block using the derived MV and the coded reference picture (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a slice, the inter prediction of the slice using the FRUC mode is completed. For example, when the processes in steps Si_1 to Si_5 are performed for all blocks in a picture, the inter prediction of the picture using the FRUC mode is completed. It should be noted that not all blocks included in a slice can go through these processes in steps Si_1 to Si_5, and when part of the blocks go through these processes, the inter prediction of the slice using the FRUC mode can be completed. When the processes in steps Si_1 to Si_5 are performed for part of the blocks included in a picture in a similar manner, the inter prediction of the picture using the FRUC mode can be completed.
[0432] Similar processes can be performed in sub-block units.
[0433] The evaluation value can be calculated according to various methods. For example, a comparison is made between a reconstructed image in a region in a reference picture corresponding to the motion vector and a reconstructed image in a determined region (which can be, for example, a region in another reference picture or a region in a neighboring block of the current picture, as shown below). The determined region can be predetermined.
[0434] The difference between the pixel values of the two reconstructed images can be used for the evaluation value of the motion vector. It should be noted that information other than the difference value can be used to calculate the evaluation value.
[0435] Next, an example of pattern matching is described in detail. First, one MV candidate included in an MV candidate list (e.g., a merge list) is selected as a starting point of estimation by pattern matching. For example, as the pattern matching, a first pattern matching or a second pattern matching can be used. The first pattern matching and the second pattern matching can be referred to as bilateral matching and template matching, respectively.
[0436] (MV derivation > FRUC > bilateral matching)
[0437] In the first pattern matching, pattern matching is performed between two blocks distributed along a motion trajectory of a current block and included in two different reference pictures. Thus, in the first pattern matching, a region in another reference picture along the motion trajectory of the current block is used as a determined region for calculating an evaluation value of the above-described candidate. The determined region can be predetermined.
[0438] Figure 44 is a conceptual diagram for illustrating one example of the first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in Figure 44 In the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating a best matching pair between pairs included in two different reference pictures (Ref0, Ref1) and the two motion vectors are distributed along a motion trajectory of a current block (Cur block). More specifically, a difference between a reconstructed image at a specified position in a first coded reference picture (Ref0) specified by an MV candidate and a reconstructed image at a specified position in a second coded reference picture (Ref1) specified by a symmetric MV obtained by scaling the candidate MV with a display time interval is derived for the current block, and an evaluation value is calculated using the obtained difference value. An MV candidate that produces a best evaluation value and can produce a good result among a plurality of MV candidates can be selected as a final MV.
[0439] Under the assumption of a continuous motion trajectory, motion vectors (MV0, MV1) of two reference blocks are proportional to temporal distances (TD0, TD1) between a current picture (Cur Pic) and two reference pictures (Ref0, Ref1). For example, when the current picture is located in time between the two reference pictures and the temporal distances of the current picture to the respective two reference pictures are equal to each other, bidirectional motion vectors that are mirror-symmetric are derived in the first pattern matching.
[0440] (MV derivation > FRUC > template matching)
[0441] In the second pattern matching (template matching), pattern matching is performed between a block in the reference picture and a template in the current picture (the template is a block neighboring the current block in the current picture (e.g., the neighboring block is an above neighboring block and / or a left neighboring block)). Thus, in the second pattern matching, the block neighboring the current block in the current picture is used as a determination region for calculating an evaluation value of the above-mentioned MV candidate.
[0442] Figure 45 is a conceptual diagram for illustrating one example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in Figure 45 , in the second pattern matching, a motion vector of the current block (Cur block) is derived by estimating a block in the reference picture (Ref0) that best matches the block neighboring the current block in the current picture (Cur Pic). More specifically, a difference between a reconstructed image in a coding region that is left neighboring and above neighboring or left neighboring or above neighboring and a reconstructed image in a corresponding region in the coding reference picture (Ref0) and specified by an MV candidate is derived, and an evaluation value is calculated using the obtained difference value. An MV candidate that produces the best evaluation value among a plurality of MV candidates can be selected as a best MV candidate.
[0443] This information indicating whether the FRUC mode is applied (e.g., referred to as a FRUC flag) can be signaled at a CU level. In addition, when the FRUC mode is applied (e.g., when the FRUC flag is true), information indicating an applicable pattern matching method (e.g., the first pattern matching or the second pattern matching) can be signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the CU level, but can also be performed at other levels (e.g., at a sequence level, a picture level, a slice level, a tile level, a CTU level, or a sub-block level).
[0444] (MV derivation > affine mode)
[0445] The affine mode is a mode in which an MV is generated using an affine transformation. For example, the MV can be derived in sub-block units based on motion vectors of a plurality of neighboring blocks. This mode is also referred to as an affine motion compensated prediction mode.
[0446] Figure 46A is a conceptual diagram for illustrating one example of MV derivation in sub-block units based on motion vectors of a plurality of neighboring blocks. In Figure 46AIn the example, the current block includes sixteen 4×4 sub-blocks. Here, the motion vector V0 of the upper left corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and the motion vector V1 of the upper right corner control point in the current block is also derived based on the motion vectors of the adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (v x , v y ).
[0447] [Mathematics.1]
[0448]
[0449] Here, x and y represent the horizontal position and vertical position of the sub-block, respectively, and w represents a determined weighting coefficient. The determined weighting coefficient may be predetermined.
[0450] Such information indicating the affine mode (e.g., referred to as an affine flag) may be signaled at the CU level. It should be noted that the signaling of the information indicating the affine mode does not necessarily need to be performed at the CU level, but may also be performed at other levels (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0451] In addition, the affine mode can include several modes of different methods for deriving motion vectors at the upper left and upper right control points. For example, the affine mode includes two modes: affine inter mode (also called affine normal inter mode) and affine merge mode.
[0452] (MV Derivation > Affine Mode)
[0453] Figure 46B : is a conceptual diagram for illustrating an example of MV derivation in units of sub-blocks in an affine mode in which three control points are used. Figure 46B In the example, the current block includes sixteen 4×4 blocks. Here, the motion vector V0 of the upper left corner control point in the current block is derived based on the motion vectors of the adjacent blocks. Here, the motion vector V1 of the upper right corner control point of the current block is derived based on the motion vectors of the adjacent blocks, and the motion vector V2 of the lower left corner control point of the current block is also derived based on the motion vectors of the adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v x , v y ).
[0454] [Mathematics.2]
[0455]
[0456] Here, x and y represent horizontal and vertical positions of a sub-block, respectively, and w and h can be weighting factors, which can be predetermined weighting factors. In one embodiment, w can represent a width of the current block, and h can represent a height of the current block.
[0457] The affine mode using different number of control points (e.g., two and three control points) can be switched and signaled at the CU level. It should be noted that the information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).
[0458] In addition, such affine mode using three control points can include different methods for deriving the motion vectors at the top-left, top-right, and bottom-left control points. For example, as in the case of affine mode using two control points, the affine mode using three control points can include both affine inter mode and affine merge mode.
[0459] It should be noted that in the affine mode, the size of each sub-block contained in the current block can not be limited to 4x4 pixels, but can be other sizes. For example, the size of each sub-block can be 8x8 pixels.
[0460] (MV derivation > affine mode > control points)
[0461] Figure 47A , Figure 47B and Figure 47C are conceptual diagrams for illustrating examples of MV derivation at control points in the affine mode.
[0462] As shown in Figure 47A , in the affine mode, the motion vector predictor at each control point of the current block is calculated, for example, based on the multiple motion vectors corresponding to the blocks coded according to the affine mode among the neighboring coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) of the current block. More specifically, the coded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) are examined in the listed order, and the first valid block coded according to the affine mode is identified. The motion vector predictor at the control point of the current block is calculated based on the multiple motion vectors corresponding to the identified block.
[0463] For example, as Figure 47BAs shown in FIG. 10, when a block A adjacent to the left side of the current block has been coded according to the affine mode using two control points, motion vectors v3 and v4 projected at the top-left position and the top-right position of the coded block including the block A are derived. Then, the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block are calculated according to the derived motion vectors v3 and v4.
[0464] For example, as shown in FIG. 11, when a block A adjacent to the left side of the current block has been coded according to the affine mode using three control points, motion vectors v3, v4 and v5 projected at the top-left position, the top-right position and the bottom-left position of the coded block including the block A are derived. Then, the motion vector v0 at the top-left control point of the current block, the motion vector v1 at the top-right control point of the current block and the motion vector v2 at the bottom-left control point of the current block are calculated according to the derived motion vectors v3, v4 and v5. Figure 47C
[0465] The MV derivation method shown in FIG. 10 can be used for the MV derivation of each control point of the current block in step Sk_1 shown in FIG. 9, or can be used for the MV predictor derivation at each control point of the current block in step Sj_1 shown in FIG. 10. Figures 47A to 47C Figure 50 The MV derivation method shown in FIG. 11 can be used for the MV derivation of each control point of the current block in step Sk_1 shown in FIG. 9, or can be used for the MV predictor derivation at each control point of the current block in step Sj_1 shown in FIG. 10. Figure 51
[0466] Figure 48A and 48B are conceptual diagrams for showing examples of the MV derivation at the control points in the affine mode.
[0467] Figure 48A is a conceptual diagram for showing an example affine mode in which two control points are used.
[0468] In the affine mode, as shown in FIG. 10, a MV selected from MVs at coded blocks A, B and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Likewise, a MV selected from MVs at coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block. Figure 48A
[0469] is a conceptual diagram for showing an example affine mode in which three control points are used. Figure 48B In the affine mode, as shown in FIG. 11, MVs selected from MVs at coded blocks A, B, C and D adjacent to the current block are used as the motion vector v0 at the top-left control point of the current block, the motion vector v1 at the top-right control point of the current block and the motion vector v2 at the bottom-left control point of the current block.
[0470] Figure 48B As shown, the MV selected from the MVs of the neighboring coded blocks A, B and C of the current block is used as the motion vector v0 at the top-left corner control point of the current block. Likewise, the MV selected from the MVs of the neighboring coded blocks D and E of the current block is used as the motion vector v1 at the top-right corner control point of the current block. In addition, the MV selected from the MVs of the neighboring coded blocks F and G of the current block is used as the motion vector v2 at the bottom-left corner control point of the current block.
[0471] It should be noted that, Figure 48A and Figure 48B The MV derivation method shown can be used in the MV predictor derivation at each control point of the current block in the step Sj_1 described later, or can be used in the MV predictor derivation at each control point of the current block in the step Sj_1 described later. Figure 50 Figure 51 The MV predictor derivation at each control point of the current block in the step Sj_1 described later.
[0472] Here, when it is possible to switch and signal the use of affine modes using different numbers of control points (e.g., two and three control points) at the CU level, the number of control points of the coded block and the number of control points of the current block can be different from each other.
[0473] Figure 49A and Figure 49B is a conceptual diagram for illustrating an example of a method of MV derivation at control points when the number of control points used for a coded block and the number of control points used for a current block are different from each other.
[0474] For example, as shown in Figure 49A the current block has three control points at the top-left corner, the top-right corner and the bottom-left corner, and the bottom-right corner and the block A adjacent to the left side of the current block have been coded according to an affine mode in which two control points are used. In this case, the motion vectors v3 and v4 projected at the top-left corner position and the top-right corner position in the coded blocks including the block A are derived. Then the motion vector v0 at the top-left corner control point of the current block and the motion vector v1 at the top-right corner control point of the current block are calculated from the derived motion vectors v3 and v4. In addition, the motion vector v2 at the bottom-left corner control point is calculated from the derived motion vectors v0 and v1.
[0475] For example, as shown in Figure 49B As shown, the current block has two control points at the upper left and upper right corners, and the upper right corner and block A adjacent to the left of the current block have been encoded according to an affine mode using two control points. In this case, motion vectors v3, v4, and v5 are derived for the upper left corner position, the upper right corner position, and the lower left corner position of the coded block including block A. Motion vector v0 for the control point at the upper left corner of the current block and motion vector v1 for the control point at the upper right corner of the current block are then calculated based on the derived motion vectors v3, v4, and v5.
[0476] It should be pointed out that Figure 49A and Figure 49B The MV derivation method shown can be used in the following description Figure 50 The MV derivation of each control point of the current block in step Sk_1 shown in FIG. 1 is described later, or can be used in the following description. Figure 51 The MV predictor at each control point of the current block is derived in step Sj_1 as shown.
[0477] (MV derivation > affine mode > affine merge mode)
[0478] Figure 50 is a flowchart showing one example of a process in affine merge mode.
[0479] In the affine merge mode as shown in the figure, first, the inter-frame predictor 126 derives the MV at each control point of the current block (step Sk_1). The control point is the upper left corner point of the current block and the upper right corner point of the current block, such as Figure 46A As shown, or the upper left corner point of the current block, the upper right corner point of the current block and the lower left corner point of the current block, as shown Figure 46B The inter-frame predictor 126 may encode MV selection information to identify two or three derived MVs in the stream.
[0480] For example, when using Figures 47A to 47C When the MV derivation method shown is used, Figure 47A As shown, the inter-frame predictor 126 examines coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) and identifies the first valid block coded according to the affine mode.
[0481] The inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine mode to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as Figure 47BAs shown, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector v1 at the top-right corner control point of the current block from the motion vectors v3 and v4 including the top-left corner and the top-right corner of the coded block of the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block and the motion vector v1 at the top-right corner control point of the current block by projecting the motion vectors v3 and v4 at the top-left corner and the top-right corner of the coded block onto the current block.
[0482] Alternatively, when the block A is identified and the block A has three control points, as shown in FIG. 6B, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block, the motion vector v1 at the top-right corner control point of the current block, and the motion vector v2 at the bottom-left corner control point of the current block from the motion vectors v3, v4, and v5 including the top-left corner, the top-right corner, and the bottom-left corner of the coded block of the block A. For example, the inter predictor 126 calculates the motion vector v0 at the top-left corner control point of the current block, the motion vector v1 at the top-right corner control point of the current block, and the motion vector v2 at the bottom-left corner control point of the current block by projecting the motion vectors v3, v4, and v5 at the top-left corner, the top-right corner, and the bottom-left corner of the coded block onto the current block. Figure 47C It should be noted that, as shown in FIG. 6A, when the block A is identified and the block A has two control points, the MVs at the three control points can be calculated, and as shown in FIG. 6B, when the block A is identified and the block A has three control points, the MVs at two control points can be calculated.
[0483] Figure 49A It should be noted that, as shown in FIG. 6A, when the block A is identified and the block A has two control points, the MVs at the three control points can be calculated, and as shown in FIG. 6B, when the block A is identified and the block A has three control points, the MVs at two control points can be calculated. Figure 49B
[0484] Next, the inter predictor 126 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter predictor 126 calculates the MV of each of the plurality of sub-blocks as an affine MV using, for example, the two motion vectors v0 and v1 and the above expression (1A) or the three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter predictor 126 then performs motion compensation of the sub-blocks using these affine MVs and the coded reference picture (step Sk_3). When the processes in steps Sk_2 and Sk_3 are performed for each of all the sub-blocks included in the current block, the process for generating a prediction image using the affine merge mode for the current block ends. In other words, motion compensation of the current block is performed to generate a prediction image of the current block.
[0485] It should be noted that the above MV candidate list can be generated in step Sk_1. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, Figures 47A to 47C the MV derivation method shown in Figure 48A and 48B the MV derivation method shown in Figure 49A and 49B any combination of the MV derivation method shown in
[0486] It should be noted that, in addition to the affine mode, the MV candidate list can include MV candidates in modes in which prediction is performed in units of sub-blocks.
[0487] It should be noted that, for example, the MV candidate list including MV candidates can be generated as the MV candidate list in the affine merge mode using two control points and the affine merge mode using three control points. Alternatively, the MV candidate list including MV candidates in the affine merge mode using two control points and the MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, the MV candidate list including MV candidates can be generated in one of the affine merge mode using two control points and the affine merge mode using three control points. The MV candidates can be, for example, MVs for the coding blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), or MVs of effective blocks among these blocks.
[0488] It should be noted that an index indicating one of the MVs in the MV candidate list can be transmitted as the MV selection information.
[0489] (MV derivation > affine mode > affine inter mode)
[0490] Figure 51 is a flowchart showing one example of a procedure in the affine inter mode.
[0491] In the affine inter mode, first, the inter predictor 126 derives MV predictor values (v0, v1) or (v0, v1, v2) of the respective two or three control points of the current block (step Sj_1). As shown in Figure 46A or Figure 46B The control points can be, for example, the upper left corner point of the current block, the upper right corner point of the current block, and the lower left corner point of the current block.
[0492] For example, when the MV derivation method shown in Figure 48A and 48B is used, the inter predictor 126 derives the MV predictor values (v0, v1) or (v0, v1, v2) of the respective two or three control points of the current block by performing the following steps. Figure 48A or Figure 48BThe MV of an arbitrary block is selected from among the encoded blocks in the vicinity of each of the control points of the current block, and the MV predictors (v0, v1) or (v0, v1, v2) are derived at the corresponding two or three control points of the current block. At this time, the inter predictor 126 encodes, in the stream, MV predictor selection information for identifying the selected two or three MV predictors.
[0493] For example, the inter predictor 126 can determine, from among the encoded blocks adjacent to the current block, a block from which one MV is selected as an MV predictor at a control point, using cost evaluation or the like, and can write, in the bitstream, a flag indicating which MV predictor has been selected. In other words, the inter predictor 126 outputs, to the entropy encoder 110, the MV predictor selection information (e.g., the flag) as a prediction parameter through the prediction parameter generator 130.
[0494] Next, the inter predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictors selected or derived in step Sj_1 (step Sj_2). In other words, the inter predictor 126 calculates, using the above-described expression (1A) or expression (1B), the MV of each sub-block corresponding to the updated MV predictor as an affine MV (step Sj_3). The inter predictor 126 then performs motion compensation of the sub-blocks using these affine MVs and the encoded reference picture (step Sj_4). The processes in steps Sj_3 and Sj_4 are performed for all blocks in the current block while updating the MV predictors in step Sj_2. As a result, for example, the inter predictor 126 determines the MV predictor that yields the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter predictor 126 also encodes, in the stream, the difference between the determined MV and the MV predictor as an MV difference. In other words, the inter predictor 126 outputs, to the entropy encoder 110, the MV difference as a prediction parameter through the prediction parameter generator 130.
[0495] Finally, the inter predictor 126 generates a prediction image of the current block by performing motion compensation of the current block using the determined MV and the encoded reference picture (step Sj_6).
[0496] It should be noted that the above-described MV candidate list can be generated in step Sj_1. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example, Figures 47A-47C the MV derivation method shown in FIG. 2, Figure 48A the MV derivation method shown in FIG. 3, 48B the MV derivation method shown in FIG. 4, Figure 49A the MV derivation method shown in FIG. 5, 49B the MV derivation method shown in FIG. 6, and any combination of other MV derivation methods.
[0497] It should be noted that, in addition to the affine mode, the MV candidate list can include MV candidates in modes in which prediction is performed in sub-block units.
[0498] It should be noted that, for example, the MV candidate list including MV candidates can be generated as the MV candidate list in the affine inter mode using two control points and the affine inter mode using three control points. Alternatively, the MV candidate list including MV candidates in the affine inter mode using two control points and the MV candidate list including MV candidates in the affine inter mode using three control points can be generated respectively. Alternatively, the MV candidate list including MV candidates can be generated in one of the affine inter mode using two control points and the affine inter mode using three control points. The MV candidates can be, for example, MVs for the blocks A (left), B (above), C (upper right), D (lower left), and E (upper left), or MVs of effective blocks among these blocks.
[0499] It should be noted that an index indicating one of the MV candidates in the MV candidate list can be transmitted as the MV predictor selection information.
[0500] (MV derivation > triangle mode)
[0501] In the above example, the inter predictor 126 generates one rectangular prediction image for the current rectangular block. However, the inter predictor 126 can generate a plurality of prediction images each having a shape different from the rectangle of the current rectangular block, and can combine the plurality of prediction images to generate a final rectangular prediction image. The shape different from the rectangle can be, for example, a triangle.
[0502] Figure 52A is a conceptual diagram for illustrating generation of two triangular prediction images.
[0503] The inter predictor 126 generates a triangular prediction image by performing motion compensation on a first partition having a triangular shape in the current block using a first MV of the first partition to generate a triangular prediction image. Likewise, the inter predictor 126 generates a triangular prediction image by performing motion compensation on a second partition having a triangular shape in the current block using a second MV of the second partition to generate a triangular prediction image. The inter predictor 126 then generates a prediction image having a rectangular shape identical to that of the current block by combining these prediction images.
[0504] It should be noted that a first prediction image having a rectangular shape corresponding to the current block can be generated using the first MV as the prediction image for the first partition. Further, a second prediction image having a rectangular shape corresponding to the current block can be generated using the second MV as the prediction image for the second partition. The prediction image for the current block can be generated by performing a weighted addition of the first prediction image and the second prediction image. It should be noted that the portion of the weighted addition can be a partial region that spans the first partition and the second partition boundary.
[0505] Figure 52B is a conceptual diagram for illustrating an example of a first portion of a first partition that overlaps with a second partition and a first and second set of samples that can be weighted as part of a correction process. The first portion can be, for example, one quarter of the width or height of the first partition. In another example, the first portion can have a width that corresponds to N samples adjacent to an edge of the first partition, where N is an integer greater than zero, for example, N can be the integer 2. As shown, Figure 52B the left example of illustrates a rectangular partition that has a width that is one quarter of the width of the first partition, where the first set of samples includes samples outside of the first portion and samples inside of the first portion and the second set of samples includes samples inside of the first portion. Figure 52B the center example of illustrates a rectangular partition that has a height that is one quarter of the height of the first partition, where the first set of samples includes samples outside of the first portion and samples inside of the first portion and the second set of samples includes samples inside of the first portion. Figure 52B the right example of is a triangular partition that has a polygonal portion in addition to a triangular portion that has a height that corresponds to two samples, where the first set of samples includes samples outside of the first portion and samples inside of the first portion and the second set of samples includes samples inside of the first portion.
[0506] The first portion can be a portion of the first partition that overlaps with an adjacent partition. Figure 52C is a conceptual diagram for illustrating a first portion of a first partition that is a portion of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition is shown that has an overlapping portion with a rectangular partition that is spatially adjacent. Partitions having other shapes, such as triangular partitions, can be employed and the overlapping portion can overlap with a spatially or temporally adjacent partition.
[0507] Further, while examples are given of generating prediction images for each of the two partitions using inter prediction, prediction images can be generated for at least one of the partitions using intra prediction.
[0508] Figure 53 is a flowchart illustrating one example of a process in a triangular mode.
[0509] In the triangular mode, first, the inter-frame predictor 126 splits the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode partition information as prediction parameters in the stream as information related to the splitting into the partitions. In other words, the inter-frame predictor 126 can output the partition information as the prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0510] First, the inter-frame predictor 126 obtains a plurality of MV candidates for the current block based on information such as MVs of a plurality of coded blocks temporally or spatially surrounding the current block (step Sx_2). In other words, the inter-frame predictor 126 generates an MV candidate list.
[0511] The inter-frame predictor 126 then selects an MV candidate for the first partition and an MV candidate for the second partition from the plurality of MV candidates obtained in step Sx_1 as a first MV and a second MV, respectively (step Sx_3). At this time, the inter-frame predictor 126 encodes MV selection information for identifying the selected MV candidates in the stream as prediction parameters. In other words, the inter-frame predictor 126 outputs the MV selection information as the prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.
[0512] Next, the inter-frame predictor 126 generates a first prediction image by performing motion compensation using the selected first MV and the coded reference picture (step Sx_4). Likewise, the inter-frame predictor 126 generates a second prediction image by performing motion compensation using the selected second MV and the coded reference picture (step Sx_5).
[0513] Finally, the inter-frame predictor 126 generates a prediction image for the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_6).
[0514] It should be noted that, although the first partition and the second partition are triangular in the example shown, Figure 52A the first partition and the second partition can be trapezoidal, or other shapes different from each other. In addition, although the current block includes two partitions in the example shown, Figure 52A and 52C the current block can include three or more partitions.
[0515] In addition, the first partition and the second partition can overlap each other. In other words, the first partition and the second partition can include the same pixel region. In this case, a prediction image in the first partition and a prediction image in the second partition can be used to generate a prediction image for the current block.
[0516] Furthermore, although an example has been shown in which a prediction image is generated for each of the two partitions using inter prediction, a prediction image can be generated for at least one partition using intra prediction.
[0517] It should be noted that the MV candidate list used for selecting the first MV and the MV candidate list used for selecting the second MV can be different from each other, or the MV candidate list used for selecting the first MV can also be used as the MV candidate list used for selecting the second MV.
[0518] It should be noted that the partition information can include an index indicating a split direction in which the at least current block is split into a plurality of partitions. The MV selection information can include an index indicating the selected first MV and an index indicating the selected second MV. One index can indicate multiple pieces of information. For example, one index that collectively indicates part or all of the partition information and part or all of the MV selection information can be encoded.
[0519] (MV derivation > ATMVP mode)
[0520] Figure 54 is a conceptual diagram for showing one example of an advanced temporal motion vector prediction (ATMVP) mode in which an MV is derived in units of sub-blocks.
[0521] The ATMVP mode is a mode that is categorized into the merge mode. For example, in the ATMVP mode, MV candidates of each sub-block are registered in an MV candidate list for the normal merge mode.
[0522] More specifically, in the ATMVP mode, first, as shown in Figure 54 , a temporal MV reference block associated with the current block is identified in a coded reference picture specified by an MV (MV0) of a neighboring block located at a lower-left position with respect to the current block. Next, in each sub-block in the current block, an MV used for coding a corresponding region of a sub-block in the temporal MV reference block is identified. The MV identified in this way is included in the MV candidate list as an MV candidate of the sub-block in the current block. When an MV candidate of each sub-block is selected from the MV candidate list, the sub-block is motion-compensated with the MV candidate serving as an MV of the sub-block. In this way, a prediction image of each sub-block is generated.
[0523] Although in the example shown in Figure 54 , a block located at a lower-left position with respect to the current block is used as a surrounding MV reference block, it should be noted that another block can be used. In addition, the size of the sub-block can be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block can be switched in units of a slice, a block, a picture, or the like.
[0524] (MV derivation > DMVR)
[0525] Figure 55 is a flowchart illustrating a relationship between the merge mode and the decoding motion vector refinement DMVR.
[0526] The inter predictor 126 derives a motion vector of the current block according to the merge mode (step S1_1). Next, the inter predictor 126 determines whether or not to perform estimation of the motion vector, i.e., motion estimation (step S1_2). Here, when it is determined not to perform the motion estimation (NO in step S1_2), the inter predictor 126 determines the motion vector derived in step S1_1 as a final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merge mode.
[0527] When it is determined to perform the motion estimation in step S1_1 (YES in step S1_2), the inter predictor 126 derives a final motion vector of the current block by estimating a surrounding area of a reference picture specified by the motion vector derived in step S1_1 (step S1_3). In other words, in this case, the motion vector of the current block is determined according to the DMVR.
[0528] Figure 56 is a conceptual diagram for illustrating one example of a DMVR process for determining an MV.
[0529] First, MV candidates (L0 and L1) are selected for the current block, e.g., in the merge mode. Reference pixels are identified from a first reference picture (L0), which is an encoded picture in the L0 list, according to the MV candidate (L0). Likewise, reference pixels are identified from a first reference picture (L1), which is an encoded picture in the L1 list, according to the MV candidate (L1). A template is generated by computing an average of these reference pixels.
[0530] Next, the template is used to estimate each surrounding area of the MV candidates of the first reference picture (L0) and the second reference picture (L1), and the MV that results in the smallest cost is determined as the final MV. It should be noted that the cost can be computed, e.g., with a difference between each pixel value in the template and a corresponding pixel value in the estimated area, a value of the MV candidate, etc.
[0531] It is not always necessary to perform the exact same process described here. Other processes for achieving derivation of the final MV through estimation in the surrounding area of the MV candidate can be used.
[0532] Figure 57 is a conceptual diagram for illustrating another example of DMVR for determining an MV. Unlike the example of DMVR illustrated in Figure 56 , in the example illustrated in Figure 57 , the cost is computed without generating a template.
[0533] First, the inter predictor 126 estimates a surrounding area of a reference block in each reference picture included in the L0 list and the L1 list based on initial MVs (which are MV candidates obtained from each MV candidate list). For example, as shown in Figure 57 the initial MV corresponding to the reference block in the L0 list is InitMV L0, and the initial MV corresponding to the reference block in the L1 list is InitMV L1. In the motion estimation, the inter predictor 126 first sets a search position for the reference picture in the L0 list. Based on the position indicated by the vector difference indicating the search position, specifically, the initial MV (i.e., InitMV L0, with the vector difference of the search position being MVd L0). The inter predictor 126 then determines an estimated position in the reference picture in the L1 list. The search position is represented by the vector difference of the search position from the position indicated by the initial MV (i.e., InitMV L1). More specifically, the inter predictor 126 determines the vector difference as MVd L1 by mirroring MVd L0. In other words, the inter predictor 126 determines a position symmetrical about the position indicated by the initial MV as the search position in each reference picture in the L0 list and the L1 list. The inter predictor 126 calculates the sum of absolute differences (SAD) between pixel values at the search position in the block as a cost with respect to each search position, and finds the search position that produces the minimum cost.
[0534] Figure 58A is a conceptual diagram for illustrating one example of motion estimation in DMVR, and Figure 58B is a flowchart illustrating one example of a motion estimation process.
[0535] First, in step 1, the inter predictor 126 calculates the cost between the search position (also referred to as the starting point) indicated by the initial MV and eight surrounding search positions. The inter predictor 126 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at the search position other than the starting point is the minimum, the inter predictor 126 changes the target to the search position that obtains the minimum cost, and performs the process in step 2. When the cost at the starting point is the minimum, the inter predictor 126 skips the process in step 2 and performs the process in step 3.
[0536] In step 2, the inter predictor 126 performs a search similar to the process in step 1, with the changed target search position as a new starting point according to the result of the process in step 1. The inter predictor 126 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at the search position other than the starting point is the minimum, the inter predictor 126 performs the process in step 4. When the cost at the starting point is the minimum, the inter predictor 126 performs the process in step 3.
[0537] In step 4, the inter predictor 126 takes the search position at the origin as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0538] In step 3, the inter predictor 126 determines the pixel position with sub-pixel accuracy based on the costs at the four points at the up, down, left, right positions relative to the origin in step 1 or step 2 to obtain the minimum cost, and takes the pixel position as the final search position. The pixel position with sub-pixel accuracy is determined by weighted addition of each of the four vectors ((0, 1), (0, -1), (-1, 0) and (1, 0)) using the costs at the corresponding positions of the four search positions as the weights. The inter predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0539] (Motion compensation > BIO / OBMC / LIC)
[0540] Motion compensation involves modes for generating a prediction image and correcting the prediction image. The modes are, for example, bi-directional optical flow (BIO), overlapped block motion compensation (OBMC), local illumination compensation (LIC), etc., which will be described later.
[0541] Figure 59 is a flowchart showing one example of a process of generating a prediction image.
[0542] The inter predictor 126 generates a prediction image (step Sm_1), and, for example, corrects the prediction image according to, for example, any of the modes described above (step Sm_2).
[0543] Figure 60 is a flowchart showing another example of a process of generating a prediction image.
[0544] The inter predictor 126 determines a motion vector of a current block (step Sn_1). Next, the inter predictor 126 generates a prediction image using the motion vector (step Sn_2), and determines whether to perform a correction process (step Sn_3). Here, when it is determined to perform the correction process (Yes in step Sn_3), the inter predictor 126 generates a final prediction image by correcting the prediction image (step Sn_4). It should be noted that in LIC described later, the correction can be performed on the luminance and the chrominance in step Sn_4. When it is determined not to perform the correction process (No in step Sn_3), the inter predictor 126 outputs the prediction image as the final prediction image without correcting the prediction image (step Sn_5).
[0545] (Motion compensation > OBMC)
[0546] It should be noted that, in addition to the motion information of the current block obtained through motion estimation, the motion information of neighboring blocks can also be used to generate the inter prediction picture. More specifically, by weightedly adding the prediction picture (in the reference picture) based on the motion information obtained through motion estimation and the prediction picture (in the current picture) based on the motion information of the neighboring blocks, an inter prediction picture can be generated for each sub-block in the current block. Such inter prediction (motion compensation) is also referred to as overlapped block motion compensation (OBMC) or OBMC mode.
[0547] In the OBMC mode, information indicating the sub-block size for OBMC (e.g., referred to as OBMC block size) can be signaled at the sequence level. In addition, information indicating whether to apply the OBMC mode (e.g., referred to as OBMC flag) can be signaled at the CU level. It should be noted that the signaling of such information does not necessarily need to be performed at the sequence level and the CU level, but can also be performed at other levels (e.g., at the picture level, slice level, tile level, CTU level, or sub-block level).
[0548] The OBMC mode will be described in more detail. Figure 61 and Figure 62 are a flowchart and a conceptual diagram for illustrating an outline of the prediction picture correction process performed by OBMC.
[0549] First, as shown in Figure 62 , a normally motion-compensated prediction picture (Pred) is obtained using the MV assigned to the current block. In Figure 62 , the arrow “MV” points to the reference picture and indicates what the current block of the current picture refers to in order to obtain the prediction picture.
[0550] Next, a prediction picture (Pred_L) is obtained by applying a motion vector (MV_L) that has been derived for a coding block neighboring the left side of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow “MV_L”, which indicates the reference picture from the current block. The first correction of the prediction picture is performed by overlapping the two prediction pictures Pred and Pred_L. This provides the effect of blending the boundaries between neighboring blocks.
[0551] Likewise, a prediction picture (Pred_U) is obtained by applying an MV (MV_U) that has been derived for a coding block neighboring above the current block to the current block (reusing the MV of the current block). The MV (MV_U) is indicated by the arrow "MV_U" which indicates the reference picture from the current block. A second correction of the prediction picture is performed by overlapping the prediction picture Pred_U to the prediction picture (e.g. Pred and Pred_L) to which the first correction has been performed. This provides the effect of blending the boundaries between neighboring blocks. The prediction picture resulting from the second correction is a picture in which the boundaries between neighboring blocks have been blended (smoothed), and is thus the final prediction picture for the current block.
[0552] While the above example is a two-path correction method using a left neighboring block and an upper neighboring block, it is to be noted that the correction method can be a three-path or more path correction method using a right neighboring block and / or a lower neighboring block at the same time.
[0553] It is to be noted that the area in which this overlapping is performed can be only a part of the area near the block boundary, not the entire pixel area of the block.
[0554] It is to be noted that the prediction picture correction process for obtaining one prediction picture Pred from one reference picture by overlapping the additional prediction pictures Pred_L and Pred_U according to OBMC has been described above. However, when a prediction picture is corrected based on multiple reference pictures, a similar process can be applied to each of the multiple reference pictures. In this case, after the corrected prediction pictures are obtained from the respective reference pictures by performing the OBMC picture correction based on the multiple reference pictures, the resulting corrected prediction pictures are further overlapped to obtain the final prediction picture.
[0555] It is to be noted that in OBMC, the current block unit can be a PU, or a sub-block unit obtained by further splitting a PU.
[0556] One example of a method for determining whether to apply OBMC is a method using obmc_flag, which is a signal indicating whether to apply OBMC. As one specific example, the encoder 100 can determine whether the current block belongs to a region with complex motion. The encoder 100 sets the obmc_flag to a "1" value and applies OBMC at the time of encoding when the block belongs to a region with complex motion, and sets the obmc_flag to a "0" value and encodes the block without applying OBMC when the block does not belong to a region with complex motion. The decoder 200 switches between the application and non-application of OBMC by decoding the obmc_flag written in the stream.
[0557] (motion compensation > BIO)
[0558] Next, the MV derivation method is described. First, a mode for deriving MVs based on a model assuming uniform linear motion is described. This mode is also referred to as the bi-directional optical flow (BIO) mode. In addition, this bi-directional optical flow can be written as BDOF instead of BIO.
[0559] Figure 63 is a conceptual diagram for showing the model assuming uniform linear motion. In Figure 63 , (v x , v y ) denotes a velocity vector, τ0and τ1denote the temporal distance between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). (MV x0 , MV y0 ) denotes the MV corresponding to the reference picture Ref0, and (MV x1 , MV y1 ) denotes the MV corresponding to the reference picture Ref1.
[0560] Here, assuming that the uniform linear motion is represented by the velocity vector (v x , v y ), (MV x0 , MV y0 ) and (MV x1 , MV y1 ) are respectively expressed as (v xτ0 , v yτ0 ) and (-v xτ1 , -v yτ1 ), and the following optical flow equation (2) is given.
[0561] [math. 3]
[0562]
[0563] Here, I(k) denotes the motion-compensated luma value of the reference picture k (k = 0, 1). This optical flow equation indicates that (i) the time derivative of the luma value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image are zero. Based on a combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, the MV candidate list can be corrected in pixel units.
[0564] It should be noted that the motion vector can be derived at the decoder side 200 using a method different from the one based on the model assuming uniform linear motion. For example, the motion vector can be derived in sub-block units based on the motion vectors of a plurality of neighboring blocks.
[0565] Figure 64 is a flowchart showing one example of an inter prediction process according to BIO. Figure 65 is a functional block diagram showing one example of a functional configuration of an inter predictor 126 that can perform inter prediction according to BIO.
[0566] As shown in Figure 65 , the inter predictor 126 includes, for example, a memory 126a, a difference image deriver 126b, a gradient image deriver 126c, an optical flow deriver 126d, a correction value deriver 126e, and a prediction image corrector 126f. It should be noted that the memory 126a can be the frame memory 122.
[0567] The inter predictor 126 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from a picture (Cur Pic) including a current block. The inter predictor 126 then derives a prediction image of the current block using the two motion vectors (M0, M1) (step Sy_1). It should be noted that the motion vector M0 is a motion vector (MV x0 , MV y0 ) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MV x1 , MV y1 ) corresponding to the reference picture Ref1.
[0568] Next, the difference image deriver 126b derives a difference image I 0 of the current block using the motion vector M0 and the reference picture L0 by referring to the memory 126a. Subsequently, the difference image deriver 126b derives a difference image I 1 of the current block using the motion vector M1 and the reference picture L1 by referring to the memory 126a (step Sy_2). Here, the difference image I 0 is an image included in the reference picture Ref0 and to be derived for the current block, and the difference image I 1 is an image included in the reference picture Ref1 and to be derived for the current block. Each of the difference image I 0 and the difference image I 1 may be the same size as the current block. Alternatively, each of the difference image I 0 and the difference image I 1 may be an image larger than the current block. Furthermore, the difference image I 0 and the difference image I 1 may include a prediction image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.
[0569] In addition, the gradient image deriver 126c obtains the difference image I 0 Sum difference image I 1 Derive the gradient image of the current block (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) (Step Sy_3). It should be noted that the gradient image in the horizontal direction is (Ix 0 , 1x 1 ), the vertical gradient image is (Iy 0 , Iy 1 ). The gradient image deriver 126c may derive each gradient image by, for example, applying a gradient filter to the difference image. The gradient image may indicate the amount of spatial variation of pixel values in the horizontal direction, in the vertical direction, or in both directions.
[0570] Next, the optical flow deriver 126d uses the interpolated image (I 0 , I 1 ) and gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) is derived as an optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_4). The optical flow indicates a coefficient for correcting the amount of spatial pixel movement and can be referred to as a local motion estimation value, a corrected motion vector, or a corrected weight vector. As an example, the sub-block can be a 4×4 pixel sub-CU. It should be noted that the optical flow derivation can be performed for each pixel unit, etc., rather than for each sub-block.
[0571] Next, the inter-frame predictor 126 uses the optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value deriver 126e uses the optical flow (vx, vy) to derive correction values for the pixel values included in the current block (step Sy_5). The predicted image corrector 126f can then use the correction values to correct the predicted image of the current block (step Sy_6). It should be noted that the correction values can be derived in units of pixels, multiple pixels, or sub-blocks.
[0572] It should be noted that the BIO process is not limited to Figure 64 For example, you can just execute Figure 64 Some of the processes disclosed in the present invention may be added or replaced with different processes, or the processes may be performed in a different processing order, etc.
[0573] (Motion Compensation > LIC)
[0574] Next, one example of a mode for generating a prediction image (prediction) using a local illumination compensation (LIC) process is described.
[0575] Figure 66A is a conceptual diagram for illustrating one example of a process of a prediction image generation method using a luminance correction process performed by LIC. Figure 66B is a flowchart illustrating one example of a process of a prediction image generation method using LIC.
[0576] First, the inter-frame predictor 126 derives MVs from the coded reference pictures and obtains a reference image corresponding to the current block (step Sz_1).
[0577] Next, the inter-frame predictor 126 extracts, for the current block, information indicating how luminance values vary between the current block and the reference picture (step Sz_2). This extraction is performed based on luminance pixel values of a coded left-adjacent reference region (surrounding reference region) and a coded upper-adjacent reference region (surrounding reference region) in the current picture, and luminance pixel values at corresponding positions in the reference picture specified by the derived MVs. The inter-frame predictor 126 uses the information indicating how the luminance values vary to calculate luminance correction parameters (step Sz_3).
[0578] The inter-frame predictor 126 generates a prediction image of the current block by performing a luminance correction process in which the luminance correction parameters are applied to the reference image in the reference picture specified by the MVs (step Sz_4). In other words, the prediction image (which is the reference image in the reference picture specified by the MVs) is corrected based on the luminance correction parameters. In this correction, luminance can be corrected, or chrominance can be corrected, or both. In other words, chrominance correction parameters can be calculated using information indicating how the chrominance varies, and a chrominance correction process can be performed.
[0579] It should be noted that, Figure 66A The shape of the surrounding reference region shown is one example; another shape can be used.
[0580] Furthermore, although the processing of generating a prediction image from a single reference picture is described here, the case of generating a prediction image from multiple reference pictures can be described in the same way. The prediction image can be generated after performing a luminance correction process on a reference image obtained from the reference pictures in the same way as described above.
[0581] One example of a method for determining whether to apply LIC is a method using a lic_flag, which is a signal indicating whether to apply the LIC. As one specific example, the encoder 100 determines whether the current block belongs to a region having luminance variation. The encoder 100 sets the lic_flag to a "1" value and applies the LIC at the time of encoding when the block belongs to a region having luminance variation, and sets the lic_flag to a "0" value and performs encoding without applying the LIC when the block does not belong to a region having luminance variation. The decoder 200 can decode the lic_flag written in the stream and decode the current block by switching between applying and not applying the LIC according to the flag value.
[0582] One example of a different method of determining whether to apply the LIC process is a determination method according to whether the LIC process has been applied to a surrounding block. As one specific example, when the current block has been processed in merge mode, the inter predictor 126 determines whether the encoded surrounding block selected in the MV derivation in the merge mode has been encoded using the LIC. The inter predictor 126 performs encoding by switching between applying and not applying the LIC according to the result. Note that the same process is applied in the process at the decoder 200 side as well in this example.
[0583] The luminance correction (LIC) process has been described with reference to Figure 66A and Figure 66B and is further described below.
[0584] First, the inter predictor 126 derives an MV for obtaining a reference picture corresponding to the current block to be encoded from a reference picture, which is an encoded picture.
[0585] Next, the inter predictor 126 extracts information indicating how the luminance values of the reference picture change to the luminance values of the current picture using the luminance pixel values of the encoded surrounding reference region neighboring the left and top of the current block and the luminance values of the corresponding positions in the reference picture specified by the MV, and calculates luminance correction parameters. For example, assume that the luminance pixel value of a given pixel in the surrounding reference region in the current picture is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region in the reference picture is p1. The inter predictor 126 calculates the coefficients A and B for optimizing Axp1+B=p0 as the luminance correction parameters for a plurality of pixels in the surrounding reference region.
[0586] Next, the inter predictor 126 performs a luminance correction process using the luminance correction parameter of the reference picture in the reference picture specified by the MV to generate a prediction picture of the current block. For example, assume that the luminance pixel value in the reference picture is p2, and the luminance pixel value of the prediction picture after the luminance correction is p3. The inter predictor 126 generates the prediction picture after going through the luminance correction process by calculating A x p2 + B = p3 for each pixel in the reference picture.
[0587] For example, a region having a determined number of pixels extracted from each of the upper and left neighboring pixels can be used as the surrounding reference region. In addition, the surrounding reference region is not limited to a region adjacent to the current block, and can also be a region that is not adjacent to the current block. In this case, the surrounding reference region can be a region in the current picture that is not adjacent to the current block. Figure 66A In the example shown, the surrounding reference region in the reference picture can be a region in the current picture specified by another MV, from the surrounding reference region in the current picture. For example, the other MV can be a MV in the surrounding reference region in the current picture.
[0588] Although operations performed by the encoder 100 are described here, it should be noted that the decoder 200 performs similar operations.
[0589] It should be noted that LIC can be applied not only to luminance but also to chrominance. At this time, the correction parameter can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any one of Y, Cb, and Cr.
[0590] Furthermore, the LIC process can be applied in units of sub-blocks. For example, the correction parameter can be derived using the surrounding reference region in the current sub-block and the surrounding reference region in the reference sub-block in the reference picture specified by the MV of the current sub-block.
[0591] (Prediction controller)
[0592] The prediction controller 128 selects one of the intra prediction signal (the image or signal output from the intra predictor 124) and the inter prediction signal (the image or signal output from the inter predictor 126), and outputs the selected prediction picture to the subtracter 104 and the adder 116 as a prediction signal.
[0593] (Prediction parameter generator)
[0594] The prediction parameter generator 130 can output information related to intra prediction, inter prediction, selection of a prediction image in the prediction controller 128, and the like, as a prediction parameter to the entropy encoder 110. The entropy encoder 110 can generate a stream based on the prediction parameter input from the prediction parameter generator 130 and the quantized coefficients input from the quantizer 108. The prediction parameter can be used in the decoder 200. The decoder 200 can receive and decode the stream, and perform the same process as the prediction process performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128. The prediction parameter can include, for example, (i) a selected prediction signal (e.g., an MV, a prediction type, or a prediction mode used by the intra predictor 124 or the inter predictor 126), or (ii) a selectable index, a flag, or a value based on the prediction process performed in each of the intra predictor 124, the inter predictor 126, and the prediction controller 128, or a value indicating the prediction process.
[0595] (Decoder)
[0596] Next, a decoder 200 capable of decoding a stream output from the above-described encoder 100 will be described. Figure 67 is a block diagram showing a functional structure of the decoder 200 of the present embodiment. The decoder 200 is a device that decodes a stream that is an encoded image in units of blocks.
[0597] As shown in Figure 67 , the decoder 200 includes an entropy decoder 202, an inverse quantizer 204, an inverse transformer 206, an adder 208, a block memory 210, a loop filter 212, a frame memory 214, an intra predictor 216, an inter predictor 218, a prediction controller 220, a prediction parameter generator 222, and a split determiner 224. It should be noted that the intra predictor 216 and the inter predictor 218 are configured as part of a prediction performer.
[0598] (Installation example of decoder)
[0599] Figure 68 is a functional block diagram showing an installation example of the decoder 200. The decoder 200 includes a processor b1 and a memory b2. For example, as shown in Figure 67 , a plurality of constituent elements of the decoder 200 are installed on Figure 68 the processor b1 and the memory b2.
[0600] The processor b1 is a circuit that performs information processing and is coupled to the memory b2. For example, the processor b1 is a dedicated or general electronic circuit that decodes a stream. The processor b1 can be a processor, such as a CPU. Further, the processor b1 can be an aggregation of a plurality of electronic circuits. Further, for example, the processor b1 can function as Figure 67The roles of two or more of the plurality of constituent elements of the decoder 200 shown other than the constituent elements for storing information, and the like.
[0601] The memory b2 is a dedicated or general-purpose memory for storing information used by the processor b1 to decode the stream. The memory b2 can be an electronic circuit, and can be connected to the processor b1. Further, the memory b2 can be included in the processor b1. Further, the memory b2 can be an aggregation of a plurality of electronic circuits. In addition, the memory b2 can be a magnetic disk, an optical disk, or the like, or can be expressed as a storage, a recording medium, or the like. Further, the memory b2 can be a non-volatile memory or a volatile memory.
[0602] For example, the memory b2 can store an image or a stream. Further, the memory b2 can store a program for causing the processor b1 to decode the stream.
[0603] Further, for example, the memory b2 can function as Figure 67 two or more of the plurality of constituent elements of the decoder 200 shown for storing information. More specifically, the memory b2 can function as Figure 67 the block memory 210 and the frame memory 214 shown in FIG. 1. More specifically, the memory b2 can store a reconstructed image (specifically, a reconstructed block, a reconstructed picture, or the like).
[0604] It should be noted that not all of the plurality of constituent elements shown in the decoder 200 are implemented, and not all of the processes described herein are performed. Figure 67 the constituent elements shown in FIG. 1 are implemented, and not all of the processes described herein are performed. Figure 67 Part of the constituent elements shown in FIG. 1 can be included in another device, or part of the processes described herein can be performed by another device.
[0605] Hereinafter, the overall flow of the processes performed by the decoder 200 is described, and then each of the constituent elements included in the decoder 200 is described. Note that some of the constituent elements included in the decoder 200 perform the same processes as those performed by some of the constituent elements included in the encoder 100, and thus the same processes are not described in detail again. For example, the inverse quantizer 204, the inverse transformer 206, the adder 208, the block memory 210, the frame memory 214, the intra predictor 216, the inter predictor 218, the prediction controller 220, and the loop filter 212 included in the decoder 200 distribute the processes similar to those performed by the inverse quantizer 112, the inverse transformer 114, the adder 116, the block memory 118, the frame memory 122, the intra predictor 124, the inter predictor 126, the prediction controller 128, and the loop filter 120 included in the decoder 200.
[0606] (Overall flow of decoding process)
[0607] Figure 69 is a flowchart showing one example of the overall decoding process performed by the decoder 200.
[0608] First, the split determiner 224 in the decoder 200 determines a split pattern of each of a plurality of fixed-size blocks (128 x 128 pixels) included in a picture, based on parameters input from the entropy decoder 202 (step Sp_1). The split pattern is the split pattern selected by the encoder 100. The decoder 200 then performs the processes of steps Sp_2 to Sp_6 for each of the plurality of blocks of the split pattern.
[0609] The entropy decoder 202 decodes (specifically, entropy-decodes) the encoded quantized coefficients and the prediction parameters of the current block (step Sp_2).
[0610] Next, the inverse quantizer 204 inverse-quantizes the plurality of quantized coefficients, and the inverse transformer 206 inverse-transforms the result to restore the prediction residual (i.e., the difference block) (step Sp_3).
[0611] Next, the prediction executor including all or a part of the intra predictor 216, the inter predictor 218, and the prediction controller 220 generates a prediction signal of the current block (step Sp_4).
[0612] Next, the adder 208 adds the prediction image and the prediction residual to generate a reconstructed image (also referred to as a decoded image block) of the current block (step Sp_5).
[0613] When the reconstructed image is generated, the loop filter 212 performs filtering of the reconstructed image (step Sp_6).
[0614] The decoder 200 then determines whether or not the decoding of the entire picture has been completed (step Sp_7). When it is determined that the decoding has not been completed (NO in step Sp_7), the decoder 200 repeats the processes starting from step Sp_1.
[0615] It should be noted that the processes of these steps Sp_1 to Sp_7 can be sequentially performed by the decoder 200, or two or more of the processes can be performed in parallel. The processing order of two or more of the processes can be modified.
[0616] (Split determiner)
[0617] Figure 70 is a conceptual diagram for showing the relationship between the split determiner 224 and other constituent elements in the embodiment. As an example, the split determiner 224 can perform the following processes.
[0618] For example, the split determiner 224 collects block information from the block memory 210 or the frame memory 214, and further obtains parameters from the entropy decoder 202. The split determiner 224 can then determine a split pattern of the fixed size block based on the block information and the parameters. The split determiner 224 can then output information indicating the determined split pattern to the inverse transformer 206, the intra predictor 216, and the inter predictor 218. The inverse transformer 206 can perform inverse transform of the transform coefficients based on the split pattern indicated by the information from the split determiner 224. The intra predictor 216 and the inter predictor 218 can generate the predicted image based on the split pattern indicated by the information from the split determiner 224.
[0619] (Entropy decoder)
[0620] Figure 71 is a block diagram showing one example of a functional configuration of the entropy decoder 202.
[0621] The entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to a split pattern by entropy-decoding the stream. For example, CABAC is used for the entropy-decoding. More specifically, the entropy decoder 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a dequantizer 202c. The binary arithmetic decoder 202a uses a context value derived by the context controller 202b to arithmetic-decode the stream into a binary signal. The context controller 202b derives the context value in the same manner as the context controller 110b of the encoder 100 does, depending on the characteristics or surrounding state of the syntax element, i.e., the appearance probability of the binary signal. The dequantizer 202c performs dequantization so as to transform the binary signal output from the binary arithmetic decoder 202a into a multi-level signal representing the quantized coefficients as described above. This dequantization can be performed according to the dequantization method described above.
[0622] In this way, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can output the prediction parameters included in the stream to the intra predictor 216, the inter predictor 218, and the prediction controller 220 (see Figure 1 ). The intra predictor 216, the inter predictor 218, and the prediction controller 220 can perform the same prediction processes as those performed by the intra predictor 124, the inter predictor 126, and the prediction controller 128 on the encoder 100 side.
[0623] Figure 72 is a conceptual diagram for showing a flow of an example CABAC process in the entropy decoder 202.
[0624] First, initialization is performed in CABAC in the entropy decoder 202. In the initialization, initialization in the binary arithmetic decoder 202a and setting of initial context values are performed. The binary arithmetic decoder 202a and the dequantizer 202c then perform arithmetic decoding and dequantization of encoded data of, for example, a CTU. At this time, the context controller 202b updates the context values each time the arithmetic decoding is performed. The context controller 202b then saves the context values as a post-process. For example, the saved context values are used to initialize the context values for the next CTU.
[0625] (Inverse quantizer)
[0626] The inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block, which are input from the entropy decoder 202. More specifically, the inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block based on the quantization parameter corresponding to the quantized coefficients. The inverse quantizer 204 then outputs the inverse-quantized transform coefficients (i.e., transform coefficients) of the current block to the inverse transformer 206.
[0627] Figure 73 is a block diagram illustrating one example of a functional configuration of the inverse quantizer 204.
[0628] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predicted quantization parameter generator 204b, a quantization parameter storage 204d, and an inverse quantization executor 204e.
[0629] Figure 74 is a flowchart illustrating one example of an inverse quantization process performed by the inverse quantizer 204.
[0630] The inverse quantizer 204 can perform the inverse quantization process based on Figure 74 The flowchart illustrated in FIG. 13 performs the inverse quantization process for each CU as one example. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined to perform inverse quantization (Yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameter of the current block from the entropy decoder 202 (step Sv_12).
[0631] Next, the predicted quantization parameter generator 204b then obtains the quantization parameters of the processing units different from the current block from the quantization parameter storage 204d (step Sv_13). The predicted quantization parameter generator 204b generates the predicted quantization parameter of the current block based on the obtained quantization parameters (step Sv_14).
[0632] The quantization parameter generator 204a then generates the quantization parameter of the current block based on the delta quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the delta quantization parameter of the current block obtained from the entropy decoder 202 and the predicted quantization parameter of the current block generated by the predicted quantization parameter generator 204b can be added together to generate the quantization parameter of the current block. Further, the quantization parameter generator 204a stores the quantization parameter of the current block in the quantization parameter storage 204d (step Sv_16).
[0633] Next, the inverse quantization executor 204e inverse quantizes the quantized coefficients of the current block into transform coefficients using the quantization parameter generated in step Sv_15 (step Sv_17).
[0634] It should be noted that different quantization parameters can be decoded at the bit sequence level, the picture level, the slice level, the block level, or the CTU level. Further, an initial value of the quantization parameter can be decoded at the sequence level, the picture level, the slice level, the block level, or the CTU level. At this time, the initial value of the quantization parameter and the delta quantization parameter can be used to generate the quantization parameter.
[0635] It should be noted that the inverse quantizer 204 can include a plurality of inverse quantizers, and an inverse quantization method selected from a plurality of inverse quantization methods can be used to inverse quantize the quantized coefficients.
[0636] (Inverse transformer)
[0637] The inverse transformer 206 recovers the prediction residual by inverse transforming the transform coefficients input from the inverse quantizer 204.
[0638] For example, when the information parsed from the stream indicates that EMT or AMT is to be applied (for example, when the AMT flag is true), the inverse transformer 206 inverse transforms the transform coefficients of the current block based on the information indicating the parsed transform type.
[0639] Further, for example, when the information parsed from the stream indicates that NSST is to be applied, the inverse transformer 206 applies a quadratic inverse transform to the transform coefficients.
[0640] Figure 75 is a flowchart showing one example of a process performed by the inverse transformer 206.
[0641] For example, the inverse transformer 206 determines whether there is information in the stream indicating that no orthogonal transform was performed (step St ll). Here, when it is determined that there is no such information (NO in step St ll) (e.g., there is no indication as to whether an orthogonal transform was performed; there is an indication that an orthogonal transform is to be performed); the inverse transformer 206 obtains information indicating the transform type decoded by the entropy decoder 202 (step St 12). Next, based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step St 13). The inverse transformer 206 then performs inverse orthogonal transform using the determined transform type (step St 14). As shown in FIG. 8, when it is determined that there is information indicating that no orthogonal transform was performed (YES in step St ll) (e.g., an explicit indication that no orthogonal transform was performed; there is no indication that an orthogonal transform was performed), no orthogonal transform is performed. Figure 75
[0642] Figure 76 is a flowchart showing one example of the process performed by the inverse transformer 206.
[0643] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a determined value (step Su ll). The determined value can be predetermined. Here, when it is determined that the transform size is less than or equal to the determined value (YES in step Su ll), the inverse transformer 206 obtains information from the entropy decoder 202 indicating which transform type of at least one transform type included in the first transform type group was used by the encoder 100 (step Su 12). Note that this information is decoded by the entropy decoder 202 and output to the inverse transformer 206.
[0644] Based on this information, the inverse transformer 206 determines the transform type used for the orthogonal transform in the encoder 100 (step Su 13). The inverse transformer 206 then performs inverse orthogonal transform on the transform coefficients of the current block using the determined transform type (step Su 14). When it is determined that the transform size is not less than or equal to the determined value (NO in step Su ll), the inverse transformer 206 performs inverse transform on the transform coefficients of the current block using the second transform type group (step Su 15).
[0645] Note that, as one example, the transform size can be determined based on the size of the current block and the size of the block to which the current block is divided. Figure 75 or Figure 76 The illustrated process performs inverse orthogonal transformation of the inverse transformer 206 for each TU. In addition, the inverse orthogonal transformation can be performed by using a defined transform type without decoding information indicating a transform type used for the orthogonal transformation. The defined transform type can be a pre-defined transform type or a default transform type. In addition, the transform type can be specifically DST7, DCT8, or the like. In the inverse orthogonal transformation, an inverse transform basis function corresponding to the transform type is used.
[0646] (adder)
[0647] The adder 208 reconstructs the current block by adding the prediction residual input from the inverse transformer 206 and the prediction image input from the prediction controller 220. In other words, a reconstructed image of the current block is generated. The adder 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter 212.
[0648] (block memory)
[0649] The block memory 210 is a memory for storing blocks included in the current picture and which can be referred to in the intra prediction. More specifically, the block memory 210 stores the reconstructed image output from the adder 208.
[0650] (loop filter)
[0651] The loop filter 212 applies a loop filter to the reconstructed image generated by the adder 208 and outputs the filtered reconstructed image to the frame memory 214 and provides an output of the decoder 200, e.g., to a display device or the like.
[0652] When the information indicating the on or off of the ALF parsed from the stream indicates that the ALF is on, one filter is selected from a plurality of filters, e.g., based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0653] Figure 77 is a block diagram illustrating one example of a functional configuration of the loop filter 212. It should be noted that the configuration of the loop filter 212 is similar to that of the loop filter 120 of the encoder 100.
[0654] For example, as Figure 77 illustrated, the loop filter 212 includes a deblocking filter enforcer 212a, an SAO enforcer 212b, and an ALF enforcer 120c. The deblocking filter enforcer 212a performs deblocking filter processing on the reconstructed image. The SAO enforcer 212b performs the SAO process on the reconstructed image after the deblocking filter process. The ALF enforcer 212c performs the ALF process on the reconstructed image after the SAO process. It should be noted that the loop filter 212 does not always need to includeFigure 77 All the constituent elements disclosed in the above description can be included, and only a part of the constituent elements can be included. Further, the loop filter 212 can be configured to perform the above-described processes in a processing order different from the processing order disclosed in the above description, and can not perform Figure 77 All the processes and the like shown in the above description can be performed. Figure 77
[0655] (Frame memory)
[0656] The frame memory 214 is a memory for storing, for example, a reference picture used in inter prediction, and can also be referred to as a frame buffer. More specifically, the frame memory 214 stores a reconstructed image filtered by the loop filter 212.
[0657] (Predictor (intra predictor, inter predictor, prediction controller))
[0658] Figure 78 is a flowchart showing one example of a process performed by the predictor of the decoder 200. It should be noted that the prediction performer can include all or a part of the following constituent elements: the intra predictor 216; the inter predictor 218; and the prediction controller 220. The prediction performer includes, for example, the intra predictor 216 and the inter predictor 218.
[0659] The predictor generates a prediction image of the current block (step Sq_1). The prediction image can also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra prediction signal or an inter prediction signal. More specifically, the predictor generates the prediction image of the current block using a reconstructed image that has been obtained for another block, by prediction image generation, prediction residual recovery, and prediction image addition. The predictor of the decoder 200 generates the same prediction image as the prediction image generated by the predictor of the encoder 100. In other words, the prediction image is generated according to a method common to or corresponding to each other between the predictors.
[0660] The reconstructed image can be, for example, an image in a reference picture, or an image of a decoded block (i.e., the above-described other block) in a current picture (which is a picture including the current block). The decoded block in the current picture is, for example, a neighboring block of the current block.
[0661] Figure 79 is a flowchart showing another example of a process performed by the prediction performer of the decoder 200.
[0662] The predictor determines a method or mode for generating a prediction image (step Sr_1). The method or mode can be determined, for example, based on, for example, a prediction parameter or the like.
[0663] When the first method is determined as the mode of generating the prediction image, the predictor generates the prediction image according to the first method (step Sr_2a). When the second method is determined as the mode of generating the prediction image, the predictor generates the prediction image according to the second method (step Sr_2b). When the third method is determined as the mode of generating the prediction image, the predictor generates the prediction image according to the third method (step Sr_2c).
[0664] The first method, the second method, and the third method can be different methods from each other for generating the prediction image. Each of the first to third methods can be an inter prediction method, an intra prediction method, or another prediction method. The above-described reconstructed image can be used in these prediction methods.
[0665] Figure 80 is a flowchart illustrating another example of a process performed by the prediction of the decoder 200.
[0666] As one example, the predictor can perform the prediction process according to Figure 80 the flow illustrated in FIG. 13. Note that, Figure 80 Intra block copy illustrated in FIG. 12 is a mode belonging to inter prediction in which a block included in the current picture is referred to as a reference image or a reference block. In other words, in intra block copy, no picture other than the current picture is referred to. In addition, Figure 80 The PCM mode illustrated in FIG. 14 is a mode belonging to intra prediction and in which no transform and quantization are performed.
[0667] (Intra predictor)
[0668] The intra predictor 216 performs intra prediction based on an intra prediction mode parsed from the stream by referring to a block in the current picture stored in the block memory 210 to generate a prediction image of the current block (i.e., an intra predicted block). More specifically, the intra predictor 216 performs intra prediction by referring to pixel values (e.g., luma and / or chroma values) of one or more blocks adjacent to the current block to generate an intra predicted image, and then outputs the intra predicted image to the prediction controller 220.
[0669] It should be noted that when an intra prediction mode of referring to a luma block in intra prediction of a chroma block is selected, the intra predictor 216 can predict a chroma component of the current block based on a luma component of the current block.
[0670] Further, when information parsed from the stream indicates that PDPC is to be applied, the intra predictor 216 corrects intra predicted pixel values based on horizontal / vertical reference pixel gradients.
[0671] Figure 81 is a diagram illustrating one example of a process performed by the predictor 216 of the decoder 200.
[0672] The intra predictor 216 first determines whether to use MPM. Figure 81 As shown, the intra-frame predictor 216 determines whether an MPM flag indicating 1 is present in the stream (step Sw_11). Here, when it is determined that the MPM flag indicating 1 is present (yes in step Sw_11), the intra-frame predictor 216 obtains information indicating the intra-frame prediction mode selected in the encoder 100 from the entropy decoder 202. It should be noted that such information is decoded by the entropy decoder 202 and output to the intra-frame predictor 216. Next, the intra-frame predictor 216 determines the MPM (step Sw_3). The MPM includes, for example, six intra-frame prediction modes. The intra-frame predictor 216 then determines the intra-frame prediction mode included in the multiple intra-frame prediction modes included in the MPM and indicated by the information obtained in step Sw_12 (step Sw_14).
[0673] When it is determined that there is no MPM flag indicating 1 (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information indicating the intra prediction mode selected in the encoder 100 from at least one intra prediction mode not included in the MPM from the entropy decoder 202. It should be noted that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. The intra predictor 216 then determines the intra prediction mode that is not included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15 (step Sw_17).
[0674] The intra predictor 216 generates a predicted image according to the intra prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18 ).
[0675] (Inter-frame predictor)
[0676] The inter-frame predictor 218 predicts the current block by referencing the reference picture stored in the frame memory 214. Prediction is performed in units of the current block or the current subblock within the current block. It should be noted that a subblock is included in a block and is a smaller unit than a block. The size of a subblock can be 4×4 pixels, 8×8 pixels, or other sizes. The size of a subblock can be switched between slices, blocks, pictures, and other units.
[0677] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from the stream (e.g., prediction parameters output from the entropy decoder 202), and outputs the inter-frame prediction image to the prediction controller 220.
[0678] When the information parsed from the stream indicates that the OBMC mode is to be applied, the inter predictor 218 generates an inter prediction image using the motion information of the neighboring blocks in addition to the motion information of the current block obtained by the motion estimation.
[0679] Further, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter predictor 218 derives the motion information by performing the motion estimation according to the pattern matching method (e.g., bilateral matching or template matching) parsed from the stream. The inter predictor 218 then performs the motion compensation (prediction) using the derived motion information.
[0680] Further, when the BIO mode is to be applied, the inter predictor 218 derives the MV based on a model assuming a uniform linear motion. Further, when the information parsed from the stream indicates that the affine mode is to be applied, the inter predictor 218 derives the MV of each sub-block based on the MVs of a plurality of neighboring blocks.
[0681] (MV derivation flow)
[0682] Figure 82 is a flowchart illustrating one example of the MV derivation process in the decoder 200.
[0683] For example, the inter predictor 218 determines whether or not to decode the motion information (e.g., MV). For example, the inter predictor 218 can determine according to the prediction mode included in the stream, or can determine based on other information included in the stream. Here, when it is determined to decode the motion information, the inter predictor 218 derives the MV of the current block in a mode in which the motion information is decoded. When it is determined not to decode the motion information, the inter predictor 218 derives the MV in a mode in which no motion information is decoded.
[0684] Here, the MV derivation mode includes the normal inter mode, the normal merge mode, the FRUC mode, the affine mode, and the like described later. The mode in which the motion information is decoded includes the normal inter mode, the normal merge mode, the affine mode (specifically, the affine inter mode and the affine merge mode), and the like. It should be noted that the motion information can include not only the MV but also the MV predictor selection information described later. The mode in which no motion information is decoded includes the FRUC mode, and the like. The inter predictor 218 selects the mode for deriving the MV of the current block from among the plurality of modes, and derives the MV of the current block using the selected mode.
[0685] Figure 83 is a flowchart illustrating one example of the MV derivation process in the decoder 200.
[0686] For example, the inter predictor 218 can determine whether to decode the MV difference, i.e., for example, can determine in accordance with a prediction mode included in the stream, or can determine based on other information included in the stream. Here, when it is determined to decode the MV difference, the inter predictor 218 can derive the MV of the current block in a mode in which the MV difference is decoded. In this case, for example, the MV difference included in the stream is decoded as the prediction parameter.
[0687] When it is determined not to decode any MV difference, the inter predictor 218 derives the MV in a mode in which no MV difference is decoded. In this case, no encoded MV difference is included in the stream.
[0688] Here, as described above, the MV derivation mode includes a normal inter mode, a normal merge mode, an FRUC mode, an affine mode, and the like described later. The mode in which the MV difference is encoded includes the normal inter mode and the affine mode (specifically, affine inter mode), and the like. The mode in which no MV difference is encoded includes the FRUC mode, the normal merge mode, the affine mode (specifically, affine merge mode), and the like. The inter predictor 218 selects a mode for deriving the MV of the current block from among the plurality of modes, and derives the MV of the current block using the selected mode.
[0689] (MV derivation > normal inter mode)
[0690] For example, when information parsed from the stream indicates that the normal inter mode is to be applied, the inter predictor 218 derives the MV based on the information parsed from the stream and performs motion compensation (prediction) using the MV.
[0691] Figure 84 is a flowchart showing an example of an inter prediction process by the normal inter mode in the decoder 200.
[0692] The inter predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter predictor 218 obtains a plurality of MV candidates of the current block based on information such as MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sg_11). In other words, the inter predictor 218 generates an MV candidate list.
[0693] Next, the inter predictor 218 extracts N (N is an integer of 2 or more) MV candidates from the plurality of MV candidates obtained in step Sg_11 as motion vector prediction candidates (also referred to as MV prediction candidates) in accordance with the ranking in the priority order (step Sg_12). Note that the ranking in the priority order can be determined in advance for the respective N MV prediction candidates, and the ranking can be predetermined.
[0694] Next, the inter predictor 218 decodes the MV predictor selection information from the input stream, and selects one of the N MV predictor candidates as the MV predictor of the current block using the decoded MV predictor selection information (step Sg_13).
[0695] Next, the inter predictor 218 decodes the MV difference from the input stream, and derives the MV of the current block by adding the decoded MV difference to the selected MV predictor (step Sg_14).
[0696] Finally, the inter predictor 218 generates the prediction image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Sh_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed for each of all the blocks in a slice, the inter prediction of the slice using the normal inter mode is completed. For example, when the processes in steps Sg_11 to Sg_15 are performed for each of all the blocks in a picture, the inter prediction of the picture using the normal inter mode is completed. It should be noted that not all the blocks included in a slice can undergo the processes in steps Sh_11 to Sh_15, and the inter prediction of the slice using the normal inter mode can be completed when some of the blocks undergo the processes. The same applies to the picture in steps Sh_11 to Sh_15. The inter prediction of the picture using the normal inter mode can be completed when the processes are performed for some of the blocks in the picture.
[0697] (MV derivation > normal merge mode)
[0698] For example, when the information parsed from the stream indicates that the normal merge mode is to be applied, the inter predictor 218 derives the MV and performs motion compensation (prediction) using the MV.
[0699] Figure 85 is a flowchart showing an example of the inter prediction process by the normal merge mode in the decoder 200.
[0700] First, the inter predictor 218 obtains a plurality of MV candidates of the current block based on information such as MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sh_11). In other words, the inter predictor 218 generates an MV candidate list.
[0701] Next, the inter predictor 218 selects one of the plurality of MV candidates obtained in step Sh_11, thereby deriving the MV of the current block (step Sh_12). More specifically, the inter predictor 218 obtains MV selection information included in the stream as a prediction parameter, and selects the MV candidate identified by the MV selection information as the MV of the current block.
[0702] Finally, the inter predictor 218 generates a prediction image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed for each of all blocks in a slice, the inter prediction of the slice using the normal merge mode is completed. Further, when the processes in steps Sh_11 to Sh_13 are performed for each of all blocks in a picture, the inter prediction of the picture using the normal merge mode is completed. It should be noted that not all blocks included in a slice undergo these processes in steps Sh_11 to Sh_13, and the inter prediction of the slice using the normal merge mode can be completed when part of the blocks undergoes these processes. The same applies to the picture in steps Sh_11 to Sh_13. When these processes are performed for part of the blocks in a picture, the inter prediction of the picture using the normal merge mode can be completed.
[0703] (MV derivation > FRUC mode)
[0704] For example, when the information parsed from the stream indicates that the FRUC mode is to be applied, the inter predictor 218 derives the MV in the FRUC mode and performs motion compensation (prediction) using the MV. In this case, the motion information is derived at the decoder 200 side without being signaled from the encoder 100 side. For example, the decoder 200 can derive the motion information by performing motion estimation. In this case, the decoder 200 performs the motion estimation without using any pixel value in the current block.
[0705] Figure 86 is a flowchart showing an example of the inter prediction process by the FRUC mode in the decoder 200.
[0706] First, the inter predictor 218 indicates a list of MVs of decoded blocks spatially or temporally neighboring the current block by using the MVs as MV candidates (the list is an MV candidate list, and for example, can also be used as an MV candidate list for the normal merge mode) (step Si_11). Next, a best MV candidate is selected from among the plurality of MV candidates registered in the MV candidate list (step Si_12). For example, the inter predictor 218 calculates an evaluation value of each of the MV candidates included in the MV candidate list, and selects one of the MV candidates as the best MV candidate based on the evaluation values. Based on the selected best MV candidate, the inter predictor 218 then derives the MV of the current block (step Si_14). More specifically, for example, the selected best MV candidate is directly derived as the MV of the current block. In addition, for example, the MV of the current block can be derived using pattern matching in a certain position surrounding region, which is included in the reference picture and corresponds to the selected best MV candidate. In other words, estimation using pattern matching and evaluation values in the reference picture can be performed in the surrounding region of the best MV candidate, and when there is an MV that produces a better evaluation value, the best MV candidate can be updated to the MV that produces the better evaluation value, and the updated MV can be determined as the final MV of the current block. In an embodiment, the update of the MV that produces the better evaluation value can not be performed.
[0707] Finally, the inter predictor 218 generates a prediction image of the current block by performing motion compensation of the current block using the derived MV and the decoded reference picture (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each of all the blocks in a slice, the inter prediction of the slice using the FRUC mode is completed. For example, when the processes in steps Si_11 to Si_15 are performed for each of all the blocks in a picture, the inter prediction of the picture using the FRUC mode is completed. Each sub-block can be processed in a similar manner to the case of each block.
[0708] (MV derivation > FRUC mode)
[0709] For example, when the information parsed from the stream indicates that the affine merge mode is to be applied, the inter predictor 218 derives the MV in the affine merge mode and performs motion compensation (prediction) using the MV.
[0710] Figure 87 is a flowchart showing an example of the inter prediction process by the affine merge mode in the decoder 200.
[0711] In the affine merge mode, first, the inter predictor 218 derives MVs at respective control points of the current block (step Sk_11). The control points are a left-top corner point of the current block and a right-top corner point of the current block, as shown in FIG. 12.Figure 46A or a top-left corner point of the current block, a top-right corner point of the current block, and a bottom-left corner point of the current block, as Figure 46B indicated.
[0712] For example, when the MV derivation method illustrated in FIG. 6 is used, the inter predictor 218 checks the decoded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) in this order and identifies the first valid block decoded according to the affine mode, as Figures 47A-47C indicated. Figure 47A For example, when the MV derivation method illustrated in FIG. 6 is used, the inter predictor 218 checks the decoded blocks A (left), B (top), C (top-right), D (bottom-left), and E (top-left) in this order and identifies the first valid block decoded according to the affine mode, as Figure 47B indicated.
[0713] It should be noted that, as Figure 49A indicated, when the block A is identified and the block A has two control points, the MVs at three control points can be calculated, and as Figure 49B indicated, when the block A is identified and when the block A has three control points, the MVs at two control points can be calculated.
[0714] Further, when the MV selection information is included in the stream as the prediction parameter, the inter predictor 218 can use the MV selection information to derive the MVs at each control point of the current block.
[0715] Next, the inter predictor 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter predictor 218 calculates the MVs of each of the plurality of sub-blocks as affine MVs using the two motion vectors v0 and v1 and the above expression (1A) or the three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_12). The inter predictor 218 then performs motion compensation of the sub-blocks using these affine MVs and the decoded reference picture (step Sk_13). When the processes in steps Sk_12 and Sk_13 are performed for each of the sub-blocks included in the current block, the inter prediction using the affine merge mode for the current block ends. In other words, the motion compensation of the current block is performed to generate a prediction image of the current block.
[0716] It should be noted that the above MV candidate list can be generated in step Sk_11. The MV candidate list can be, for example, a list including MV candidates derived using a plurality of MV derivation methods for each control point. The plurality of MV derivation methods can be, for example,Figures 47A-47C the MV derivation method shown in Figure 48A and 48B the MV derivation method shown in Figure 49A and 49B any combination of the MV derivation method shown in
[0717] It should be noted that, in addition to the affine mode, the MV candidate list can include MV candidates in modes that perform prediction in sub-block units.
[0718] It should be noted that, for example, the MV candidate list including MV candidates can be generated as the MV candidate list in the affine merge mode using two control points and the affine merge mode using three control points. Alternatively, the MV candidate list including MV candidates in the affine merge mode using two control points and the MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, the MV candidate list including MV candidates can be generated in one of the affine merge mode using two control points and the affine merge mode using three control points.
[0719] (MV derivation > affine inter mode)
[0720] For example, when information parsed from the stream indicates that the affine inter mode is to be applied, the inter predictor 218 derives the MV in the affine inter mode and performs motion compensation (prediction) using the MV.
[0721] Figure 88 is a flowchart showing an example of the inter prediction process by the affine inter mode in the decoder 200.
[0722] In the affine inter mode, first, the inter predictor 218 derives MV predictors (v0, v1) or (v0, v1, v2) of the respective two or three control points of the current block (step Sj_11). As shown in Figure 46A or Figure 46B The control points are the top-left corner point of the current block, the top-right corner point of the current block, and the bottom-left corner point of the current block.
[0723] The inter predictor 218 obtains MV predictor selection information included in the stream as a prediction parameter, and derives the MV predictor at each control point of the current block using the MV identified by the MV predictor selection information. For example, when the MV derivation method shown in Figure 48A and Figure 48B Figure 48A or Figure 48B The MV of the block identified by the MV predictor selection information is selected from among the decoded blocks in the vicinity of the respective control points of the current block. The MV of the block identified by the MV predictor selection information is selected from among the decoded blocks in the vicinity of the respective control points of the current block. The MV of the block identified by the MV predictor selection information is selected from among the decoded blocks in the vicinity of the respective control points of the current block.
[0724] Next, the inter predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block and the MV difference corresponding to the MV predictor (step Sj_12). This derives the MV of the current block at each control point.
[0725] Next, the inter predictor 218 performs motion compensation on each of the plurality of sub-blocks included in the current block. In other words, the inter predictor 218 calculates the MV of each of the plurality of sub-blocks as an affine MV using the two motion vectors v0 and v1 and the above expression (1A) or the three motion vectors v0, v1, and v2 and the above expression (1B) (step Sj_13). The inter predictor 218 then performs motion compensation of the sub-blocks using these affine MVs and the decoded reference picture (step Sj_14). The inter prediction using the affine merge mode for the current block ends when the processes in steps Sj_13 and Sj_14 are performed for each of the sub-blocks included in the current block. In other words, motion compensation of the current block is performed to generate a predicted picture of the current block.
[0726] It should be noted that the above MV candidate list can be generated in step Sj_11 as in step Sk_11.
[0727] (MV derivation > triangle mode)
[0728] For example, when the information parsed from the stream indicates that the triangle mode is to be applied, the inter predictor 218 derives the MV in the triangle mode and performs motion compensation (prediction) using the MV.
[0729] Figure 89 is a flowchart showing an example of the process of inter prediction by the affine triangle mode in the decoder 200.
[0730] In the triangle mode, first, the inter predictor 218 splits the current block into a first partition and a second partition (step Sx_11). For example, the inter predictor 218 can obtain partition information as a prediction parameter from the stream, which is information related to the splitting. The inter predictor 218 can then split the current block into the first partition and the second partition according to the partition information.
[0731] Next, the inter predictor 218 obtains a plurality of MV candidates of the current block based on information such as MVs of a plurality of decoded blocks temporally or spatially surrounding the current block (step Sx_12). In other words, the inter predictor 218 generates a list of MV candidates.
[0732] The inter predictor 218 then selects a MV candidate of the first partition and a MV candidate of the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, the inter predictor 218 can obtain MV selection information for identifying each of the selected MV candidates as a prediction parameter from the stream. The inter predictor 218 can then select the first MV and the second MV in accordance with the MV selection information.
[0733] Next, the inter predictor 218 generates a first prediction image by performing motion compensation using the selected first MV and a decoded reference picture (step Sx_14). Likewise, the inter predictor 218 generates a second prediction image by performing motion compensation using the selected second MV and a decoded reference picture (step Sx_15).
[0734] Finally, the inter predictor 218 generates a prediction image of the current block by performing weighted addition of the first prediction image and the second prediction image (step Sx_16).
[0735] (MV estimation > DMVR)
[0736] For example, information parsed from the stream indicates that DMVR is to be applied, and the inter predictor 218 performs motion estimation using DMVR.
[0737] Figure 90 is a flowchart illustrating an example of a process of motion estimation by DMVR in the decoder 200.
[0738] The inter predictor 218 derives a MV of the current block in accordance with the merge mode (step S1_11). Next, the inter predictor 218 derives a final MV of the current block by searching a region around a reference picture indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined in accordance with DMVR.
[0739] Figure 91 is a flowchart illustrating an example of a process of motion estimation by DMVR in the decoder 200, and is Figure 58B identical to
[0740] First, in Figure 58AIn Step 1 shown, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 changes the target to the search position at which the minimum cost is obtained, and executes the process in Step 2. Figure 58A The process in Step 2 shown. When the cost at the starting point is the minimum, the inter-frame predictor 218 skips the process in Step 2 Figure 58A The process in Step 2 shown and executes the process in Step 3.
[0741] In Step 2 shown, the inter-frame predictor 218 executes a search similar to the process in Step 1, and changes the target to the search position after the target is changed according to the result of the process in Step 1 as a new starting point. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is the minimum. Here, when it is determined that the cost at one of the search positions other than the starting point is the minimum, the inter-frame predictor 218 executes the process in Step 4. When the cost at the starting point is the minimum, the inter-frame predictor 218 executes the process in Step 3. Figure 58A In Step 4, the inter-frame predictor 218 regards the search position at the starting point as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0742] In Step 3 shown, the inter-frame predictor 218 determines a pixel position at sub-pixel accuracy based on the costs at the four points at the upper, lower, left, and right positions with respect to the starting point in Step 1 or Step 2 to obtain the minimum cost, and regards the pixel position as the final search position.
[0743] Figure 58A The pixel position at sub-pixel accuracy is determined by weighted addition of each of the four vectors ((0, 1), (0, -1), (-1, 0), and (1, 0)) using the cost at the corresponding position of the four search positions as the weight. The inter-frame predictor 218 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.
[0744] (Motion compensation > BIO / OBMC / LIC)
[0745] For example, when the information parsed from the stream indicates that correction of the prediction image is to be performed, the inter-frame predictor 218 corrects the prediction image based on a correction mode at the time of generation of the prediction image. The mode is, for example, one of BIO, OBMC, and LIC described above.
[0746]
[0747] Figure 92 is a flowchart showing one example of a process of generating a prediction image in the decoder 200.
[0748] The inter predictor 218 generates a prediction image (step Sm_11), and corrects the prediction image according to any of the above-described modes (step Sm_12).
[0749] Figure 93 is a flowchart showing another example of a process of generating a prediction image in the decoder 200.
[0750] The inter predictor 218 derives the MV of the current block (step Sn_11). Next, the inter predictor 218 generates a prediction image using the MV (step Sn_12), and determines whether to perform a correction process (step Sn_13). For example, the inter predictor 218 obtains a prediction parameter included in the stream, and determines whether to perform the correction process based on the prediction parameter. For example, the prediction parameter is a flag indicating whether one or more of the above-described modes are to be applied. Here, when it is determined to perform the correction process (Yes in step Sn_13), the inter predictor 218 generates a final prediction image by correcting the prediction image (step Sn_14). Note that, in LIC, the correction can be performed on the luminance and the chrominance in step Sn_14. When it is determined not to perform the correction process (No in step Sn_13), the inter predictor 218 outputs the final prediction image without correcting the prediction image (step Sn_15).
[0751] (Motion compensation > OBMC)
[0752] For example, when information parsed from the stream indicates that OBMC is to be performed, the inter predictor 218 corrects the prediction image according to OBMC when generating the prediction image.
[0753] Figure 94 is a flowchart showing an example of a process of correction of a prediction image by OBMC in the decoder 200. Note that, Figure 94 The flowchart in Figure 62 shows a correction flow of a prediction image using the current picture and the reference picture shown in
[0754] First, as shown in Figure 62 , the inter predictor 218 obtains a prediction image (Pred) by normal motion compensation using the MV assigned to the current block.
[0755] Next, the inter predictor 218 obtains a prediction picture (Pred_L) by applying a motion vector (MV_L) that has been derived for a decoded block neighboring the left side of the current block to the current block (reusing the motion vector of the current block). The inter predictor 218 then performs a first correction of the prediction picture by overlaying the two prediction pictures Pred and Pred_L. This provides an effect of blending the boundary between the neighboring blocks.
[0756] Likewise, the inter predictor 218 obtains a prediction picture (Pred_U) by applying a motion vector (MV_U) that has been derived for a decoded block neighboring the top of the current block to the current block (reusing the motion vector of the current block). The inter predictor 218 then performs a second correction of the prediction picture by overlaying the prediction picture Pred_U to the prediction picture (e.g., Pred and Pred_L) to which the first correction has been performed. This provides an effect of blending the boundary between the neighboring blocks. The prediction picture resulting from the second correction is a picture in which the boundary between the neighboring blocks has been blended (smoothed), and is thus the final prediction picture for the current block.
[0757] (Motion compensation > BIO)
[0758] For example, when information parsed from the stream indicates that BIO is to be performed, the inter predictor 218 corrects the prediction picture according to the BIO when generating the prediction picture.
[0759] Figure 95 is a flowchart showing an example of the process of correction of the prediction picture by BIO in the decoder 200.
[0760] As Figure 63 indicated, the inter predictor 218 derives two motion vectors (M0, M1) using two reference pictures (Ref0, Ref1) different from the picture (Cur Pic) including the current block. The inter predictor 218 then derives a prediction picture for the current block using the two motion vectors (M0, M1) (step Sy_11). It should be noted that the motion vector M0 is a motion vector (MVx0, MVy0) corresponding to the reference picture Ref0, and the motion vector M1 is a motion vector (MVx1, MVy1) corresponding to the reference picture Ref1.
[0761] Next, the inter predictor 218 derives a difference image I0 of the current block using the motion vector M0 and the reference picture L0. Also, the inter predictor 218 derives a difference image I1 of the current block using the motion vector M1 and the reference picture L1 (step Sy_12). Here, the difference image I0 is an image included in the reference picture Ref0 and to be derived for the current block, and the difference image I1 is an image included in the reference picture Ref1 and to be derived for the current block. Each of the difference image I0 and the difference image I1 can be the same size as the current block. Alternatively, each of the difference image I0 and the difference image I1 can be an image larger than the current block. Also, the difference image I0 and the difference image I1 can include a predicted image obtained by using the motion vectors (M0, M1) and the reference pictures (L0, L1) and applying a motion compensation filter.
[0762] Also, the inter predictor 218 derives gradient images (Ix0, Ix1, Iy0, Iy1) of the current block from the difference image I0 and the difference image I1 (step Sy_13). Note that the gradient images in the horizontal direction are (Ix0, Ix1), and the gradient images in the vertical direction are (Iy0, Iy1). The inter predictor 218 can derive the gradient images by, for example, applying a gradient filter to the difference images. The gradient images can be images each indicating an amount of spatial variation of pixel values in the horizontal direction or an amount of spatial variation of pixel values in the vertical direction.
[0763] Next, the inter predictor 218 derives, for each sub-block of the current block, an optical flow (vx, vy) that is a velocity vector using the interpolated images (I0, I1) and the gradient images (Ix0, Ix1, Iy0, Iy1) (step Sy_14). As one example, the sub-block can be a 4x4-pixel sub-CU.
[0764] Next, the inter predictor 218 corrects the predicted image of the current block using the optical flow (vx, vy). For example, the inter predictor 218 derives correction values of pixel values included in the current block using the optical flow (vx, vy) (step Sy_15). The inter predictor 218 can then correct the predicted image of the current block using the correction values (step Sy_16). Note that the correction values can be derived in units of pixels, in units of multiple pixels, or in units of sub-blocks, and so on.
[0765] Note that the BIO process flow is not limited to the process disclosed in Figure 95 . Only part of the process disclosed in Figure 95 may be executed, or different processes can be added or used instead, or the processes can be executed in a different processing order.
[0766] (Motion compensation > LIC)
[0767] For example, when the information parsed from the stream indicates that LIC is to be performed, in generating the prediction image, the inter-predictor 218 corrects the prediction image according to the LIC.
[0768] Figure 96 is a flowchart illustrating an example of a process of correction of a prediction image by LIC in the decoder 200.
[0769] First, the inter-predictor 218 obtains a reference image corresponding to the current block from a decoded reference picture using the MV (step Sz_11).
[0770] Next, the inter-predictor 218 extracts, for the current block, information indicating how luminance values change between the current picture and the reference picture (step Sz_12). This extraction can be performed based on luminance pixel values of a decoded left neighboring reference region (surrounding reference region) and a decoded upper neighboring reference region (surrounding reference region), and luminance pixel values at corresponding positions in the reference picture specified by the derived MV. The inter-predictor 218 uses the information indicating how the luminance values change to calculate luminance correction parameters (step Sz_13).
[0771] The inter-predictor 218 generates the prediction image of the current block by performing a luminance correction process in which the luminance correction parameters are applied to the reference image in the reference picture specified by the MV (step Sz_14). In other words, the prediction image, which is the reference image in the reference picture specified by the MV, is corrected based on the luminance correction parameters. In this correction, luminance can be corrected, or chrominance can be corrected.
[0772] (Prediction controller)
[0773] The prediction controller 220 selects either the intra-predicted image or the inter-predicted image, and outputs the selected image to the adder 208. Overall, the configuration, functions, and processes of the prediction controller 220, the intra-predictor 216, and the inter-predictor 218 on the decoder 200 side can correspond to the configuration, functions, and processes of the prediction controller 128, the intra-predictor 124, and the inter-predictor 218 on the encoder 100 side.
[0774] (Decoding using predicted chroma samples)
[0775] In a first aspect, it is determined whether a block of chroma samples of a current block can be predicted using luminance samples, where the block is decoded using the predicted chroma samples.
[0776] Figure 97 is a flowchart illustrating one example of a process 1000 of decoding a block using predicted chroma samples, which can be performed by, for example,Figure 7 encoder 100 or Figure 67 decoder 200 performs. For convenience, the decoding of Figure 67 decoder 200 will be described. Figure 97
[0777] At S1001, the decoder 200 determines whether the current chroma block is within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid. Figure 98 and Figure 99 are conceptual diagrams illustrating examples of determining whether a current chroma block is within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid. In certain formats, such as the YUV420 format, a 16x16 pixel chroma region corresponds to a 32x32 pixel luma region. As shown in Figure 98 and Figure 99 , a chroma block within a 32x32 luma region that is aligned with a 16x16 chroma grid is determined to be within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid. Chroma blocks that are not within the 32x32 luma region are not determined to be within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid. As shown in Figure 99 , the chroma samples of the illustrated chroma block can be predicted using luma samples of the chroma block because the chroma block is included in a grid (as shown, a 16x16 grid) and the collocated luma block is also within the collocated 32x32 region.
[0778] In some embodiments, the chroma samples of a block that is not determined to be within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid can not be predicted using luma samples, while the chroma samples of a block that is determined to be within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid can be predicted using luma samples, e.g., by default, when other conditions are met (as discussed below with reference to S1002).
[0779] As shown in Figure 97 , when the current chroma block is not determined to be within an MxN non-overlapping region that is aligned with the chroma samples of an MxN grid at S1001, the process 1000 proceeds from S1001 to S1004, where the decoder 200 predicts the chroma samples of the block without using luma samples. The process 1000 proceeds from S1004 to S1005, where the decoder 200 decodes the block using the predicted chroma samples. When the current chroma block is determined to be within an MxN non-overlapping region at S1001, the process 1000 proceeds from S1001 to S1002.
[0780] At S1002, the decoder 200 determines whether to split the current luma VPDU into smaller blocks. Whether to split the current luma VPDU into smaller blocks can be determined in various ways, discussed below with reference toFigure 102 and Figure 103 Some examples are discussed in more detail. When it is not determined at S1002 that the current luma VPDU will be split into smaller blocks, the process 1000 proceeds from S1002 to S1004, where the decoder 200 predicts the chroma samples of the block without using the luma samples. The process 1000 proceeds from S1004 to S1005, where the decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1002 that the current luma VPDU will be split into smaller blocks, the process 1000 proceeds from S1002 to S1003, where the decoder 200 predicts the chroma samples of the block using the luma samples. The process 1000 proceeds from S1003 to S1005, where the decoder 200 decodes the block using the predicted chroma samples. In some embodiments, additional considerations may be taken into account to determine whether to use luma samples to decode the chroma samples of the block, for example, as described below with reference to Figure 103 discussed.
[0781] Figure 100 A conceptual diagram illustrating a VPDU is shown in FIG. VPDU is non-overlapping and represents the buffer size of a pipeline stage. Figure 100 The left side of FIG (labeled a) shows an example of a 128x128 CTU with four 64x64 VPDUs. Figure 100 The right side of FIG (labeled b) shows an example of a 128x128 CTU with 16 32x32 VPDUs.
[0782] Figure 101 This conceptual diagram illustrates an example of determining whether the current VPDU can use luma samples to predict blocks of chroma samples based on whether the luma VPDU is split into blocks. The left side shows a luma CTU, and the right side shows the corresponding chroma CTU. As shown in the figure, luma VPDU0 is split into blocks, while luma VPDU1 is not. Therefore, luma samples can be used to predict chroma samples in VPDU0, while luma samples are not used to predict chroma samples in VPDU1.
[0783] Figure 102 is a conceptual diagram illustrating two example ways of determining whether a luma VPDU is to be split into smaller blocks. Figure 102In a first example shown on the left (labeled a), a determination can be made as to whether the luma VPDU is to be split based on a split flag associated with the luma VPDU. As shown, when the value of the split flag is 1, the VPDU is to be split (and luma samples can be used to predict chroma samples of the block). When the value of the split flag is 0, the VPDU is not split (and luma samples can not be used to predict chroma samples of the block). Other split flag values can be employed to determine whether the luma VPDU can be split.
[0784] In Figure 102 In a second example shown on the right (labeled b), a determination can be made as to whether the luma VPDU is to be split based on a quadtree split depth of the luma block of the VPDU. As shown, the quadtree split depth of the luma block of VPDUO is greater than 1, and thus luma samples can be used to predict chroma samples when decoding the block of VPDUO. In contrast, the quadtree split depth of the block of VPDUl is less than or equal to 1, and thus luma samples can not be used to predict chroma samples when decoding the block of VPDUO. Other split depth values can be employed to determine whether the luma VPDU can be split.
[0785] Figure 103 is a conceptual diagram illustrating additional considerations that can be taken into account to determine whether to use luma samples to predict chroma samples of a block. As shown, whether the current block size is equal to or less than a threshold block size can be used as an additional consideration in determining whether to use luma samples to predict chroma samples of a block.
[0786] The threshold block size can be a default block size, a signaled block size, or a determined block size, and can be a luma or chroma block size. For example, if the threshold block size is a 16x16 luma block size, the luma block size of VPDUO is greater than 16x16, and thus it can be determined that luma samples are not to be used to determine chroma samples of the block. The threshold block size can be employed at S1002 to determine whether the current luma VPDU is to be split into smaller blocks.
[0787] Aspects of the process 1000 of Figure 97 may be modified in various ways. For example, the process 1000 can be modified to perform more actions than shown, can be modified to perform fewer actions than shown, can be modified to perform actions in various orders, can be modified to combine or split actions, etc. For example, the process 1000 can be modified to determine whether to use luma samples to predict chroma samples of a block based on other considerations (e.g., the size of the current block as discussed with reference to Figure 103 In another example, the process 1000 can be modified to omit S1001.
[0788] The blocks described in each aspect can be replaced with rectangular or non-rectangular shaped partitions. Figure 104 Examples of non-rectangular shaped partitions are shown, such as triangular partitions, L-shaped partitions, pentagonal partitions, hexagonal partitions, and polygonal partitions. Other non-rectangular shaped partitions can be employed, and combinations of various shapes can be employed. The term "partition" described in each aspect can be replaced with the term "prediction unit." The term "partition" described in each aspect can also be replaced with the term "sub-prediction unit." The term "partition" described in each aspect can also be replaced with the term "coding unit."
[0789] Among other advantages, determining whether luma samples can be used to predict chroma samples for decoding a current block helps reduce reconstruction latency and increase flexibility of hardware implementation.
[0790] One or more aspects disclosed herein can be performed in combination with at least a portion of other aspects in the disclosure. Also, one or more aspects disclosed herein can be performed by combining a portion of a process indicated in any flowchart according to the aspects, a portion of a configuration of any device, a portion of syntax, etc. with other aspects. Aspects described with reference to constituent elements of an encoder can be similarly performed by corresponding constituent elements of a decoder.
[0791] (Implementation and Application)
[0792] As described in each of the above-described embodiments, for example, each functional block or operation block can be realized as an MPU (Micro Processing Unit) and a memory. Furthermore, a process performed by each functional block can be realized as a program execution unit, such as a processor that reads and executes software (a program) recorded on a recording medium such as a ROM. The software can be distributed. The software can be recorded on various recording media, such as a semiconductor memory. It should be noted that each functional block can also be realized as hardware (a dedicated circuit). Various combinations of hardware and software can be employed.
[0793] The processes described in each embodiment can be realized by integrated processing using a single device (system), or can be realized by distributed processing using a plurality of devices. Furthermore, the processor that executes the above-described program can be a single processor or a plurality of processors. In other words, integrated processing can be performed, or distributed processing can be performed.
[0794] Embodiments of the disclosure are not limited to the above-described exemplary embodiments; various modifications can be made to the exemplary embodiments, and the results thereof are also included in the scope of the embodiments of the disclosure.
[0795] Next, application examples of the moving image encoding method (image encoding method) and the moving image decoding method (image decoding method) described in each of the above-described embodiments will be described, as well as various systems that implement these application examples. The feature of such a system can be an image encoder that employs the image encoding method, an image decoder that employs the image decoding method, or an image encoder-decoder that includes both the image encoder and the image decoder. Other configurations of such a system can be modified as appropriate.
[0796] (Usage Examples)
[0797] Figure 105 The overall configuration of a content providing system ex100 suitable for implementing a content distribution service is shown. Areas in which communication services are provided are divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the example shown, are located in the respective cells.
[0798] In the content providing system ex100, devices including a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via the Internet service provider ex102 or the communication network ex104 and the base stations ex106 to ex110. The content providing system ex100 can combine and connect any combination of the above-described devices. In various embodiments, the devices can be connected together directly or indirectly via a telephone network or near field communication, rather than via the base stations ex106 to ex110. Further, the streaming server ex103 can be connected to the devices including the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via, for example, the Internet ex101. The streaming server ex103 can also be connected to terminals in a hotspot in, for example, an airplane ex117 via a satellite ex116.
[0799] Note that, instead of the base stations ex106 to ex110, wireless access points or hotspots can be used. The streaming server ex103 can be directly connected to the communication network ex104 without passing through the Internet ex101 or the Internet service provider ex102, and can be directly connected to the airplane ex117 without passing through the satellite ex116.
[0800] The camera ex113 can be a device capable of capturing still images and videos, such as a digital camera. The smartphone ex115 can be a smartphone device, a cellular phone, or a Personal Handyphone System (PHS) phone, and can operate under 2G, 3G, 3.9G, and 4G system mobile communication system standards, as well as the next-generation 5G system.
[0801] The home appliance ex114 is, for example, a refrigerator or an apparatus included in a domestic fuel cell cogeneration system.
[0802] In the content providing system ex100, a terminal including an image and / or video capturing function can perform real-time streaming, for example, by connecting to the streaming server ex103 via, for example, the base station ex106. At the time of real-time streaming, the terminal (for example, the computer ex111, the game device ex112, the camera ex113, the home appliance ex114, the smart phone ex115, and the terminal in the airplane ex117) can perform the encoding process described in the above-described embodiments on still image or video content captured by the user via the terminal, can multiplex video data obtained via the encoding and audio data obtained by encoding audio corresponding to the video, and can transmit the obtained data to the streaming server ex103. In other words, the terminal functions as an image encoder according to an aspect of the present disclosure.
[0803] The streaming server ex103 streams content data to clients that request streaming. Examples of the clients include the computer ex111, the game device ex112, the camera ex113, the home appliance ex114, the smart phone ex115, and the terminal ex117 inside the airplane, which can decode the above-described encoded data. The devices that receive the streaming data can decode and reproduce the received data. In other words, the devices can each function as an image decoder according to an aspect of the present disclosure.
[0804] (Distributed processing)
[0805] The streaming server ex103 can be implemented as a plurality of servers or computers that divide tasks such as processing, recording, and streaming of data among them. For example, the streaming server ex103 can be implemented as a content delivery network (CDN) that streams content via a network connecting a plurality of edge servers located around the world. In the CDN, edge servers that are physically close to clients can be dynamically assigned to the clients. Content is cached and streamed to the edge servers to reduce loading time. For example, when some type of error or connection change occurs, for example, due to a surge in traffic, data can be stably streamed at high speed because, for example, an affected part of the network can be avoided by dividing processing among a plurality of edge servers, or the streaming task is switched to a different edge server and streaming is continued.
[0806] The distribution is not limited to division of processing for streaming; encoding of captured data can be divided and performed among terminals, on the server side, or both. In one example, in typical encoding, processing is performed in two loops. The first loop is used to detect the complexity of an image frame by frame or scene by scene, or to detect the encoding load. The second loop is used for processing that maintains image quality and improves encoding efficiency. For example, by having a terminal perform the first encoding loop and having the server side that receives content perform the second encoding loop, the processing load of the terminal can be reduced and the quality and encoding efficiency of the content can be improved. In this case, upon receiving a decoding request, the encoded data produced by the first loop performed by one terminal can be received and reproduced on another terminal in near real time. This makes it possible to achieve smooth real-time streaming.
[0807] In another example, a camera ex113 or the like extracts a feature amount (feature amount or characteristic amount) from an image, compresses data related to the feature amount as metadata, and transmits the compressed metadata to a server. For example, the server determines the importance of an object based on the feature amount and changes the quantization accuracy accordingly to perform compression that is appropriate to the image meaning (or content importance). In the second compression process performed by the server, the feature amount data is particularly effective in improving the accuracy and efficiency of motion vector prediction. Furthermore, encoding with a relatively low processing load, such as variable length coding (VLC), can be processed by the terminal, while encoding with a relatively high processing load, such as context adaptive binary arithmetic coding (CABAC), can be processed by the server.
[0808] In yet another example, there is a case where a plurality of terminals in, for example, a stadium, a shopping mall, or a factory capture a plurality of videos of substantially the same scene. In this case, for example, encoding can be distributed by dividing processing tasks by unit among the plurality of terminals that capture videos and, if necessary, among other terminals that do not capture videos and the server. The unit can be, for example, a group of pictures (GOP), a picture, or a tile produced by dividing a picture. This makes it possible to reduce the loading time and achieve streaming closer to real time.
[0809] Since the videos have substantially the same scene, they can be managed and / or instructed by the server so that the videos captured by the terminals can refer to each other. Furthermore, the server can receive encoded data from the terminals, change the reference relationship between data items, or correct or replace pictures on its own, and then perform encoding. This makes it possible to generate a stream with higher quality and efficiency for individual data items.
[0810] Furthermore, the server can stream video data after performing transcoding to convert the encoding format of the video data. For example, the server can convert the encoding format from MPEG to VP (e.g., VP9), can convert H.264 to H.265, and the like.
[0811] In this way, the encoding can be performed by the terminal or one or more servers. Thus, although the device performing the encoding is referred to as a "server" or a "terminal" in the following description, part or all of the processes performed by the server can be performed by the terminal, and likewise, part or all of the processes performed by the terminal can be performed by the server. This also applies to the decoding processes.
[0812] (3D, multi-angle)
[0813] The use of images or videos that are composed of images or videos of different scenes that are simultaneously captured by a plurality of terminals such as a camera ex113 and / or a smartphone ex115, or images or videos of the same scene that are captured from different angles is increasing. The videos captured by the terminals can be composed based on, for example, the relative positional relationship between the terminals that are obtained individually, or regions having matching feature points in the videos.
[0814] In addition to the encoding of two-dimensional moving pictures, the server can encode still images based on the scene analysis of the moving pictures (for example, automatically or at a user-specified time point) and transmit the encoded still images to the receiving terminal. Furthermore, when the server is able to obtain the relative positional relationship between the video capture terminals, in addition to the two-dimensional moving pictures, the server can generate a three-dimensional geometry of the scene based on the videos of the same scene that are captured from different angles. The server can separately encode the three-dimensional data generated from, for example, a point cloud, and based on the results of recognizing or tracking a person or an object using the three-dimensional data, it can select or reconstruct and generate a video to be transmitted to the receiving terminal from the videos captured by the plurality of terminals.
[0815] This allows the user to enjoy the scene by freely selecting a video corresponding to a video capture terminal, and allows the user to enjoy the content by extracting a video at a selected viewpoint from three-dimensional data reconstructed from a plurality of images or videos. Furthermore, for the video, sound can be recorded from relatively different angles, and the server can multiplex audio from a specific angle or space with the corresponding video, and transmit the multiplexed video and audio.
[0816] In recent years, content that is compounded from the real world and the virtual world, such as virtual reality (VR) and augmented reality (AR) content, has also become popular. In the case of a VR image, the server can create an image from two viewpoints of the left eye and the right eye, and perform encoding that allows referencing between the two viewpoint images, such as multi-view encoding (MVC), or can also encode the images as separate streams without referencing. When the images are decoded as separate streams, the streams can be synchronized at the time of reproduction, thereby reconstructing a virtual three-dimensional space according to the viewpoint of the user.
[0817] In the case of an AR image, the server can superimpose virtual object information existing in a virtual space onto camera information representing a real world space, for example, based on a three-dimensional position or motion from the user's perspective. The decoder can obtain or store virtual object information and three-dimensional data, generate a two-dimensional image from the user's perspective of the motion, and then generate superimposition data through seamless connection of the images. Alternatively, in addition to a request for virtual object information, the decoder can transmit motion from the user's perspective to the server. The server can generate superimposition data based on three-dimensional data stored ...
Claims
1. An encoder, comprising: Circuit; as well as a memory coupled to the circuit; The circuit performs the following operations during operation: determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks, wherein a chroma virtual pipeline decoding unit corresponding to the current luma virtual pipeline decoding unit is split into smaller blocks; In response to determining not to split the current luma virtual pipeline decoding unit into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining to split the luma virtual pipeline decoding unit into smaller blocks, using luma samples to predict blocks of chroma samples; and encoding the block using the predicted chroma samples, The circuit determines whether to split the current luma virtual pipeline decoding unit into smaller blocks based on a split flag, a quadtree split depth, or a threshold block size during operation.
2. A decoder comprising: Circuit; a memory coupled to the circuit; The circuit performs the following operations during operation: determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks, wherein a chroma virtual pipeline decoding unit corresponding to the current luma virtual pipeline decoding unit is split into smaller blocks; In response to determining not to split the current luma virtual pipeline decoding unit into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining to split the luma virtual pipeline decoding unit into smaller blocks, using luma samples to predict blocks of chroma samples; and decoding the block using the predicted chroma samples, The circuit determines whether to split the current luma virtual pipeline decoding unit into smaller blocks based on a split flag, a quadtree split depth, or a threshold block size during operation.
3. A non-transitory medium that stores a bit stream and can be read by a computer, The bitstream includes syntax for causing the computer to perform a decoding process, The decoding process includes: determining whether to split a current luma virtual pipeline decoding unit (VPDU) into smaller blocks, wherein a chroma virtual pipeline decoding unit corresponding to the current luma virtual pipeline decoding unit is split into smaller blocks; In response to determining not to split the current luma virtual pipeline decoding unit into smaller blocks, predicting a block of chroma samples without using luma samples; In response to determining to split the luma virtual pipeline decoding unit into smaller blocks, using luma samples to predict blocks of chroma samples; and decoding the block using the predicted chroma samples, The determining whether to split the current brightness virtual pipeline decoding unit into smaller blocks is based on a split flag or a quadtree split depth or a threshold block size.