Systems and methods for video coding

By splitting the luminance VPDU in the encoder and decoder as needed and using or not using luminance samples to predict chrominance samples, the problems of low coding efficiency and large circuit size in existing video coding technologies are solved, achieving more efficient image processing and resource utilization.

CN120956892APending Publication Date: 2025-11-14PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511206143.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-05-17
Filing Date
2020-05-15
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video coding technologies offer limited improvement in coding efficiency and image quality when processing ever-increasing volumes of digital video data, and their large circuit size makes it difficult to meet the ever-growing demands.

Method used

The encoding and decoding process is achieved by determining in the encoder and decoder whether to split the luminance virtual pipeline decoding unit (VPDU) into smaller blocks, and using or not using luminance samples to predict chrominance sample blocks under different conditions.

Benefits of technology

It improves encoding efficiency, enhances image quality, reduces the utilization of processing resources, reduces circuit size, and increases the processing speed of encoding/decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956892A_ABST
    Figure CN120956892A_ABST
Patent Text Reader

Abstract

An encoder includes a circuit and a memory coupled to the circuit. The circuitry determines whether to split a current luma virtual pipeline decode unit (VPDU) into smaller blocks. When it is determined that the current luma VPDU is not split into smaller blocks, the circuitry predicts a block of chroma samples without using luma samples. When it is determined that the luma VPDU is split into smaller blocks, circuitry predicts blocks of chroma samples using luma samples. The circuitry encodes the block using the predicted chroma samples.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the same patent application, filed on May 15, 2020, with application number 202080028161.5. Technical Field

[0002] This invention relates to video coding, and more particularly, to video coding and decoding systems, components and methods in video coding and decoding, such as for performing encoding of blocks using predicted chroma samples. Background Technology

[0003] With advancements in video coding technologies, from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High-Efficiency Video Coding), and H.266 / VVC (Multi-Functional Video Codec), there remains a continuous need to improve and optimize video coding technologies to handle the ever-increasing volumes of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video coding, particularly in the use of predicted chroma samples to perform block coding. Summary of the Invention

[0004] In one aspect, the encoder includes circuitry and memory coupled to that circuitry. The circuitry determines whether to break down the current luminance virtual pipeline decoding unit (VPDU) into smaller blocks. When it is determined that the current luminance VPDU should not be broken down into smaller blocks, the circuitry predicts blocks of chrominance samples without using luminance samples. When it is determined that the luminance VPDU should be broken down into smaller blocks, the circuitry uses luminance samples to predict blocks of chrominance samples. The circuitry then encodes the blocks using the predicted chrominance samples.

[0005] In one aspect, an encoder includes: a block splitter that splits a first image into multiple blocks in operation; an intra-frame predictor that predicts blocks included in the first image using reference blocks included in the first image in operation; an inter-frame predictor that predicts blocks included in the first image using reference blocks included in a second image different from the first image in operation; a cyclic filter that filters the blocks included in the first image in operation; a transformer that transforms a prediction error between the original signal and a prediction signal generated by the intra-frame predictor or the inter-frame predictor in operation to generate transform coefficients; a quantizer that quantizes the transform coefficients in operation to generate quantized coefficients; and an entropy encoder that variablely encodes the quantized coefficients in operation to generate an encoded bitstream including the encoded quantized coefficients and control information. Predicting blocks includes: determining whether to split the current luminance virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining that the current luminance VPDU should not be split into smaller blocks, blocks of chrominance samples are predicted without using luminance samples. In response to determining that the luminance VPDU is to be split into smaller blocks, luminance samples are used to predict blocks of chrominance samples.

[0006] On one hand, the decoder includes circuitry and memory coupled to that circuitry. The circuitry determines whether to break the current luminance virtual pipeline decoding unit (VPDU) into smaller blocks. When it is determined that the current luminance VPDU should not be broken into smaller blocks, the circuitry predicts blocks of chrominance samples without using luminance samples. When it is determined that the luminance VPDU should be broken into smaller blocks, the circuitry uses luminance samples to predict blocks of chrominance samples. The circuitry then decodes the blocks using the predicted chrominance samples.

[0007] In one aspect, a decoding device includes: a decoder that decodes an encoded bitstream in operation to output quantized coefficients; an inverse quantizer that in operation inverse-quantizes the quantized coefficients to output transform coefficients; an inverse transformer that in operation inversely transforms the transform coefficients to output a prediction error; an intra-frame predictor that in operation uses a reference block included in a first image to predict a block included in the first image; an inter-frame predictor that in operation uses a reference block included in a second image different from the first image to predict a block included in the first image; a cyclic filter that in operation filters the block included in the first image; and an output that in operation outputs an image including the first image. Predicting a block includes: determining whether to split a current luminance virtual pipeline decoding unit (VPDU) into smaller blocks. In response to determining that the current luminance VPDU should not be split into smaller blocks, predicting blocks of chrominance samples without using luminance samples. In response to determining that the luminance VPDU should be split into smaller blocks, predicting blocks of chrominance samples using luminance samples.

[0008] In one aspect, an encoding method includes: determining whether to split the current luminance virtual pipeline decoding unit (VPDU) into smaller blocks; in response to determining not to split the current luminance VPDU into smaller blocks, predicting blocks of chrominance samples without using luminance samples; in response to determining to split the luminance VPDU into smaller blocks, using luminance samples to predict blocks of chrominance samples; and encoding the blocks using the predicted chrominance samples.

[0009] In one aspect, a decoding method includes: determining whether to split the current luminance virtual pipeline decoding unit (VPDU) into smaller blocks; in response to determining not to split the current luminance VPDU into smaller blocks, predicting blocks of chrominance samples without using luminance samples; in response to determining to split the luminance VPDU into smaller blocks, using luminance samples to predict blocks of chrominance samples; and decoding the blocks using the predicted chrominance samples.

[0010] In video coding technology, there is a desire to propose new methods to improve coding efficiency, enhance image quality, and reduce circuit size. Some embodiments of this disclosure, including constituent elements of embodiments of this disclosure considered individually or in various combinations, can facilitate one or more of the following: improved coding efficiency, enhanced image quality, reduced utilization of processing resources associated with encoding / decoding, reduced circuit size, increased encoding / decoding processing speed, and so on.

[0011] Furthermore, some implementations of embodiments of this disclosure, including constituent elements of embodiments of this disclosure considered individually or in various combinations, can facilitate encoding and decoding, and the appropriate selection of one or more elements (e.g., filters, blocks, sizes, motion vectors, reference pictures, reference blocks, or operations). It should be noted that this disclosure includes disclosures regarding configurations and methods that can provide advantages beyond those described above. Examples of such configurations and methods include configurations or methods for improving encoding efficiency while reducing the increase in processing resource usage.

[0012] Additional benefits and advantages of the disclosed embodiments will become apparent from the specification and accompanying drawings. Benefits and / or advantages may be obtained individually from the various embodiments and features in the specification and drawings, and it is not necessary to provide all of these embodiments and features to obtain one or more of such benefits and / or advantages.

[0013] It should be noted that general or specific embodiments may be implemented as systems, methods, integrated circuits, computer programs, storage media, or any alternative combination thereof. Attached Figure Description

[0014] Figure 1This is a schematic diagram illustrating an example of the functional configuration of a transmission system according to an embodiment.

[0015] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream.

[0016] Figure 3 This is a conceptual diagram used to illustrate an example of slice configuration.

[0017] Figure 4 This is a conceptual diagram used to illustrate an example of tile configuration.

[0018] Figure 5 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.

[0019] Figure 6 This is a conceptual diagram used to illustrate an example of the coding structure in scalable coding.

[0020] Figure 7 This is a block diagram illustrating the functional configuration of an encoder according to an embodiment.

[0021] Figure 8 This is a functional block diagram illustrating an example of encoder installation.

[0022] Figure 9 This is a flowchart illustrating an example of the overall encoding process performed by the encoder.

[0023] Figure 10 This is a conceptual diagram used to illustrate an example of block splitting.

[0024] Figure 11 This is a block diagram illustrating an example of the functional configuration of a splitter according to an embodiment.

[0025] Figure 12 This is a conceptual diagram used to illustrate an example of a splitting pattern.

[0026] Figure 13A This is a conceptual diagram used to illustrate an example of a syntax tree for splitting patterns.

[0027] Figure 13B This is a conceptual diagram used to illustrate another example of a syntax tree for splitting patterns.

[0028] Figure 14 It is a graph representing example transformation basis functions used for various transformation types.

[0029] Figure 15 This is a conceptual diagram used to illustrate an example spatial transformation (SVT).

[0030] Figure 16This is a flowchart illustrating an example of a process performed by a converter.

[0031] Figure 17 This is a flowchart illustrating another example of a process performed by a converter.

[0032] Figure 18 This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment.

[0033] Figure 19 This is a flowchart illustrating an example of the quantization process performed by a quantizer.

[0034] Figure 20 This is a block diagram illustrating an example of the functional configuration of an entropy encoder according to an embodiment.

[0035] Figure 21 This is a conceptual diagram illustrating an example flow of the context-based adaptive binary arithmetic coding (CABAC) process in an entropy encoder.

[0036] Figure 22 This is a block diagram illustrating an example of the functional configuration of a cyclic filter according to an embodiment.

[0037] Figure 23A This is a conceptual diagram used to illustrate an example of the filter shape used in an adaptive cyclic filter (ALF).

[0038] Figure 23B This is a conceptual diagram used to illustrate another example of the filter shape used in ALF.

[0039] Figure 23C This is a conceptual diagram used to illustrate another example of the filter shape used in ALF.

[0040] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF).

[0041] Figure 23E This is a conceptual diagram used to illustrate an example of the filter shape used in CC-ALF.

[0042] Figure 23F This is a conceptual diagram used to illustrate an example process of Joint Chromaticity CCALF (JC-CCALF).

[0043] Figure 23G This is a table showing example weight index candidates that can be used in JC-CCALF.

[0044] Figure 24 This is a block diagram illustrating an example of a specific configuration of a cyclic filter used as a deblocking filter (DBF).

[0045] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics about the block boundaries.

[0046] Figure 26 It is a conceptual diagram used to illustrate the block boundaries for which the deblocking filtering process is performed.

[0047] Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.

[0048] Figure 28 This is a flowchart illustrating an example of the process performed by the encoder's predictor.

[0049] Figure 29 This is a flowchart illustrating another example of the process performed by the encoder's predictor.

[0050] Figure 30 This is a flowchart illustrating another example of the process performed by the encoder's predictor.

[0051] Figure 31 This is a conceptual diagram used to illustrate the sixty-seven intra-prediction modes used in the intra-prediction in the embodiments.

[0052] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor.

[0053] Figure 33 This is a concept diagram used to illustrate a reference image.

[0054] Figure 34 This is a concept diagram used to illustrate a list of reference images.

[0055] Figure 35 This is a flowchart illustrating the basic processing flow of an example inter-frame prediction.

[0056] Figure 36 This is a flowchart illustrating an example of the process of deriving the motion vector.

[0057] Figure 37 This is another example of a flowchart illustrating the process of deriving the motion vector.

[0058] Figure 38A This is a conceptual diagram used to illustrate an example representation of the MV derivation pattern.

[0059] Figure 38B This is a conceptual diagram used to illustrate an example representation of the MV derivation pattern.

[0060] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in normal inter-frame mode.

[0061] Figure 40 This is a flowchart illustrating an example of the inter-frame prediction process in normal merging mode.

[0062] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process in the merging mode.

[0063] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process of the HMVP merging pattern for the current image.

[0064] Figure 43 This is a flowchart illustrating an example of the Frame Rate Upconversion (FRUC) process.

[0065] Figure 44 This is a conceptual diagram used to illustrate an example of pattern matching (bilateral matching) between two blocks along a motion trajectory.

[0066] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image.

[0067] Figure 46A This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block based on the motion vectors of multiple adjacent blocks.

[0068] Figure 46B This is a conceptual diagram used to illustrate an example of deriving the motion vector of each sub-block in an affine pattern that uses three control points.

[0069] Figure 47A This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.

[0070] Figure 47B This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.

[0071] Figure 47C This is a conceptual diagram used to illustrate the example MV derivation at the control point in an affine mode.

[0072] Figure 48A This is a conceptual diagram used to illustrate an affine pattern in which two control points are used.

[0073] Figure 48B This is a conceptual diagram used to illustrate an affine pattern in which three control points are used.

[0074] Figure 49AThis is a conceptual diagram illustrating an example of a method for deriving the MV at a control point when the number of control points used for the encoded block and the number of control points used for the current block are different from each other.

[0075] Figure 49B This is a conceptual diagram illustrating another example of a method for deriving the MV at a control point when the number of control points used for the encoded block and the number of control points used for the current block are different from each other.

[0076] Figure 50 This is a flowchart illustrating an example of the process in the affine merge pattern.

[0077] Figure 51 This is a flowchart illustrating an example of a process in affine inter-frame mode.

[0078] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangular predicted images.

[0079] Figure 52B This is a conceptual diagram used to illustrate an example of a first portion of a first partition that overlaps with a second partition, and first and second sets of samples that can be weighted as part of a correction process.

[0080] Figure 52C This is a conceptual diagram used to illustrate the first part of the first partition, which is the portion of the first partition that overlaps with a portion of an adjacent partition.

[0081] Figure 53 This is a flowchart illustrating an example of a process in a triangle pattern.

[0082] Figure 54 This is a conceptual diagram used to illustrate an example of an advanced temporal motion vector prediction (ATMVP) pattern in which the MV is derived on a sub-block basis.

[0083] Figure 55 This is a flowchart illustrating the relationship between merge mode and Dynamic Motion Vector Refresh (DMVR).

[0084] Figure 56 This is a conceptual diagram used to illustrate an example of DMVR.

[0085] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV.

[0086] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in DMVR.

[0087] Figure 58B This is a flowchart illustrating an example of the motion estimation process in a DMVR.

[0088] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.

[0089] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.

[0090] Figure 61 This is a flowchart illustrating an example of the correction process performed on the predicted image by Overlapping Block Motion Compensation (OBMC).

[0091] Figure 62 This is a conceptual diagram illustrating an example of the predictive image correction process performed by OBMC.

[0092] Figure 63 It is a conceptual diagram used to illustrate a model that assumes uniform linear motion.

[0093] Figure 64 This is a flowchart illustrating an example of the inter-frame prediction process based on BIO.

[0094] Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor that can perform inter-frame prediction based on BIO.

[0095] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses a brightness correction process performed by a LIC.

[0096] Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.

[0097] Figure 67 This is a block diagram illustrating the functional configuration of the decoder according to an embodiment.

[0098] Figure 68 This is a functional block diagram illustrating an example of decoder installation.

[0099] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by the decoder.

[0100] Figure 70 It is a conceptual diagram used to illustrate the relationship between the splitting determinant and other constituent elements.

[0101] Figure 71 This is a block diagram illustrating an example of the functional configuration of an entropy decoder.

[0102] Figure 72 This is a conceptual diagram used to illustrate an example flow of the CABAC process in an entropy decoder.

[0103] Figure 73 This is a block diagram illustrating an example of the functional configuration of an inverse quantizer.

[0104] Figure 74 This is a flowchart illustrating an example of the inverse quantization process performed by an inverse quantizer.

[0105] Figure 75 This is a flowchart illustrating an example of a process performed by an inverse transformer.

[0106] Figure 76 This is a flowchart illustrating another example of the process performed by the inverse converter.

[0107] Figure 77 This is a block diagram illustrating an example of the functional configuration of a cyclic filter.

[0108] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of the decoder.

[0109] Figure 79 This is a flowchart illustrating another example of the process performed by the predictor of the decoder.

[0110] Figure 80 This is a flowchart illustrating another example of the process performed by the predictor of the decoder.

[0111] Figure 81 This is a diagram illustrating an example of the process performed by the decoder's intra-frame predictor.

[0112] Figure 82 This is a flowchart illustrating an example of the MV derivation process in the decoder.

[0113] Figure 83 This is a flowchart illustrating another example of the MV derivation process in the decoder.

[0114] Figure 84 This is a flowchart illustrating an example of the inter-frame prediction process in the decoder using normal inter-frame mode.

[0115] Figure 85 This is a flowchart illustrating an example of the inter-frame prediction process in the decoder using normal merging mode.

[0116] Figure 86 This is a flowchart illustrating an example of the inter-frame prediction process in the decoder using FRUC mode.

[0117] Figure 87 This is a flowchart illustrating an example of the inter-frame prediction process in the decoder using affine merging mode.

[0118] Figure 88This is a flowchart illustrating an example of the inter-frame prediction process in the decoder using affine inter-frame modes.

[0119] Figure 89 This is a flowchart illustrating an example of the inter-frame prediction process using triangle patterns in the decoder.

[0120] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed via DMVR in the decoder. Figure 91 This is a flowchart illustrating an example process of motion estimation via DMVR in the decoder.

[0121] Figure 92 This is a flowchart illustrating an example of the process of generating a predicted image in a decoder.

[0122] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in the decoder.

[0123] Figure 94 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via OBMC.

[0124] Figure 95 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via BIO.

[0125] Figure 96 This is a flowchart illustrating an example of the correction process for the predicted image in the decoder via LIC.

[0126] Figure 97 This is a flowchart illustrating an example of the process of decoding a block using predicted chroma samples.

[0127] Figure 98 This is a conceptual diagram used to illustrate an example of determining whether the current chroma block is inside an MxN non-overlapping region aligned with the chroma samples of the MxN grid.

[0128] Figure 99 This is a conceptual diagram used to illustrate an example of determining whether the current chroma block is inside an MxN non-overlapping region aligned with the chroma samples of the MxN grid.

[0129] Figure 100 This is a conceptual diagram used to illustrate a Virtual Pipeline Decoding Unit (VPDU).

[0130] Figure 101 This is a conceptual diagram used to illustrate an example of determining whether the current VPDU can be used to predict chromaticity samples.

[0131] Figure 102This is a conceptual diagram used to illustrate an example of how to determine whether a luminance VPDU will be split into smaller blocks.

[0132] Figure 103 This is a conceptual diagram used to illustrate how a threshold size is used to determine whether to use luminance samples to predict the chrominance samples of a block.

[0133] Figure 104 This is a conceptual diagram used to illustrate an example of a non-rectangular partition.

[0134] Figure 105 This is a diagram illustrating an example overall configuration of a content delivery system used to implement a content distribution service. Figure 106 This is a conceptual diagram illustrating an example of a display screen used to show a webpage.

[0135] Figure 107 This is a conceptual diagram illustrating an example of a display screen used to show a webpage.

[0136] Figure 108 This is a block diagram showing an example of a smartphone.

[0137] Figure 109 This is a block diagram illustrating an example of the functional configuration of a smartphone. Detailed Implementation

[0138] In the accompanying drawings, unless the context otherwise indicates, the same reference numerals denote the same elements. The size and relative position of the elements in the drawings are not necessarily drawn to scale.

[0139] In the following description, embodiments will be illustrated with reference to the accompanying drawings. Note that the embodiments described below each illustrate a general or specific example. The numerical values, shapes, materials, components, arrangements and connections of components, steps, relationships and sequences of steps referred to in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0140] Embodiments of the encoder and decoder will now be described. The embodiments are examples of encoders and decoders to which the processes and / or configurations presented in the description of aspects of this disclosure may be applied. The processes and / or configurations may also be implemented in encoders and decoders different from those according to the embodiments. For example, with respect to the processes and / or configurations applied to the embodiments, any of the following may be implemented:

[0141] (1) Any component of the encoder or decoder of the embodiments presented in the description of this disclosure may be replaced or combined with another component presented anywhere in the description of this disclosure.

[0142] (2) In the encoder or decoder according to the embodiment, any function or process performed by one or more components of the encoder or decoder may be arbitrarily changed, such as by adding, replacing, or removing a function or process. For example, any function or process may be replaced or combined with another function or process presented anywhere in the description of this disclosure.

[0143] (3) In the method implemented by the encoder or decoder according to the embodiment, variations may be made as appropriate, such as adding, replacing, and removing one or more processes included in the method. For example, any process in the method may be replaced by or combined with another process presented anywhere in the description of an aspect of this disclosure.

[0144] One or more components included in the encoder or decoder according to the embodiment may be combined with components presented anywhere in the description of this disclosure, may be combined with components presenting one or more functions presented anywhere in the description of this disclosure, and may be combined with components that implement one or more processes implemented by components presented in the description of this disclosure.

[0145] A component that includes one or more functions of an encoder or decoder according to an embodiment, or a component that implements one or more processes of an encoder or decoder according to an embodiment, may be combined with or replaced by a component presented anywhere in the description of this disclosure, or may be combined with or replaced by a component that implements one or more processes presented anywhere in the description of this disclosure.

[0146] In a method implemented by an encoder or decoder according to an embodiment, any process included in the method may be replaced by or combined with a process presented anywhere in the description of this disclosure, or by or combined with any corresponding or equivalent process.

[0147] One or more processes included in a method implemented by an encoder or decoder according to an embodiment may be combined with processes presented anywhere in the description of aspects of this disclosure.

[0148] The implementation of the processes and / or configurations presented in the description of aspects of this disclosure is not limited to the encoder or decoder according to the embodiments. For example, the processes and / or configurations may be implemented in devices for purposes different from the motion picture encoder or motion picture decoder disclosed in the embodiments.

[0149] (Terminology Definition)

[0150] The terms can be defined as follows as an example.

[0151] An image is a data unit composed of a set of pixels; it is a picture, or a block of data smaller than a pixel. Besides video, images also include still images.

[0152] An image is an image processing unit configured with a set of pixels, and can also be referred to as a frame or field. For example, an image can be an array of luminance samples in monochrome format, or an array of luminance samples and two corresponding chrominance sample arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0153] A block is a processing unit, which is a defined set of pixels. Blocks can have any number of different shapes. For example, a block can be a rectangle of M×N (M columns × N rows) pixels, a square of M×M pixels, a triangle, a circle, etc. Examples of blocks include slices, tiles, blocks, CTUs, superblocks, basic splitting units, VPDUs, hardware-specific processing splitting units, CUs, processing block units, prediction block units (PUs), orthogonal transform block units (TUs), units, and subblocks. Blocks can take the form of an M×N array of samples or an M×N array of transform coefficients. For example, a block can be a square or rectangular region of pixels comprising a luma matrix and two chroma matrices.

[0154] A pixel or sample is the smallest point in an image. A pixel or sample includes pixels at integer positions and pixels at sub-pixel positions, such as pixels generated based on pixels at integer positions.

[0155] Pixel values ​​or sample values ​​are the feature values ​​of a pixel. Pixel values ​​or sample values ​​can include one or more of the following: luminance value, chrominance value, RGB grayscale level, depth value, binary value 0 or 1, etc.

[0156] Chromaticity, or color intensity, is the intensity of a color and is usually represented by the symbols Cb and Cr. These symbols specify the value of an array of samples or a single sample value representing the value of one of two color difference signals associated with the primary color.

[0157] Brightness or luminance is the lightness of an image, usually represented by the symbol or subscript Y or L, which specifies: the value of a sample array or a single sample value representing the value of a monochromatic signal associated with the primary color.

[0158] Flags consist of one or more bits that indicate the value of a parameter or index, for example. Flags can be binary flags, which represent the binary value of the flag, or they can represent the non-binary value of a parameter.

[0159] Signals convey information, which is symbolized or encoded into signals. Signals include discrete digital signals and continuous analog signals.

[0160] A stream or bitstream is a string of digital data. A stream or bitstream can be a single stream or can be configured with multiple streams having multiple hierarchical layers. A stream or bitstream can be sent using a single transmission path in a serial communication manner, or it can be sent using multiple transmission paths in a packet communication manner.

[0161] Difference refers to various mathematical differences, such as simple difference (xy), absolute value of difference (|xy|), difference of squares (x^2-y^2), square root of difference (√(xy)), weighted difference (ax-by: a and b are constants), offset difference (x-y+a: a is the offset), etc. In the case of scalars, simple difference is sufficient and includes difference calculation.

[0162] The term "sum" refers to various mathematical sums, such as the simple sum (x+y), the absolute value of the sum (|x+y|), the sum of squares (x^2+y^2), the square root of the sum (√(x+y)), the weighted sum (ax+by: a and b are constants), and the offset sum (x+y+a: a is the offset), etc. In the case of scalars, the simple sum is sufficient and includes sum calculations.

[0163] A frame is a combination of a top field and a bottom field, where sample lines 0, 2, 4, ... originate from the top field, and sample lines 1, 3, 5, ... originate from the bottom field.

[0164] A slice is an integer number of coded tree units contained in an independent slice and all subsequent dependent slices (if any) before the next independent slice (if any) within the same access unit.

[0165] A tile is a rectangular region of a coded tree block within a specific tile column and row in an image. A tile can be a rectangular region of a frame designed to be decoded and encoded independently, although cyclic filtering can still be applied across the tile edges.

[0166] A coding tree unit (CTU) can be a coding tree block of luminance samples from an image with three sample arrays, or two corresponding coding tree blocks of chrominance samples. Alternatively, a CTU can be a coding tree block of samples from a monochrome image and one of the images encoded using three separate color planes and a syntax structure for encoding the samples. A superblock can be a 64×64 pixel square block consisting of one or two pattern information blocks, or recursively divided into four 32×32 blocks, which themselves can be further subdivided.

[0167] (System Configuration)

[0168] First, the transmission system according to an embodiment will be described. Figure 1 This is a schematic diagram illustrating an example configuration of a transmission system 400 according to an embodiment.

[0169] The transmission system 400 is a system for transmitting a stream generated by encoding an image and for decoding the transmitted stream. As shown in the figure, the transmission system 400 includes... Figure 1 The encoder 100, network 300, and decoder 200 are shown.

[0170] An image is input to encoder 100. Encoder 100 generates a stream by encoding the input image and outputs the stream to network 300. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed through encoding.

[0171] It should be noted that the image before encoding by encoder 100 is also referred to as the original image, original signal, or original sample. The image can be video or a still image. Image is a general concept encompassing sequences, pictures, and blocks, and therefore, unless otherwise stated, is not limited to spatial regions of a specific size or temporal regions of a specific size. An image is an array of pixels or pixel values, and the signal representing the image or pixel values ​​is also called a sample. The stream can be referred to as a bitstream, an encoded bitstream, a compressed bitstream, or an encoded signal. Furthermore, encoder 100 can be referred to as an image encoder or a video encoder. The encoding method performed by encoder 100 can be referred to as an encoding method, an image encoding method, or a video encoding method.

[0172] Network 300 sends the stream generated by encoder 100 to decoder 200. Network 200 can be any combination of the Internet, wide area network (WAN), local area network (LAN), or network. Network 300 is not limited to a two-way communication network and can be a one-way communication network that transmits broadcast waves such as digital terrestrial broadcasts and satellite broadcasts. Alternatively, network 300 can be replaced by a recording medium such as a digital multifunction disc (DVD) and Blu-ray disc (BD) on which the stream is recorded.

[0173] Decoder 200 generates a decoded image, such as an uncompressed image, by decoding the stream sent by network 300. For example, the decoder decodes the stream according to a decoding method corresponding to the encoding method used by encoder 100.

[0174] It should be noted that the decoder 200 can also be called an image decoder or a video decoder, and the decoding method performed by the decoder 200 can also be called a decoding method, an image decoding method, or a video decoding method.

[0175] (Data Structures)

[0176] Figure 2 This is a conceptual diagram used to illustrate an example of the hierarchical structure of data in a stream. For convenience, reference will be made to... Figure 1 The transmission system 400 is used to describe Figure 2Streams include, for example, video sequences. (e.g.) Figure 2 As shown in (a), the video sequence includes one or more video parameter sets (VPS), one or more sequence parameter sets (SPS), one or more picture parameter sets (PPS), supplementary enhancement information (SEI), and multiple pictures.

[0177] In a video with multiple layers, a VPS may include encoding parameters shared between some of the layers, as well as encoding parameters associated with some of the layers included in the video or with a single layer.

[0178] The SPS includes parameters for the sequence, that is, the encoding parameters that the decoder 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the image. It should be noted that multiple SPSs may exist.

[0179] The PPS includes parameters for the images, that is, encoding parameters that the decoder 200 references in order to decode each image in the sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the image and a flag indicating the application of weighted prediction. It should be noted that multiple SPSs may exist. Each of the SPS and PPS can be simply referred to as a parameter set.

[0180] like Figure 2 As shown in (b), an image may include an image header and one or more slices. The image header includes encoding parameters referenced by the decoder 200 for decoding the one or more slices.

[0181] like Figure 2 As shown in (c), a slice includes a slice header and one or more blocks. The slice header includes encoding parameters that the decoder 200 references in order to decode one or more blocks.

[0182] like Figure 2 As shown in (d), a block comprises one or more coding tree units (CTUs).

[0183] It should be noted that an image may not include any slices and may include groups of tiles instead of slices. In this case, a group of tiles includes at least one tile. Furthermore, a block may include slices.

[0184] CTU can also be called a superblock or basic splitting unit. For example... Figure 2 As shown in (e), the CTU includes a CTU header and at least one coding unit (CU). As shown, the CTU includes four coding units CU (10), CU (11), CU (12), and CU (13). The CTU header includes coding parameters referenced by the decoder 200 for decoding at least one CU.

[0185] A CU can be subdivided into multiple smaller CUs. As shown in the figure, CU(10) is not subdivided into smaller coding units; CU(11) is subdivided into four smaller coding units CU(110), CU(111), CU(112), and CU(113); CU(12) is not subdivided into smaller coding units; while CU(13) is subdivided into seven smaller coding units CU(1310), CU(1311), CU(1312), CU(1313), CU(132), CU(133), and CU(134). Figure 2 As shown in (f), the CU includes a CU header, prediction information, and residual coefficient information. The prediction information is used to predict the CU, while the residual coefficient information represents the prediction residuals, which will be described later. Although the CU is essentially the same as the prediction unit (PU) and the transform unit (TU), it should be noted that, for example, the subblock transform (SBT) described later may include multiple TUs smaller than the CU. Furthermore, the CU can be processed for each virtual pipeline decoding unit (VPDU) included in the CU. The VPDU is, for example, a fixed unit that can be processed at one stage when pipeline processing is performed in hardware.

[0186] It should be noted that the flow may not include Figure 2 All hierarchical layers are shown. The order of the hierarchical layers can be interchanged, or any hierarchical layer can be replaced by another hierarchical layer. Here, the image that is the target of a process to be performed by a device such as encoder 100 or decoder 200 is called the current image. When the process is an encoding process, the current image refers to the current image to be encoded; and when the process is a decoding process, the current image refers to the current image to be decoded. Similarly, for example, a CU or CU block that is the target of a process to be performed by a device such as encoder 100 or decoder 200 is called the current block. When the process is an encoding process, the current block refers to the current block to be encoded; and when the process is a decoding process, the current block refers to the current block to be decoded.

[0187] (Image structure: slices / tiles)

[0188] An image can be configured with one or more slice units or one or more tile units to facilitate parallel encoding / decoding of the image.

[0189] A slice is a basic coding unit included in an image. An image may include, for example, one or more slices. Furthermore, a slice includes one or more coding tree units (CTUs).

[0190] Figure 3 This is a conceptual diagram used to illustrate an example of slice configuration. For example, in Figure 3The image contains 11×8 CTUs and is divided into four slices (slices 1 to 4). Slice 1 contains 16 CTUs, slice 2 contains 21 CTUs, slice 3 contains 29 CTUs, and slice 4 contains 22 CTUs. Each CTU in the image belongs to one of the slices. The shape of each slice is obtained by horizontally splitting the image. The boundaries of each slice do not need to coincide with the edges of the image and can coincide with any boundaries between CTUs in the image. The processing order (encoding or decoding order) of the CTUs in the slice is, for example, the raster scan order. Each slice includes a slice header and encoded data. Slice characteristics can be written into the slice header. Characteristics may include the CTU address of the top CTU in the slice, the slice type, etc.

[0191] A tile is a rectangular area unit included in an image. The tiles of an image can be assigned a number called TileId according to the raster scan order.

[0192] Figure 4 This is a conceptual diagram used to illustrate an example of tile configuration. For example, in Figure 4 In the image, there are 11×8 CTUs, divided into four tiles (tiles 1 to 4) within a rectangular region. When using tiles, the processing order of the CTUs can differ from that without tiles. When not using tiles, multiple CTUs in the image are typically processed in raster scan order. When using multiple tiles, at least one CTU in each of the multiple tiles is processed in raster scan order. For example, as... Figure 4 As shown, the processing order of the CTU included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then continues from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0193] It should be noted that a tile may include one or more slices, and a slice may include one or more tiles.

[0194] It should be noted that an image can be configured with one or more tile sets. A tile set can include one or more tile groups, or one or more individual tiles. An image can be configured with one of the following: tile sets, tile groups, and individual tiles. For example, suppose that the order in which multiple tiles are scanned for each tile set in a raster scan order is the basic coding order of the tiles. Suppose that the set of one or more consecutive tiles in the basic coding order within each tile set is a tile group. Such an image can be processed by a splitter 102 described later (see [link to splitter 102]). Figure 7 Configure it using ).

[0195] (Scalable encoding)

[0196] Figure 5 and Figure 6This is a conceptual diagram illustrating an example of a scalable flow structure, and for convenience, references are provided. Figure 1 Describe it.

[0197] like Figure 5 As shown, encoder 100 can generate a temporally / spatially scalable stream by dividing each of multiple images into any of multiple layers and encoding the images within those layers. For example, encoder 100 encodes the images for each layer, thereby achieving scalability even when enhancement layers exist on top of the base layer. This encoding of each image is also referred to as scalable encoding. In this way, decoder 200 is able to switch the image quality of the images displayed by decoding the stream. In other words, decoder 200 can determine which layer to decode based on internal factors such as the processing power of decoder 200 and external factors such as the state of communication bandwidth. Therefore, decoder 200 is able to decode content while freely switching between low and high resolutions. For example, a user of the stream might watch half of a video stream on their smartphone on their way home and continue watching it at home on an internet-connected device (such as a television). It should be noted that each of the aforementioned smartphones and devices includes decoder 200 with the same or different performance. In this case, the user can watch high-quality video at home as the device decodes layers to higher layers in the stream. In this way, encoder 100 does not need to generate multiple streams with different image qualities having the same content, and thus the processing load can be reduced.

[0198] Furthermore, the enhancement layer may include metadata based on statistical information about the image. The decoder 200 may generate a video whose image quality has been enhanced by performing super-resolution imaging on the images in the base layer based on the metadata. Super-resolution imaging may include, for example, an increase in the SN ratio at the same resolution, an increase in resolution, etc. The metadata may include, for example, information used to identify linear or nonlinear filter coefficients (as used in the super-resolution process), or information identifying parameter values ​​in the filtering process, or information such as machine learning, least squares methods, etc., used in the super-resolution processing.

[0199] In an embodiment, a configuration can be provided in which an image is divided into, for example, tiles based on the meaning of objects within it. In this case, decoder 200 can decode only a portion of the image by selecting tiles to decode. Furthermore, the attributes of objects (people, cars, balls, etc.) and their positions within the image (coordinates within the same image) can be stored as metadata. In this case, decoder 200 can identify the location of a desired object based on the metadata and determine the tiles that include that object. For example, as... Figure 6As shown, metadata can be stored using a different data storage structure than image data (e.g., SEI (Supplemental Enhancement Information) messages in HEVC). This metadata indicates, for example, the location, size, or color of the primary object.

[0200] Metadata can be stored in units of multiple images (e.g., streams, sequences, or random access units). In this way, decoder 200 can obtain, for example, the time when a specific person appears in the video, and by fitting the time information with the image unit information, can identify the image in which the object (person) exists and determine the object's position in the image.

[0201] (encoder)

[0202] An encoder according to an embodiment will be described. Figure 7 This is a block diagram illustrating the functional configuration of an encoder 100 according to an embodiment. The encoder 100 is a video encoder that encodes video in blocks.

[0203] like Figure 7 As shown, encoder 100 is a device for encoding images in blocks, and it includes a splitter 102, a subtractor 104, a transformer 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse transformer 114, an adder 116, a block memory 118, a cyclic filter 120, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, a prediction controller 128, and a prediction parameter generator 130. As shown, the intra-frame predictor 124 and the inter-frame predictor 126 are part of the prediction controller.

[0204] The encoder 100 is implemented, for example, as a general-purpose processor and memory. In this case, when the software program stored in memory is executed by the processor, the processor acts as a splitter 102, a subtractor 104, a converter 106, a quantizer 108, an entropy encoder 110, an inverse quantizer 112, an inverse converter 114, an adder 116, a cyclic filter 120, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. Alternatively, the encoder 100 may be implemented as one or more dedicated electronic circuits corresponding to the splitter 102, subtractor 104, converter 106, quantizer 108, entropy encoder 110, inverse quantizer 112, inverse converter 114, adder 116, cyclic filter 120, intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128.

[0205] (Encoder installation example)

[0206] Figure 8 This is a functional block diagram illustrating an example installation of encoder 100. Encoder 100 includes a processor a1 and a memory a2. For example, Figure 7 The encoder 100 shown is mounted on multiple components. Figure 8 The processor a1 and memory a2 shown are shown.

[0207] Processor a1 is a circuit that performs information processing and is coupled to memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that encodes images. Processor a1 can be a processor, such as a CPU. Furthermore, processor a1 can be an aggregation of multiple electronic circuits. Additionally, for example, processor a1 can function as... Figure 7 The function of two or more of the constituent elements of the encoder 100 shown, etc.

[0208] Memory a2 is a dedicated or general-purpose memory used by processor a1 to encode images. Memory a2 can be an electronic circuit and can be connected to processor a1. Furthermore, memory a2 can be included within processor a1. Additionally, memory a2 can be an aggregation of multiple electronic circuits. Furthermore, memory a2 can be a disk, optical disk, etc., or can be represented as a storage or recording medium. Furthermore, memory a2 can be non-volatile memory or volatile memory.

[0209] For example, memory a2 can store the image to be encoded or the bitstream corresponding to the encoded image. Furthermore, memory a2 can store the program used to enable processor a1 to encode the image.

[0210] Furthermore, for example, memory a2 can serve as Figure 7 The encoder 100 shown has multiple constituent elements, including two or more elements used for storing information. For example, memory a2 can serve as... Figure 7 The block memory 118 and frame memory 122 shown herein serve a specific purpose. More specifically, memory a2 can store reconstructed blocks, reconstructed images, etc.

[0211] It should be noted that in encoder 100, it is not necessary to implement... Figure 7 It specifies all the multiple constituent elements, etc., and does not perform all the processes described herein. Figure 7 Some of the constituent elements shown may be contained in another device, or some of the processes described herein may be performed by another device.

[0212] The following describes the overall flow of the process performed by encoder 100, and then describes each constituent element included in encoder 100.

[0213] (The overall flow of the coding process)

[0214] Figure 9This is a flowchart illustrating an example of the overall encoding process performed by encoder 100, and for convenience, references... Figure 7 Describe it.

[0215] First, the splitter 102 of the encoder 100 splits each image included in the input image into multiple blocks of a fixed size (e.g., 128 × 128 pixels) (step Sa_1). The splitter 102 then selects a splitting pattern for the fixed-size blocks (also referred to as block shapes) (step Sa_2). In other words, the splitter 102 further splits the fixed-size blocks into multiple blocks forming the selected splitting pattern. For each of the multiple blocks, the encoder 100 performs steps Sa_3 to Sa_9 for that block (i.e., the current block to be encoded).

[0216] The prediction controller 128 and the prediction actuator (which includes an intra-frame predictor 124 and an inter-frame predictor 126) generate a prediction image of the current block (step Sa-3). The prediction image may also be referred to as a prediction signal, a prediction block, or a prediction sample.

[0217] Next, subtractor 104 generates the difference between the current block and the predicted image as the prediction residual (step Sa_4). The prediction residual can also be referred to as the prediction error.

[0218] Next, the transformer 106 transforms the predicted image, and the quantizer 108 quantizes the result to generate multiple quantized coefficients (step Sa_5). The multiple quantized coefficients can sometimes be referred to as a coefficient block.

[0219] Next, the entropy encoder 110 encodes (specifically, entropy coding) multiple quantized coefficients and prediction parameters related to the generation of the predicted image to generate a stream (step Sa_6). This stream may sometimes be referred to as an encoded bitstream or a compressed bitstream.

[0220] Next, the inverse quantizer 112 inverse quantizes the multiple quantized coefficients, and the inverse transformer 114 inverse transforms the result to recover the prediction residual (step Sa_7).

[0221] Next, adder 116 adds the predicted image to the recovered prediction residual to reconstruct the current block (step Sa_8). This generates the reconstructed image. The reconstructed image can also be referred to as a reconstructed block or a decoded image block.

[0222] When generating the reconstructed image, the cyclic filter 120 performs filtering on the reconstructed image as needed (step Sa_9).

[0223] Encoder 100 then determines whether the encoding of the entire image has been completed (step Sa_10). If it is determined that the encoding has not been completed (no in step Sa_10), the process starting from step Sa_2 is repeated for the next block of the image.

[0224] Although in the example above, encoder 100 selects a splitting pattern for fixed-size blocks and encodes each block according to the splitting pattern, it should be noted that each block can be encoded according to a corresponding splitting pattern from a plurality of splitting patterns. In this case, encoder 100 can evaluate the cost of each of the plurality of splitting patterns and, for example, select the stream obtained by encoding according to the splitting pattern that produces the minimum cost as the output stream.

[0225] As shown in the figure, the processes in steps Sa_1 to Sa_10 are executed sequentially by encoder 100. Alternatively, two or more processes can be executed in parallel, processes can be reordered, and so on.

[0226] The encoder 100 employs a hybrid encoding process using predictive coding and transform coding. Furthermore, predictive coding is performed by an encoding loop configured with a subtractor 104, a transformer 106, a quantizer 108, an inverse quantizer 112, an inverse transformer 114, an adder 116, a cyclic filter 120, a block memory 118, a frame memory 122, an intra-frame predictor 124, an inter-frame predictor 126, and a prediction controller 128. In other words, the prediction executor configured with the intra-frame predictor 124 and the inter-frame predictor 126 is part of the encoding loop.

[0227] (Splitter)

[0228] Splitter 102 splits each image included in the original image into multiple blocks and outputs each block to subtractor 104. For example, splitter 102 first splits the image into fixed-size blocks (e.g., 128×128 pixels). Other fixed block sizes may be used. Fixed-size blocks are also referred to as coding tree units (CTUs). Splitter 102 then splits each fixed-size block into variable-size blocks (e.g., 64×64 pixels or smaller) based on recursive quadtree and / or binary tree block splitting. In other words, splitter 102 selects a splitting mode. Variable-size blocks may also be referred to as coding units (CUs), prediction units (PUs), or transform units (TUs). It should be noted that in various processing examples, there is no need to distinguish between CUs, PUs, and TUs; all or part of the blocks in the image can be processed in units of CUs, PUs, or TUs.

[0229] Figure 10 This is a conceptual diagram used to illustrate an example of block splitting according to an embodiment. Figure 10In the diagram, solid lines represent the block boundaries of blocks split by quadtree block splitting, while dashed lines represent the block boundaries of blocks split by binary tree block splitting.

[0230] Here, block 10 is a square block with 128×128 pixels (128×128 block). This 128×128 block 10 is first split into four square blocks of 64×64 pixels (quadtree block split).

[0231] The 64×64 pixel block in the top left corner is further vertically divided into two rectangular 32×64 pixel blocks, and the 32×64 pixel block on the left is further vertically divided into two rectangular 16×64 pixel blocks (binary tree block splitting). As a result, the 64×64 pixel block in the top left corner is split into two 16×64 pixel blocks 11 and 12, and a 32×64 pixel block 13.

[0232] The 64×64 pixel block in the upper right corner is horizontally divided into two rectangular 64×32 pixel blocks, 14 and 15 (binary tree block splitting).

[0233] The 64×64 pixel block in the lower left corner is first divided into four 32×32 pixel square blocks (quadtree block splitting). The top-left and bottom-right blocks of these four 32×32 pixel square blocks are further split. The top-left 32×32 pixel block is vertically split into two 16×32 pixel rectangular blocks, and the right-hand 16×32 pixel block is further horizontally split into two 16×16 pixel rectangular blocks (binary tree block splitting). The top-right 32×32 pixel square block is horizontally split into two 32×16 pixel rectangular blocks (binary tree block splitting). As a result, the 64×64 pixel block in the lower left corner was split into a rectangular 16×32 pixel block 16, two square 16×16 pixel blocks 17 and 18, two square 32×32 pixel blocks 19 and 20, and two rectangular 32×16 pixel blocks 21 and 22.

[0234] The 64×64 pixel block 23 in the bottom right corner was not split.

[0235] As mentioned above, in Figure 10 In this example, based on recursive quadtree and binary tree block splitting, block 10 is split into 13 variable-size blocks 11 to 23. This type of splitting is also known as quadtree plus binary tree (QTBT) splitting.

[0236] It should be pointed out that, in Figure 10In this context, a block is split into four or two blocks (quadtree or binary tree block split), but the split is not limited to these examples. For instance, a block can be split into three blocks (ternary block split). Splits that include this type of ternary block split are also known as multi-type tree (MBT) splits.

[0237] Figure 11 This is a block diagram illustrating an example of the functional configuration of a splitter according to one embodiment. Figure 11 As shown, the splitter 102 may include a block split determiner 102a. As an example, the block split determiner 102a may perform the following process.

[0238] For example, block splitting determiner 102a can obtain or retrieve block information from block memory 118 and / or frame memory 122, and determine a splitting mode (e.g., the splitting mode described above) based on the block information. Splitter 102 splits the original image according to the splitting mode and outputs at least one block obtained from the splitting to subtractor 104.

[0239] Furthermore, for example, the block splitting determiner 102a outputs one or more parameters indicating the determined splitting pattern (e.g., the splitting pattern described above) to the transformer 106, the inverse transformer 114, the intra-frame predictor 124, the inter-frame predictor 126, and the entropy encoder 110. The transformer 106 can predict the residual based on one or more parameters. The intra-frame predictor 124 and the inter-frame predictor 126 can generate a predicted image based on one or more parameters. Furthermore, the entropy encoder 110 can entropy encode one or more parameters.

[0240] Parameters related to the splitting mode can be written to the stream as shown in one example.

[0241] Figure 12 This is a conceptual diagram used to illustrate examples of splitting patterns. Examples of splitting patterns include: splitting into four regions (QT), where one block is divided into two regions, one horizontal and one vertical; splitting into three regions (HT or VT), where one block is split in the same direction at a 1:2:1 ratio; splitting into two regions (HB or VB), where one block is split in the same direction at a 1:1 ratio; and no splitting (NS).

[0242] It should be noted that the splitting mode does not have a block splitting direction when splitting into four regions or not splitting at all, while the splitting mode has splitting direction information when splitting into two or three regions.

[0243] Figure 13A This is a conceptual diagram used to illustrate an example of a syntax tree for splitting patterns.

[0244] Figure 13BThis is a conceptual diagram used to illustrate another example of a syntax tree for splitting patterns.

[0245] Figure 13A and Figure 13B This is a conceptual diagram used to illustrate an example of a syntax tree for splitting patterns. Figure 13A In the example, the first information indicates whether to split (S: split flag), followed by information indicating whether to split into four regions (QT: QT flag). Next is information indicating whether to split into two or three regions (TT: TT flag, or BT: BT flag), and then information indicating the splitting direction (Ver: vertical flag, or Hor: horizontal flag). It should be noted that each block in at least one block obtained by such a division can be further repeated in a similar process. In other words, as an example, whether to split, whether to split into four regions, which direction (horizontal or vertical) is the direction of the splitting method, and whether to split into three or two regions, can be determined recursively and can be determined based on the... Figure 13A The encoding order revealed by the syntax tree shown will determine the result encoded in the stream.

[0246] Additionally, although the information items indicating S, QT, TT, and Ver are respectively in Figure 13A The syntax tree shown is arranged in the listed order, but the information items indicating S, QT, Ver, and BT can be arranged in the listed order. That is, in Figure 13B In the example, the first information indicates whether to split (S: split flag), followed by information indicating whether to split into 4 regions (QT: QT flag). Next is information indicating the split direction (Ver: vertical flag, or Hor: horizontal flag), and then information indicating whether to split into two regions or three regions (BT: BT flag, or TT: TT flag).

[0247] It should be noted that the above-described splitting methods are examples, and other splitting methods besides those described above, or a portion of the above-described splitting methods, may be used.

[0248] (Subtractor)

[0249] Subtractor 104 subtracts the predicted image (predicted samples input from prediction controller 128 indicated below) from the original image in blocks, which are input from and split by splitter 102. In other words, subtractor 104 calculates the prediction residual (also known as error) of the current block. Subtractor 104 then outputs the calculated prediction residual to transformer 106.

[0250] The original image can be an image that has been input into encoder 100 as a signal (e.g., luminance signal and two chrominance signals) representing each picture included in the video. The signal representing the image can also be referred to as a sample.

[0251] (Transformer)

[0252] Transformer 106 transforms the prediction residuals in the spatial domain into transform coefficients in the frequency domain and outputs the transform coefficients to quantizer 108. More specifically, transformer 106 applies, for example, a defined Discrete Cosine Transform (DCT) or Discrete Sine Transform (DST) to predict the residuals in the spatial domain. The defined DCT or DST can be predefined.

[0253] It should be noted that the transformer 106 can adaptively select a transformation type from multiple transformation types and transform the prediction residuals into transformation coefficients by using transformation basis functions corresponding to the selected transformation type. This transformation is also known as explicit multi-kernel transformation (EMT) or adaptive multi-kernel transformation (AMT). The transformation basis functions can also be referred to as bases.

[0254] Transformation types include, for example, DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII. Note that these transformation types can also be represented as DCT2, DCT5, DCT8, DST1, and DST7. Figure 14 This is a graph indicating the example transformation basis functions of the example transformation type. Figure 14 In this context, N represents the number of input pixels. For example, the choice of transform type from multiple transform types can depend on the prediction type (either intra-prediction or inter-prediction) and can also depend on the intra-prediction mode.

[0255] Typically, at the CU level, signals indicating whether EMT or AMT is applied (e.g., referred to as EMT flags or AMT flags) and the selected transformation type are used to indicate this information. It should be noted that such signaling does not necessarily need to be performed at the CU level; it can also be performed at other levels (e.g., at the sequence level, picture level, slice level, tile level, or CTU level).

[0256] Furthermore, transformer 106 can re-transform the transform coefficients (which are the results of the transformation). This re-transformation is also known as adaptive quadratic transform (AST) or non-separable quadratic transform (NSST). For example, transformer 106 performs the re-transformation on a sub-block basis (e.g., a 4×4 pixel sub-block) included in the transform coefficient block corresponding to the intra-frame prediction residual. Typically, information indicating whether NSST is applied and information related to the transform matrix used for NSST are signaled at the CU level. It should be noted that such signaling does not necessarily need to be performed at the CU level; it can also be performed at other levels (e.g., at the sequence level, picture level, slice level, tile level, or CTU level).

[0257] Transformer 106 can employ both separable and non-separable transformations. A separable transformation is a method in which the transformation is performed multiple times by applying the transformation separately to each of the multiple directions according to the dimension of the input. A non-separable transformation is a method of performing a collective transformation, in which two or more dimensions of the multidimensional input are collectively treated as a single dimension.

[0258] In one example of an inseparable transformation, when the input is a 4×4 pixel block, the 4×4 pixel block is considered to be an array of 16 elements, and the transformation applies a 16×16 transformation matrix to the array.

[0259] In another example of an inseparable transformation, a 4×4 pixel input block is treated as a single array comprising 16 elements, and a transformation (hypercube given transformation) can then be performed on the array multiple times with a given rotation.

[0260] In the transformation within transformer 106, the type of transformation to be applied to the transform basis functions in the frequency domain based on the region transformation in the CU can be switched. Examples include spatial transformation (SVT).

[0261] Figure 15 This is a conceptual diagram used to illustrate an example of SVT.

[0262] In SVT, such as Figure 15 As shown, a CU is horizontally or vertically split into two equal regions, and only one of these regions is transformed into the frequency domain. The transform base type can be set for each region. For example, DST7 and DST8 can be used. For instance, in the two regions obtained by vertically splitting a CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in both regions, DST7 is used for the region at position 1. Similarly, in the two regions obtained by horizontally splitting a CU into two equal regions, DST7 and DCT8 are used for the region at position 0. Alternatively, in both regions, DST7 is used for the region at position 1. Although in Figure 15 In the example shown, only one of the two regions in the CU is transformed while the other is not, but each region in both regions can be transformed. Furthermore, the splitting method can include not only splitting into two regions, but also splitting into four regions. Moreover, the splitting method can be more flexible. For example, information indicating the splitting method can be encoded and signaled in the same way as CU splitting. It should be noted that SVT can also be called Subblock Transform (SBT).

[0263] The AMT and EMT described above can be referred to as MTS (Multiple Transform Selection). When applying MTS, transform types such as DST7 and DCT8 can be selected, and the information indicating the selected transform type can be encoded into the index information of each CU. There is another process called IMTS (Implicit MTS) for selecting the transform type used for orthogonal transforms performed without encoding the index information. When applying IMTS, for example, when the CU is rectangular, the rectangular orthogonal transform can be performed by using DST7 for the short side and DST2 for the long side. Alternatively, for example, when the CU is square, the rectangular orthogonal transform can be performed by using DCT2 when MTS is valid in the sequence and DST7 when MTS is invalid in the sequence. DCT2 and DST7 are just examples. Other transform types can be used, and the combination of transform types used can be changed to different combinations. IMTS can be used only for intra-prediction blocks, or it can be used for both intra-prediction blocks and inter-prediction blocks.

[0264] The three processes MTS, SBT, and IMTS have been described above as selection processes for selectively switching the transform type used for orthogonal transform. However, all three selection processes can be used, or only some selection processes can be selectively used. For example, one or more of these selection processes can be identified based on flag information such as in the header of SPS. For example, when all three selection processes are available, one of the three selection processes is selected for each CU and the orthogonal transform of the CU is performed. It should be noted that the selection process for selectively switching the transform type can be a different selection process from the three selection processes described above, or each of the three selection processes can be replaced by another process. Typically, at least one of the following four transfer functions [1] to [4] is executed. Function [1] is a function for performing the orthogonal transform of the entire CU and encoding information indicating the transform type used in the transform. Function [2] is a function for performing the orthogonal transform of the entire CU and determining the transform type based on determined rules without encoding information indicating the transform type. Function [3] is a function for performing the orthogonal transform of a portion of the CU and encoding information indicating the transform type used in the transform. Function [4] is used to perform orthogonal transformations on a portion of the CU and to determine the transformation type based on predetermined rules without encoding information indicating the type of transformation used in the transformation. The predetermined rules can be pre-determined.

[0265] It should be noted that the application of MTS, IMTS, and / or SBT can be determined for each processing unit. For example, the application of MTS, IMTS, and / or SBT can be determined for each sequence, image, block, slice, CTU, or CU.

[0266] It should be noted that the tools for selectively switching transformation types in this disclosure can be described as methods, selection procedures, or procedures for selectively selecting the basis used in the transformation process. Furthermore, the tools for selectively switching transformation types can be described as modes for adaptively selecting transformation types.

[0267] Figure 16 This is a flowchart illustrating an example of the process performed by converter 106, and for convenience, references will be made... Figure 7 Describe it.

[0268] For example, transformer 106 determines whether to perform an orthogonal transformation (step St_1). Here, when it is determined that an orthogonal transformation should be performed (Yes in step St_1), transformer 106 selects a transformation type for orthogonal transformation from multiple transformation types (step St_2). Next, transformer 106 performs the orthogonal transformation by applying the selected transformation type to the prediction residual of the current block (step St_3). Transformer 106 then outputs information indicating the selected transformation type to entropy encoder 110 to allow entropy encoder 110 to encode the information (step St_4). On the other hand, when it is determined that an orthogonal transformation should not be performed (No in step St_1), transformer 106 outputs information indicating that an orthogonal transformation was not performed to allow entropy encoder 110 to encode the information (step St_5). It should be noted that whether to perform an orthogonal transformation in step St_1 can be determined based on, for example, the size of the transform block, the prediction mode applied to the CU, etc. Alternatively, an orthogonal transformation can be performed using a defined transformation type without encoding information indicating the transformation type for orthogonal transformation. The defined transformation types can be predefined.

[0269] Figure 17 This is a flowchart illustrating an example of the process performed by converter 106, and for convenience, references will be made... Figure 7 A description is provided. It should be noted that... Figure 17 The example shown is in, for example Figure 16 The example shown illustrates an example of an orthogonal transformation in which the transformation type used for the orthogonal transformation is selectively switched.

[0270] As an example, the first transform type group may include DCT2, DST7, and DCT8. As another example, the second transform type group may include DCT2. The transform types included in the first transform type group and the transform types included in the second transform type group may partially overlap or may be completely different from each other.

[0271] Transformer 106 determines whether the transform size is less than or equal to a predetermined value (step Su_1). Here, when it is determined that the transform size is less than or equal to the predetermined value (Yes in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using the transform types included in the first transform type group (step Su_2). Next, transformer 106 outputs information indicating the transform type to be used from at least one of the transform types included in the first transform type group to the entropy encoder 110, allowing the entropy encoder 110 to encode this information (step Su_3). On the other hand, when it is determined that the transform size is not less than or equal to a predetermined value (No in step Su_1), transformer 106 performs an orthogonal transform on the prediction residual of the current block using the second transform type group (step Su_4). The predetermined value can be a threshold or a pre-determined value.

[0272] In step Su_3, the information indicating the transformation type used in the orthogonal transformation can be a combination of information indicating the transformation type applied vertically in the current block and the transformation type applied horizontally in the current block. The first type group may include only one transformation type and may not encode the information indicating the transformation type used for the orthogonal transformation. The second transformation type group may include multiple transformation types and may encode the information indicating the transformation type used for the orthogonal transformation that is included among one or more transformation types in the second transformation type group.

[0273] Alternatively, the transform type can be indicated based on the transform size without encoding information indicating the transform type. It should be noted that such a determination is not limited to whether the transform size is less than or equal to a predetermined value; other procedures can also be used to determine the transform type for orthogonal transforms based on the transform size.

[0274] (Quantizer)

[0275] Quantizer 108 quantizes the transform coefficients output from converter 106. More specifically, quantizer 108 scans the transform coefficients of the current block in a defined scan order and quantizes the scanned transform coefficients based on the quantization parameters (QP) corresponding to the transform coefficients. Quantizer 108 then outputs the quantized transform coefficients of the current block (hereinafter also referred to as quantized coefficients) to entropy encoder 110 and inverse quantizer 112. The defined scan order can be predetermined.

[0276] A defined scan order is the order in which the quantization / inverse quantization transform coefficients are applied. For example, a defined scan order can be defined as ascending frequency (from low to high frequency) or descending frequency (from high to low frequency).

[0277] The quantization parameter (QP) is a parameter that defines the quantization step size (quantization width). For example, when the value of the quantization parameter increases, the quantization step size also increases. In other words, when the value of the quantization parameter increases, the error of the quantized coefficients (quantization error) increases.

[0278] Furthermore, quantization matrices can be used for quantization. For example, several quantization matrices can be used accordingly for frequency transform sizes such as 4×4 and 8×8, prediction modes such as intra-frame prediction and inter-frame prediction, and pixel components such as luma and chroma pixel components. It should be noted that quantization means digitizing values ​​sampled at defined intervals corresponding to a determined level. In this art, quantization can refer to the use of other expressions, such as rounding and scaling, and rounding and scaling can be employed. The defined intervals and determined levels can be predetermined.

[0279] Methods for using a quantization matrix can include: using a quantization matrix directly set on the encoder 100 side, and using a quantization matrix that has been set to default (default matrix). On the encoder 100 side, a quantization matrix suitable for image features can be set by directly setting the quantization matrix. However, this may have the disadvantage of increasing the amount of coding required to encode the quantization matrix. It should be noted that a quantization matrix for quantizing the current block can be generated based on the default quantization matrix or the encoded quantization matrix, rather than directly using the default quantization matrix or the encoded quantization matrix.

[0280] There exists a method for quantizing high-frequency and low-frequency coefficients without using a quantization matrix. It should be noted that this method can be considered equivalent to using a quantization matrix (a flat matrix) whose coefficients have the same values.

[0281] Quantization matrices can be encoded at, for example, the sequence level, image level, slice level, block level, or CTU level. A quantization matrix can be specified using, for example, a Sequence Parameter Set (SPS) or a Picture Parameter Set (PPS). An SPS contains parameters for the sequence, while a PPS contains parameters for the image. Each of the SPS and PPS can be simply referred to as a parameter set.

[0282] When using a quantization matrix, quantizer 108 uses the values ​​of the quantization matrix to scale the quantization width for each transform coefficient, which can be calculated, for example, based on quantization parameters. A quantization process performed without using a quantization matrix can be a process of quantizing the transform coefficients based on a quantization width calculated according to quantization parameters. It should be noted that in quantization processing performed without using any quantization matrix, the quantization width can be multiplied by a predetermined value common to all transform coefficients in the block. This predetermined value can be pre-determined.

[0283] Figure 18 This is a block diagram illustrating an example of the functional configuration of a quantizer according to an embodiment. For example, quantizer 108 includes a differential quantization parameter generator 108a, a predictive quantization parameter generator 108b, a quantization parameter generator 108c, a quantization parameter memory 108d, and a quantization actuator 108e.

[0284] Figure 19 This is a flowchart illustrating an example of the quantization process performed by quantizer 108, and for convenience, references... Figure 7 and 18 Describe it.

[0285] As an example, the quantizer 108 can be based on Figure 19The flowchart shown performs quantization for each CU. More specifically, the quantization parameter generator 108c determines whether to perform quantization (step Sv_1). Here, when it is determined to perform quantization (yes in step Sv_1), the quantization parameter generator 108c generates the quantization parameters for the current block (step Sv_2) and stores the quantization parameters in the quantization parameter memory 108d (step Sv_3).

[0286] Next, the quantization executor 108e uses the quantization parameters generated in step Sv_2 to quantize the transform coefficients of the current block (step Sv_4). The predictive quantization parameter generator 108b then obtains the quantization parameters of the processing unit different from the current block from the quantization parameter memory 108d (step Sv_5). The predictive quantization parameter generator 108b generates the predicted quantization parameters of the current block based on the obtained quantization parameters (step Sv_6). The differential quantization parameter generator 108a calculates the difference between the quantization parameters of the current block generated by the quantization parameter generator 108c and the predicted quantization parameters of the current block generated by the predictive quantization parameter generator 108b (step Sv_7). The differential quantization parameters can be generated by calculating the difference. The differential quantization parameter generator 108a outputs the differential quantization parameters to the entropy encoder 110 to allow the entropy encoder 110 to encode the differential quantization parameters (step Sv_8).

[0287] It should be noted that different quantization parameters can be encoded at, for example, the sequence level, image level, slice level, block level, or CTU level. Furthermore, the initial values ​​of the quantization parameters can be encoded at the sequence level, image level, slice level, block level, or CTU level. During the initialization phase, the initial values ​​of the quantization parameters and the interpolation quantization parameters can be used to generate the quantization parameters.

[0288] It should be noted that the quantizer 108 may include multiple quantizers and may apply correlated quantization, in which a quantization method selected from a variety of quantization methods is used to quantize the transform coefficients.

[0289] (Entropy encoder)

[0290] Figure 20 This is a block diagram illustrating an example of the functional configuration of the entropy encoder 110 according to an embodiment, and for convenience, reference will be made to... Figure 7The entropy encoder 110 generates a stream by entropy encoding quantized coefficients input from quantizer 108 and prediction parameters input from prediction parameter generator 130. For example, context-based adaptive binary arithmetic coding (CABAC) is used as the entropy encoding. More specifically, the entropy encoder 110, as shown, includes a binarizer 110a, a context controller 110b, and a binary arithmetic encoder 110c. The binarizer 110a performs binarization, where a multi-level signal, such as quantized coefficients and prediction parameters, is transformed into a binary signal. Examples of binarization methods include truncated Ricean binarization, exponential Golomb codes, and fixed-length binarization. The context controller 110b derives a context value based on the characteristics of the syntax elements or the surrounding state, i.e., the probability of occurrence of the binary signal. Examples of methods for deriving the context value include bypassing, referencing syntax elements, referencing upper and left adjacent blocks, referencing hierarchical information, etc. The binary arithmetic encoder 110c uses the derived context to perform arithmetic encoding on the binary signal.

[0291] Figure 21 This is a conceptual diagram illustrating an example flow of the CABAC process in the entropy decoder 110. First, initialization is performed in the CABAC process of the entropy encoder 110. During initialization, initialization and setting of the initial context value are performed in the binary arithmetic encoder 110c. For example, the binarizer 110a and the binary arithmetic encoder 110c can sequentially perform binarization and arithmetic encoding of multiple quantized coefficients in the CTU. Each time arithmetic encoding is performed, the context controller 110b can update the context value. The context controller 110b can then save the context value for subsequent processes. For example, the saved context value can be used to initialize the context value for the next CTU.

[0292] (Inverse quantizer)

[0293] Inverse quantizer 112 inverse quantizes the quantized coefficients input from quantizer 108. More specifically, inverse quantizer 112 inverse quantizes the quantized coefficients of the current block in a defined scan order. Inverse quantizer 112 then outputs the inverse quantized transform coefficients of the current block to inverse transformer 114. The defined scan order can be predetermined.

[0294] (Inverse Transformer)

[0295] Inverse transformer 114 recovers the prediction residual by performing an inverse transform on the transform coefficients input from inverse quantizer 112. More specifically, inverse transformer 114 recovers the prediction residual of the current block by performing an inverse transform corresponding to the transform applied to the transform coefficients by transformer 106. Inverse transformer 114 then outputs the recovered prediction residual to adder 116.

[0296] It should be noted that, because information is typically lost during quantization, the recovered prediction residual does not match the prediction residual calculated by subtractor 104. In other words, the recovered prediction residual usually includes quantization error.

[0297] (Adder)

[0298] Adder 116 reconstructs the current block by adding the prediction residual input from inverse transformer 114 and the prediction image input from prediction controller 128. The reconstructed image is then generated. Adder 116 then outputs the reconstructed image to block memory 118 and cyclic filter 120. The reconstructed block can also be referred to as a local decoding block.

[0299] (Block memory)

[0300] Block memory 118 is used to store blocks in the current image, for example, for intra-frame prediction. More specifically, block memory 118 stores the reconstructed image output from adder 116.

[0301] (Frame Memory)

[0302] Frame memory 122 is, for example, a memory used to store reference images used in inter-frame prediction, and is also referred to as a frame buffer. More specifically, frame memory 122 stores reconstructed images filtered by cyclic filter 120.

[0303] (Loop Filter)

[0304] The recurrent filter 120 applies a recurrent filter to the reconstructed image output by the adder 116 and outputs the filtered reconstructed image to the frame memory 122. A recurrent filter is a filter used in the encoding loop. Examples of recurrent filters include, for example, an adaptive recurrent filter (ALF), a deblocking filter (DB or DBF), a sample adaptive offset (SAO) filter, etc.

[0305] Figure 22 This is a block diagram illustrating an example of the functional configuration of a cyclic filter 120 according to an embodiment. For example, as Figure 22As shown, the recurrent filter 120 includes a deblocking filter executor 120a, a SAO executor 120b, and an ALF executor 120c. The deblocking filter executor 120a performs deblocking filter processing on the reconstructed image. The SAO executor 120b performs the SAO process on the reconstructed image after the deblocking filter process. The ALF executor 120c performs the ALF process on the reconstructed image after the SAO process. ALF and the deblocking filter will be described in detail later. The SAO process improves image quality by reducing ringing (the phenomenon where pixel values ​​distort like waves around edges) and correcting pixel value deviations. Examples of SAO processes include edge offsetting and band offsetting. It should be noted that in some embodiments, the recurrent filter 120 may not include... Figure 22 All constituent elements disclosed herein may include some constituent elements and may include additional elements. Furthermore, the cyclic filter 120 may be configured to operate differently from... Figure 22 The above process can be executed according to the publicly disclosed processing order, or not all processes may be executed.

[0306] (Loop filter > Adaptive loop filter)

[0307] In ALF, a least-squares error filter is applied to remove compression artifacts. For example, a filter selected from multiple filters is applied for each 2×2 pixel sub-block in the current block, based on the direction and activity of the local gradient.

[0308] More specifically, first, each sub-block (e.g., each 2×2 pixel sub-block) is classified into one of several classes (e.g., fifteen or twenty-five classes). The classification of sub-blocks can be based on, for example, gradient directionality and activity. In one example, a class index C (e.g., C = 5D + A) is calculated or determined based on gradient directionality D (e.g., 0 to 2 or 0 to 4) and gradient activity A (e.g., 0 to 4). Then, based on the classification index C, each sub-block is classified into one of several classes.

[0309] For example, gradient directionality D is calculated by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). Furthermore, gradient activity A is calculated, for example, by summing the gradients in multiple directions and quantizing the sum.

[0310] Based on such classification results, the filter to be used for each sub-block can be determined from multiple filters.

[0311] The filter shape used in ALF is, for example, a circularly symmetrical filter shape. Figures 23A to 23C This is a conceptual diagram used to illustrate an example of the filter shape used in ALF. Figure 23A A 5×5 diamond filter is shown. Figure 23BA 7×7 diamond filter is shown. Figure 23C A 9×9 diamond-shaped filter is shown. Typically, information indicating the filter shape is signaled at the picture level. It should be noted that this signaling indicating the filter shape does not necessarily need to be performed at the picture level; it can also be performed at other levels (e.g., at the sequence level, slice level, tile level, or CTU level).

[0312] For example, the on / off state of ALF can be determined at the picture level or the CU level. For instance, a decision on whether to apply ALF to luminance can be made at the CU level, while a decision on whether to apply ALF to chrominance can be made at the picture level. The information indicating whether ALF is on or off is typically signaled at the picture level or the CU level. It should be noted that the signaling indicating whether ALF is on or off does not necessarily need to be executed at the picture level or the CU level; it can also be executed at other levels (e.g., at the sequence level, slice level, tile level, or CTU level).

[0313] Furthermore, as described above, a filter is selected from multiple filters, and ALF processing for the sub-block is performed. Typically, at the picture level, a set of coefficients for each of the multiple filters (e.g., up to the fifteenth or twenty-fifth filter) is signaled. It should be noted that signaling for the coefficient set does not necessarily need to be performed at the picture level; it can also be performed at other levels (e.g., at the sequence level, slice level, tile level, CTU level, CU level, or sub-block level).

[0314] (Loop Filter > Cross-Component Adaptive Loop Filter)

[0315] Figure 23D This is a conceptual diagram used to illustrate an example flow of the cross component ALF (CC-ALF). Figure 23E It is used to show in CC-ALF (e.g.) Figure 23D A conceptual diagram illustrating an example of the filter shape used in CC-ALF. Figure 23D and Figure 23E The example CC-ALF operates by applying a linear diamond filter to the luminance channel of each chromaticity component. For example, the filter coefficients can be sent in the APS, scaled by a factor of 2^10, and rounded against the fixed-point representation. For example, in... Figure 23D In this context, Y samples (the first component) are used for Cb's CCALF and CCALF's Cr (a component different from the first component).

[0316] The application of filters can be controlled on a variable block size and is signaled by a context-coded flag received for each sample block. The block size can be received at the slice level for each chroma component along with a CC-ALF enable flag. CC-ALF can support various block sizes, such as (in chroma samples) 16x16 pixels, 32x32 pixels, 64x64 pixels, and 128x128 pixels.

[0317] (Loop Filter > Joint Chromaticity Cross Component Adaptive Loop Filter)

[0318] An example of Union Chromaticity-CCALF is in Figure 23F and 23G As shown in the image. Figure 23F This is a conceptual diagram used to illustrate an example process for joint chromaticity CCALF. Figure 23G This is a table showing example weight index candidates. As shown, a CCALF filter is used to generate a CCALF filtered output as a chroma refinement signal for one color component, while a weighted version of the same chroma refinement signal is applied to another color component. This reduces the complexity of the existing CCALF by approximately half. Weight values ​​can be encoded as a symbol flag and a weight index. The weight index (denoted as weight_index) can be encoded as 3 bits and specifies the magnitude of the JC-CCALF weight JcCcWeight, which is non-zero. For example, the magnitude of JcCcWeight can be determined as follows:

[0319] If weight_index is less than or equal to 4, then JcCcWeight is equal to weight_index >> 2;

[0320] Otherwise, JcCcWeight equals 4 / (weight_index-4).

[0321] The block-level on / off control of Cb and Cr ALF filters can be separate. This is the same as in CCALF, and two separate sets of block-level on / off control flags can be encoded. Unlike CCALF, the block sizes for Cb and Cr on / off control are the same here, so only one block size variable needs to be encoded.

[0322] (Loop filter > Deblocking filter)

[0323] In the deblocking filtering process, the cyclic filter 120 performs filtering on the block boundaries in the reconstructed image to reduce distortion occurring at the block boundaries.

[0324] Figure 24 This shows a cyclic filter 120 used as a deblocking filter (see...). Figure 7 and Figure 22A block diagram of an example configuration of the deblocking filter actuator 120a.

[0325] The deblocking filter actuator 120a includes: a boundary determiner 1201; a filter determiner 1203; a filter actuator 1205; a process determiner 1208; a filter characteristic determiner 1207; and switches 1202, 1204, and 1206.

[0326] Boundary determiner 1201 determines whether the pixel to be deblocked (i.e., the current pixel) exists around the block boundary. Boundary determiner 1201 then outputs the determination result to switch 1202 and processing determiner 1208.

[0327] If boundary determiner 1201 has determined that the current pixel exists around the block boundary, switch 1202 outputs the unfiltered image to switch 1204. Conversely, if boundary determiner 1201 has determined that the current pixel does not exist around the block boundary, switch 1202 outputs the unfiltered image to switch 1206. It should be noted that the unfiltered image is an image configured with the current pixel and at least one surrounding pixel located around the current pixel.

[0328] The filter determiner 1203 determines whether to perform deblocking filtering on the current pixel based on the pixel values ​​of at least one surrounding pixel located around the current pixel. The filter determiner 1203 then outputs the determination result to the switch 1204 and the process determiner 1208.

[0329] If the filter determiner 1203 has determined that deblocking filtering will be performed on the current pixel, switch 1204 will output the unfiltered image obtained through switch 1202 to filter executor 1205. If the filter determiner 1203 has determined that no deblocking filtering will be performed on the current pixel, switch 1204 will output the unfiltered image obtained through switch 1202 to switch 1206.

[0330] When an unfiltered image is obtained via switches 1202 and 1204, filter actuator 1205 performs deblocking filtering on the current pixel, with filter characteristics determined by filter characteristic determiner 1207. Filter actuator 1205 then outputs the filtered pixel to switch 1206.

[0331] Under the control of the processing determinant 1208, the switch 1206 selectively outputs one of the pixels that have not yet been deblocked and filtered and the pixels that have been deblocked and filtered by the filter actuator 1205.

[0332] The processing determiner 1208 controls the switch 1206 based on the determination results made by the boundary determiner 1201 and the filter determiner 1203. In other words, when the boundary determiner 1201 determines that the current pixel exists near a block boundary and the filter determiner 1203 determines to perform deblocking filtering on the current pixel, the processing determiner 1208 causes the switch 1206 to output the pixel that has already been deblocked. Furthermore, in addition to the above cases, the processing determiner 1208 causes the switch 1206 to output pixels that have not been deblocked. By repeating the pixel output in this way, a filtered image is output from the switch 1206. It should be noted that... Figure 24 The configuration shown is an example of a configuration in the deblocking filter executor 120a. The deblocking filter executor 120a can have various configurations.

[0333] Figure 25 This is a conceptual diagram used to illustrate an example of a deblocking filter with symmetric filtering characteristics about the block boundaries.

[0334] During deblocking filtering, pixel values ​​and quantization parameters can be used to select one of two deblocking filters with different characteristics (i.e., a strong filter and a weak filter). In the case of a strong filter, when pixels p0 to p2 and q0 to q2 cross block boundaries, such as... Figure 25 As shown, by performing a calculation, for example, according to the following expression, the pixel values ​​of each pixel q0 to q2 are changed to pixel values ​​q'0 to q'2.

[0335] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8

[0336] q'1=(p0+q0+q1+q2+2) / 4

[0337] q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0338] It should be noted that in the above formula, p0~p2 and q0~q2 are the pixel values ​​of pixels p0~p2 and q0~q2, respectively. Additionally, q3 is the pixel value of the adjacent pixel q3 located on the opposite side of pixel q2 relative to the block boundary. Furthermore, the coefficients multiplied by the individual pixel values ​​of the pixels to be used for deblocking filtering on the right-hand side of each expression are the filter coefficients.

[0339] Furthermore, in deblocking filtering, clipping can be performed to ensure that the variation in calculated pixel values ​​does not exceed a threshold. For example, during cropping, a threshold determined based on quantization parameters can be used to crop the pixel values ​​calculated according to the above expression to values ​​obtained according to "calculated pixel value ± 2 × threshold". This prevents over-smoothing.

[0340] Figure 26It is a conceptual diagram used to illustrate the block boundaries for which the deblocking filtering process is performed. Figure 27 This is a conceptual diagram used to illustrate an example of boundary strength (Bs) values.

[0341] The block boundaries for which deblocking filtering is performed are, for example, the boundaries between CUs, Pus, or TUs with 8×8 pixel blocks, such as... Figure 26 As shown. For example, deblocking filtering can be performed in units of four rows or four columns. First, for Figure 26 Blocks P and Q are shown, as follows Figure 27 The boundary strength (Bs) value is determined as shown.

[0342] according to Figure 27 The Bs value determines whether to perform deblocking filtering on block boundaries belonging to the same image using different intensities. When the Bs value is 2, deblocking filtering is performed on the chroma signal. When the Bs value is 1 or greater and a certain condition is met, deblocking filtering is performed on the luminance signal. The determined condition can be predetermined. It should be noted that the conditions used to determine the Bs value are not limited to... Figure 27 The ones shown, and the Bs value can be determined based on another parameter.

[0343] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))

[0344] Figure 28 This is a flowchart illustrating an example of the process performed by the predictor of encoder 100. It should be noted that the predictor includes all or part of the following constituent elements: intra-frame predictor 124; predictor 126; and prediction controller 128. The prediction actuator includes, for example, intra-frame predictor 124 and inter-frame predictor 126.

[0345] The predictor generates a prediction image for the current block (step Sb_1). This prediction image can also be referred to as a prediction signal or a prediction block. Note that the prediction signal is, for example, an intra-frame prediction image (image prediction signal) or an inter-frame prediction image (inter-frame prediction signal). The predictor uses a reconstructed image to generate the prediction image for the current block, which has been obtained from another block through the generation of the prediction image, the generation of the prediction residual, the generation of the quantized coefficients, the recovery of the prediction residual, and the summation of the prediction images.

[0346] The reconstructed image can be, for example, an image in a reference image, or an image of a coded block (i.e., the other blocks mentioned above) in the current image (which is an image that includes the current block). A coded block in the current image can be, for example, a neighboring block of the current block.

[0347] Figure 29 This is a flowchart illustrating another example of the process performed by the predictor of encoder 100.

[0348] The predictor generates a predicted image using the first method (step Sc_1a), the second method (step Sc_1b), and the third method (step Sc_1c). The first, second, and third methods can be different from each other for generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.

[0349] Next, the prediction processor evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the predictor calculates the cost C for the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1, and evaluates the predicted images by comparing the cost C of the predicted images. It should be noted that the cost C can be calculated, for example, according to the expression of the RD optimization model (e.g., C = D + λ × R). In this expression, D represents the compression artifacts of the predicted image and is expressed as, for example, the sum of the absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. Furthermore, R represents the bit rate of the stream. Furthermore, λ represents, for example, the multiplier according to the Lagrange multiplier method.

[0350] The predictor then selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the predictor selects a method or mode for obtaining the final predicted image. For example, the predictor selects the predicted image with the lowest cost C based on the cost C calculated for the predicted image. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 can be based on parameters used in the encoding process. The encoder 100 can convert information used to identify the selected predicted image, method, or mode into a stream. This information can be, for example, flags. In this way, the decoder 200 can generate a predicted image based on this information according to the method or mode selected by the encoder 100. It should be noted that in Figure 29 In the example shown, after generating the predicted images using the appropriate method, the predictor selects any one of these predicted images. However, the predictor can select a method or mode based on the parameters used in the encoding process described above before generating the predicted images, and can generate the predicted images according to the selected method or mode.

[0351] For example, the first and second methods can be intra-frame prediction and inter-frame prediction, respectively, and the predictor can select the final predicted image of the current block from the predicted images generated according to the prediction method.

[0352] Figure 30This is a flowchart illustrating another example of the process executed by the prediction of encoder 100.

[0353] First, the predictor generates a prediction image using intra-frame prediction (step Sd_1a) and an inter-frame prediction (step Sd_1b). It should be noted that the prediction image generated by intra-frame prediction is also called the intra-frame prediction image, and the prediction image generated by inter-frame prediction is also called the inter-frame prediction image.

[0354] Next, the predictor evaluates each of the intra-frame and inter-frame predicted images (step Sd_2). The cost C described above can be used in the evaluation. The predictor can then select the predicted image for which the minimum cost C has been calculated from the intra-frame and inter-frame predicted images as the final predicted image for the current block (step Sd_3). In other words, the prediction method or mode used to generate the predicted image for the current block is selected.

[0355] The prediction processor then selects the prediction image for which the minimum cost C has been calculated from the intra-frame prediction image and the inter-frame prediction image as the final prediction image for the current block (step Sd_3). In other words, it selects the prediction method or mode used to generate the prediction image for the current block.

[0356] (Intra-frame predictor)

[0357] Intra-predictor 124 generates a prediction signal (i.e., an intra-predicted image) by performing intra-prediction (also referred to as intra-prediction) of the current block by referencing one or more blocks in the current image and storing it in block memory 118. More specifically, intra-predictor 124 generates an intra-predicted image by performing intra-prediction with reference to the pixel values ​​(e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block, and then outputs the intra-predicted image to prediction controller 128.

[0358] For example, intra-predictor 124 performs intra-prediction by using one of a plurality of predefined intra-prediction modes. Intra-prediction modes typically include one or more non-directional prediction modes and multiple directional prediction modes. The defined modes can be predefined.

[0359] One or more non-directional prediction modes include, for example, the planar prediction mode and the DC prediction mode as defined in the H.265 / High Efficiency Video Coding (HEVC) standard.

[0360] Multiple directional prediction modes include, for example, the thirty-three directional prediction modes defined in the H.265 / HEVC standard. It should be noted that, in addition to the thirty-three directional prediction modes, multiple directional prediction modes may also include thirty-two directional prediction modes (a total of sixty-five directional prediction modes). Figure 31This is a conceptual diagram illustrating the total 67 intra-prediction modes (two non-directional and 65 directional) that can be used in intra-prediction. Solid arrows represent the thirty-three directions defined in the H.265 / HEVC standard, while dashed arrows represent an additional thirty-two directions (the two non-directional prediction modes are not included in the standard). Figure 31 (as shown in the image).

[0361] In various processing examples, the luma block can be referenced in intra-frame prediction of the chroma block. In other words, the chroma component of the current block can be predicted based on the luma component of the current block. This intra-frame prediction is also known as cross-component linear model (CCLM) prediction. An intra-frame prediction mode for the chroma block that references this luma block (also known as, for example, CCLM mode) can be added as one of the intra-frame prediction modes for the chroma block.

[0362] Intra-predictor 124 can correct intra-predicted pixel values ​​based on horizontal / vertical reference pixel gradients. Intra-prediction accompanying this correction is also known as position-dependent intra-prediction combination (PDPC). Typically, information indicating whether PDPC is applied is signaled at the CU level (e.g., referred to as a PDPC flag). It should be noted that such signaling does not necessarily need to be performed at the CU level; it can also be performed at other levels (e.g., at the sequence level, picture level, slice level, tile level, or CTU level).

[0363] Figure 32 This is a flowchart illustrating an example of the process performed by the intra-frame predictor 124.

[0364] Intra-predictor 124 selects one intra-prediction mode from multiple intra-prediction modes (step Sw_1). Intra-predictor 124 then generates a predicted image based on the selected intra-prediction mode (step Sw_2). Next, intra-predictor 124 determines the most probable mode (MPM) (step Sw_3). The MPM includes, for example, six intra-prediction modes. For example, two of the six intra-prediction modes may be planar mode and DC prediction mode, and the other four modes may be directional prediction modes. Intra-predictor 124 determines whether the intra-prediction mode selected in step Sw_1 is included in the MPM (step Sw_4).

[0365] Here, when it is determined that the intra-prediction mode selected in step Sw_1 is included in the MPM (Yes in step Sw_4), the intra-predictor 124 sets the MPM flag to 1 (step Sw_5) and generates information indicating the intra-prediction mode selected among these MPMs (step Sw_6). It should be noted that the MPM flag set to 1 and the information indicating the intra-prediction mode can be encoded into prediction parameters by the entropy encoder 110.

[0366] When it is determined that the selected intra-prediction mode is not included in the MPM (No in step Sw_4), the intra-predictor 124 sets the MPM flag to 0 (step Sw_7). Alternatively, the intra-predictor 124 does not set any MPM flag. The intra-predictor 124 then generates information indicating the intra-prediction mode selected from at least one intra-prediction mode that is not included in the MPM (step Sw_8). It should be noted that the MPM flag set to 0 and the information indicating the intra-prediction mode can be encoded into prediction parameters by the entropy encoder 110. The information indicating the intra-prediction mode indicates, for example, any one of 0 to 60.

[0367] (Intra-frame predictor)

[0368] Inter-frame predictor 126 generates a predicted image (inter-frame predicted image) by performing inter-frame prediction (also called inter-frame prediction) of the current block by referencing one or more blocks in a reference image (which is different from the current image) and storing it in frame memory 122. Inter-frame prediction is performed on a unit of the current block or the current sub-block within the current block (e.g., a 4×4 block). A sub-block is included within a block and is a smaller unit than a block. The size of a sub-block can be in the form of a slice, block, image, etc.

[0369] For example, inter-frame predictor 126 performs motion estimation in a reference image of the current block or sub-block and finds a reference block or reference sub-block that best matches the current block or sub-block. Inter-frame predictor 126 then obtains motion information (e.g., motion vectors) to compensate for motion or variation from the reference block or reference sub-block to the current block or sub-block. Inter-frame predictor 126 generates an inter-frame predicted image of the current block or sub-block by performing motion compensation (or motion prediction) based on the motion information. Inter-frame predictor 126 outputs the generated inter-frame predicted image to prediction controller 128.

[0370] Motion information used in motion compensation can be transmitted as inter-frame prediction signals in various forms. For example, motion vectors can be signaled. As another example, the difference between motion vectors and their predicted values ​​can be signaled.

[0371] (List of reference images)

[0372] Figure 33 This is a concept diagram used to illustrate a reference image. Figure 34 This is a conceptual diagram used to illustrate an example of a list of reference images. The list of reference images is a list indicating at least one reference image stored in the frame memory 122. It should be noted that, in Figure 33In this diagram, each rectangle represents an image, each arrow represents an image reference relationship, the horizontal axis represents time, and the I, P, and B symbols within the rectangles represent intra-frame predicted images, single-predicted images, and bidirectional predicted images, respectively. The numbers within the rectangles indicate the decoding order. For example... Figure 33 As shown, the decoding order of the image is I0, P1, B2, B3, and B4, and the display order of the image is I0, B3, B2, B4, and P1. Figure 34 As shown, the reference image list is a list representing candidate reference images. For example, an image (or slice) may include at least one reference image list. For instance, one reference image list is used when the current image is a single-prediction image, and two reference image lists are used when the current image is a double-prediction image. Figure 33 and 34 In the example, image B3, which is the current image currPic, has two lists of reference images: the L0 list and the L1 list. When the current image currPic is image B3, the candidate reference images for the current image currPic are I0, P1, and B2, and the list of reference images (which is the L0 list and the L1 list) indicates these images. The inter-frame predictor 126 or the prediction controller 128 specifies which image in each list to actually reference in the form of a reference image index refidxLx. Figure 34 In the text, reference images P1 and B2 are specified by reference image indices refIdxL0 and refIdxL1.

[0373] Such a list of reference images can be generated for each unit, such as a sequence, picture, slice, block, CTU, or CU. Furthermore, among the reference images indicated in the list, reference image indices indicating which reference images will be referenced in inter-frame prediction can be signaled at the sequence level, picture level, slice level, block level, CTU level, or CU level. Additionally, a common list of reference images can be used across multiple inter-frame prediction modes.

[0374] (Basic process of inter-frame prediction)

[0375] Figure 35 This is a flowchart illustrating an example of the basic processing flow for inter-frame prediction.

[0376] First, the inter-frame predictor 126 generates a prediction signal (steps Se_1 to Se_3). Next, the subtractor 104 generates the difference between the current block and the prediction image as the prediction residual (step Se_4).

[0377] Here, in the generation of the predicted image, the inter-frame predictor 126 generates the predicted image by determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and motion compensation (step Se_3). Furthermore, in determining the MV, the inter-frame predictor 126 determines the MV by selecting motion vector candidates (MV candidates) (step Se_1) and deriving the MV (step Se_2). The selection of MV candidates is performed, for example, by the inter-frame predictor 126 generating a list of MV candidates and selecting at least one MV candidate from the list. It should be noted that previously derived MVs can be added to the MV candidate list. Alternatively, in the MV derivation, the inter-frame predictor 126 may also select at least one MV candidate from at least one MV candidate and determine the selected at least one MV candidate as the MV of the current block. Alternatively, the inter-frame predictor 126 may determine the MV of the current block by performing estimation in each specified reference image region from the selected at least one MV candidate. It should be noted that the estimation in the reference image region can be referred to as motion estimation.

[0378] In addition, although steps Se_1 to Se_3 are performed by the inter-frame predictor 126 in the above example, the processes of steps Se_1, Se_2, etc., can be performed by another constituent element included in the encoder 100.

[0379] It should be noted that an MV candidate list can be generated for each procedure in the inter-frame prediction mode, or a common MV candidate list can be used in multiple inter-frame prediction modes. The procedures in steps Se_3 and Se_4 correspond to... Figure 9 The steps Sa_3 and Sa_4 are shown. The process in step Sa_3 corresponds to... Figure 30 The process in step Sd_1b.

[0380] (Derivation process of motion vectors)

[0381] Figure 36 This is a flowchart illustrating an example of the process of deriving the motion vector.

[0382] The inter-frame predictor 126 can derive the MV of the current block from a mode (e.g., MV) used to encode motion information. In this case, for example, motion information can be encoded as prediction parameters and can be signaled. In other words, encoded motion information is included in the stream.

[0383] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where motion information is not encoded. In this case, motion information is not included in the stream.

[0384] Here, the MV derivation mode can include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., as described later. Modes in which motion information is encoded include the normal inter-frame mode, normal merging mode, and affine mode (specifically, affine inter-frame mode and affine merging mode). It should be noted that the motion information can include not only the MV but also the motion vector predictor selection information described later. Modes in which no motion information is encoded include the FRUC mode, etc. The inter-frame predictor 126 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.

[0385] Figure 37 This is a flowchart illustrating another example of the derivation of motion vectors.

[0386] The inter-frame predictor 126 can derive the MV of the current block in a mode in which the MV difference is encoded. In this case, for example, the MV difference can be encoded as prediction parameters and can be signaled. In other words, the encoded MV difference is included in the stream. The MV difference is the difference between the MV of the current block and the MV predictor. It should be noted that the MV predictor is a motion vector predictor.

[0387] Alternatively, the inter-frame predictor 126 can derive the MV in a mode where the MV difference is not encoded. In this case, the encoded MV difference is not included in the stream.

[0388] Here, as described above, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., as described later. Modes in which the MV difference is encoded include the normal inter-frame mode, affine mode (specifically, affine inter-frame mode), etc. Modes in which the MV difference is not encoded include FRUC mode, normal merging mode, affine mode (specifically, affine merging mode), etc. The inter-frame predictor 126 selects a mode from multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.

[0389] (Motion vector derivation mode)

[0390] Figure 38A and Figure 38B This is a conceptual diagram used to illustrate example classifications of patterns used for MV derivation. For example, such as... Figure 38A As shown, based on whether motion information and MV difference are encoded, MV derivation modes can be broadly categorized into three modes: inter-frame mode, merge mode, and frame rate up-conversion (FRUC) mode. Inter-frame mode performs motion estimation and encodes both motion information and MV difference within it. For example, as... Figure 38BAs shown, inter-frame modes include affine inter-frame mode and normal inter-frame mode. Merging mode is a mode in which motion estimation is not performed, and instead, a motion difference (MV) is selected from the encoded surrounding blocks, and the MV of the current block is derived using that MV. Merging mode is a mode that essentially encodes motion information without encoding the MV difference. For example, as... Figure 38B As shown, the merging modes include normal merging mode (also known as normal merging mode or regular merging mode), motion vector difference merging (MMVD) mode, inter-combination merging / intra-prediction (CIIP) mode, triangle mode, ATMVP mode, and affine merging mode. Here, in the MMVD mode, which is included among the merging modes, the MV difference is exceptionally encoded. It should be noted that affine merging mode and affine inter-frame mode are modes included within affine modes. An affine mode is one in which, assuming an affine transformation, the MV of each of the multiple sub-blocks included in the current block is derived as the MV of the current block. FRUC mode is one in which the MV of the current block is derived by performing estimation between coded regions, and in which neither motion information nor any MV difference is encoded. It should be noted that each mode will be described in more detail later.

[0391] It should be pointed out that, Figure 38A and Figure 38B The classification of modes shown is illustrative, and the classification is not limited to this. For example, when MV differences are encoded in CIIP mode, CIIP mode is classified as an inter-frame mode.

[0392] (MV derivation > Normal inter-frame mode)

[0393] The normal inter-frame mode is an inter-frame prediction mode in which the MV of the current block is derived from a reference picture region specified by the MV candidate based on blocks similar to the current block in the image. In this normal inter-frame mode, the MV difference is encoded.

[0394] Figure 39 This is a flowchart illustrating an example of the inter-frame prediction process in normal inter-frame mode.

[0395] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks surrounding the current block in time or space (step Sg_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.

[0396] Next, the inter-frame predictor 126 extracts N (N is an integer of 2 or greater) MV candidates from the plurality of MV candidates obtained in step Sg_1 as motion vector prediction candidates (also referred to as MV prediction candidates) according to the determined priority order (step Sg_2). It should be noted that the priority order can be predetermined for each of the N MV candidates.

[0397] Next, the inter-frame predictor 126 selects one of the N predicted motion vector candidates as the motion vector predictor (also called the MV predictor) for the current block (step Sg_3). At this time, the inter-frame predictor 126 encodes the predicted motion vector selection information used to identify the selected motion vector predictor in the stream. In other words, the inter-frame predictor 126 outputs the MV predictor selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.

[0398] Next, the inter-frame predictor 126 derives the MV of the current block using the reference encoded reference image (step Sg_4). At this time, the inter-frame predictor 126 also encodes the difference between the derived MV and the motion vector predictor as the MV difference within the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 via the prediction parameter generator 130. It should be noted that the encoded reference image is an image that includes multiple blocks that are reconstructed after encoding.

[0399] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the coded reference image (step Sg_5). The processes in steps Sg_1 to Sg_5 are performed for each block. For example, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in a slice, inter-frame prediction of the slice using normal inter-frame mode is completed. Similarly, when the processes in steps Sg_1 to Sg_5 are performed on all blocks in an image, inter-frame prediction of the image using normal inter-frame mode is completed. It should be noted that not all blocks included in a slice may undergo these processes; inter-frame prediction of the slice using normal inter-frame mode can be completed when only some blocks undergo these processes. This also applies to the processes in steps Sg_1 to Sg_5. When these processes are performed on only some blocks in an image, inter-frame prediction of the image using normal inter-frame mode can be completed.

[0400] It should be noted that the predicted image is the inter-frame prediction signal as described above. Furthermore, information representing the inter-frame prediction mode used to generate the predicted image (normal inter-frame mode in the example above) is encoded, for example, as prediction parameters in the coded signal.

[0401] It should be noted that the MV candidate list can also be used as a list in another mode. Furthermore, processes associated with the MV candidate list can be applied to list-related processes for use in another mode. Processes associated with the MV candidate list include, for example, extracting or selecting MV candidates from the MV candidate list, reordering MV candidates, or deleting MV candidates.

[0402] (MV derivation > Normal merge mode)

[0403] Normal merging mode is an inter-frame prediction mode that derives the MV by selecting MV candidates from the MV candidate list as the MV of the current block. It should be noted that normal merging mode is a merging mode, which can be simply referred to as merging mode. In this embodiment, normal merging mode and merging mode are distinguished, with merging mode having a wider range of applications.

[0404] Figure 40 This is a flowchart illustrating an example of inter-frame prediction in normal merging mode.

[0405] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks around the current block in time or space (step Sh_1). In other words, the inter-frame predictor 126 generates a list of MV candidates.

[0406] Next, the inter-frame predictor 126 selects an MV candidate from the multiple MV candidates obtained in step Sh_1, thereby deriving the MV of the current block (step Sh_2). At this time, the inter-frame predictor 126 encodes the MV selection information used to identify the selected MV candidate in the stream. In other words, the inter-frame predictor 126 outputs the MV selection information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.

[0407] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the coded reference image (step Sh_3). For example, the processes in steps Sh_1 to Sh_3 are performed for each block. For example, when the processes in steps Sh_1 to Sh_3 are performed on all blocks in a slice, inter-frame prediction of the slice using normal merging mode is completed. Furthermore, when the processes in steps Sh_1 to Sh_3 are performed on all blocks in an image, inter-frame prediction of the image using normal merging mode is completed. It should be noted that not all blocks included in a slice may undergo these processes in steps Sh_1 to Sh_3; when only some blocks undergo these processes, inter-frame prediction of the slice using normal merging mode can be completed. This also applies to the processes in steps Sh_1 to Sh_3. When these processes are performed on only some blocks in an image, inter-frame prediction of the image using normal merging mode can be completed.

[0408] Additionally, information representing the inter-frame prediction mode (normal merging mode in the example above) used to generate the predicted image and included in the coded signal is encoded, for example, as prediction parameters in the stream.

[0409] Figure 41 This is a conceptual diagram used to illustrate an example of the motion vector derivation process for the current image through normal merging mode.

[0410] First, the inter-frame predictor 126 generates a list of MV candidates registered therein. Examples of MV candidates include: spatially adjacent MV candidates, which are MVs of multiple coding blocks spatially surrounding the current block; temporally adjacent candidate MVs, which are MVs of surrounding blocks on which the position of the current block in the coding reference picture is projected; combined MV candidates, which are MVs generated by combining the MV values ​​of spatially adjacent MV predictors and temporally adjacent MV predictors; and zero MV candidates, which are MVs with a value of zero.

[0411] Next, the inter-frame predictor 126 selects an MV candidate from among the multiple MV candidates registered in the MV candidate list and determines that MV candidate as the MV of the current block.

[0412] In addition, the entropy encoder 110 writes and encodes merge_idx in the stream, which is a signal indicating which MV candidate has been selected.

[0413] It should be pointed out that, in Figure 41 The MV candidates registered in the MV candidate list described are examples. The number of MV candidates may differ from the number of MV candidates in the diagram. The MV candidate list can be configured in such a way that it may exclude certain types of candidate MVs from the diagram, or include one or more candidate MVs that are not in the types of candidate MVs in the diagram.

[0414] The final MV can be determined by performing Dynamic Motion Vector Refresh (DMVR), as described later, using the MV of the current block derived from the normal merge mode. It should be noted that in normal merge mode, motion information is encoded but the MV difference is not. In MMVD mode, an MV candidate is selected from the MV candidate list, and the MV difference is encoded as in the case of normal merge mode. Figure 38B As shown, MMVD can be classified as a merge mode along with the normal merge mode. It should be noted that the MV difference in MMVD mode does not always need to be the same as the MV difference used in inter-frame mode. For example, MV difference derivation in MMVD mode can be a process requiring less processing power than MV difference derivation in inter-frame mode.

[0415] In addition, a combined inter-frame merge / intra-frame prediction (CIIP) mode can be executed. This mode is used to overlay the prediction images generated in inter-frame prediction and the prediction images generated in intra-frame prediction to generate the prediction image for the current block.

[0416] It should be noted that the MV candidate list can be simply called a candidate list. Additionally, merge_idx contains MV selection information.

[0417] (MV Derivation > HMVP Pattern)

[0418] Figure 42 This is a conceptual diagram used to illustrate an example of the MV derivation process for the current image using the HMVP merging pattern.

[0419] In normal merge mode, the MV of a CU, for example, is determined by selecting an MV candidate from a list of MV candidates generated from a reference coding block (e.g., a CU). Here, another MV candidate can be registered in the MV candidate list. This mode of registering another MV candidate is called HMVP mode.

[0420] In HMVP mode, HMVP’s First-In-First-Out (FIFO) server is used to manage MV candidates, separate from the MV candidate list used in normal merge mode.

[0421] In a FIFO buffer, the latest motion information (e.g., MV) of the most recently processed block is stored first. In FIFO buffer management, for each block processed, the MV of the latest block (i.e., the previously processed CU) is stored in the FIFO buffer, and the MV of the oldest CU (i.e., the earliest processed CU) is deleted from the FIFO buffer. Figure 42 In the example shown, HMVP1 is the MV of the most recent block, and HMVP5 is the MV of the oldest block.

[0422] Then, for example, the inter-frame predictor 126 checks whether each MV managed in the FIFO buffer is a different MV from all MV candidates already registered in the MV candidate list for the normal merge mode starting from HMVP1. When it is determined that an MV is different from all MV candidates, the inter-frame predictor 126 can add the MV managed in the FIFO buffer as an MV candidate in the MV candidate list for the normal merge mode. At this time, one or more candidate MVs in the FIFO buffer can be registered (added to the MV candidate list).

[0423] By using the HMVP pattern in this way, not only can the MVs of blocks spatially or temporally adjacent to the current block be added, but also the MVs of blocks that have been processed in the past. As a result, the variation of MV candidates is expanded in the normal merge pattern, which increases the possibility of improving coding efficiency.

[0424] It should be noted that MV can be motion information. In other words, the information stored in the MV candidate list and FIFO buffer can include not only MV values, but also reference image information, reference direction, number of images, etc. Additionally, a block can be, for example, a CU.

[0425] It should be pointed out that, Figure 42 The MV candidate list and FIFO buffer shown are examples. The size of the MV candidate list and FIFO buffer can be... Figure 42 The differences in, or can be configured to be different from Figure 42 The order in which MV candidates are registered is determined. Furthermore, the process described herein can be common between encoder 100 and decoder 200.

[0426] It should be noted that the HMVP pattern can be applied to patterns other than the normal merge pattern. For example, the latest motion information (e.g., MV) such as blocks that were previously processed in the affine pattern can be stored first and used as MV candidates, which can improve efficiency. The pattern obtained by applying the HMVP pattern to the affine pattern can be called the historical affine pattern.

[0427] (MV Derivation > FRUC Pattern)

[0428] Motion information can be derived on the decoder side without being signaled from the encoder side. For example, motion information can be derived by performing motion estimation on the decoder 200 side. In one embodiment, motion estimation is performed on the decoder side without using any pixel values ​​in the current block. Modes for performing motion estimation on the decoder 200 side without using any pixel values ​​in the current block include Frame Rate Upconversion (FRUC) mode, Pattern Matching Motion Vector Derivation (PMMVD) mode, etc.

[0429] Figure 43 The document illustrates an example of a FRUC process in flowchart form. First, a list is created that indicates the MV of each coded block spatially or temporally adjacent to the current block as an MV candidate by referring to the MV (this list can be an MV candidate list or can be used as an MV candidate list for the normal merge mode (step Si_1).

[0430] Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_2). For example, the evaluation value of each MV candidate included in the candidate MV list is calculated, and an MV candidate is selected based on the evaluation value. Based on the selected motion vector candidate, the motion vector of the current block is then derived (step Si_4). More specifically, for example, the selected motion vector candidate (best MV candidate) is directly derived as the motion vector of the current block. Alternatively, for example, the motion vector of the current block can be derived using pattern matching in the region surrounding a location in the reference image, where that location in the reference image corresponds to the selected motion vector candidate. In other words, the estimation using pattern matching and evaluation values ​​can be performed in the region surrounding the best MV candidate, and when an MV with a better evaluation value exists, the best MV candidate can be updated to the MV with the better evaluation value, and the updated MV can be determined as the final MV of the current block. In some embodiments, updating the motion vector that produces a better evaluation value may not be performed.

[0431] Finally, the inter-frame predictor 126 generates a predicted image for the current block by performing motion compensation on the current block using the derived MV and the coded reference image (step Si_5). For example, the processes in steps Si_1 to Si_5 are performed for each block. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in a slice, inter-frame prediction of the slice using FRUC mode is completed. For example, when the processes in steps Si_1 to Si_5 are performed on all blocks in an image, inter-frame prediction of the image using FRUC mode is completed. It should be noted that not all blocks included in a slice may undergo these processes in steps Si_1 to Si_5; inter-frame prediction of the slice using FRUC mode can be completed when only some blocks undergo these processes. When the processing in steps Si_1 to Si_5 is performed on only some blocks included in an image in a similar manner, inter-frame prediction of the image using FRUC mode can be completed.

[0432] A similar process can be performed on a sub-block basis.

[0433] Evaluation values ​​can be calculated using various methods. For example, a comparison can be made between a reconstructed image in a region corresponding to a motion vector in a reference image and a reconstructed image in a determined region (which could be, for example, a region in another reference image or a region in a neighboring block of the current image, as shown below). The determined region can be predetermined.

[0434] The difference between pixel values ​​of two reconstructed images can be used as an evaluation value for motion vectors. It should be noted that information other than the difference can be used to calculate the evaluation value.

[0435] Next, an example of pattern matching is described in detail. First, a candidate MV included in the candidate MV list (e.g., a merge list) is selected as the starting point for estimation by pattern matching. For example, as pattern matching, either first pattern matching or second pattern matching can be used. First pattern matching and second pattern matching can be referred to as bilateral matching and template matching, respectively.

[0436] (MV derivation > FRUC > bilateral matching)

[0437] In the first pattern matching, pattern matching is performed between two blocks distributed along the motion trajectory of the current block and included in two different reference images. Therefore, in the first pattern matching, a region in another reference image along the motion trajectory of the current block is used as the region to determine the candidate evaluation values. The determined region can be predetermined.

[0438] Figure 44 This is a conceptual diagram used to illustrate an example of a first pattern matching (bilateral matching) between two blocks in two reference images along a motion trajectory. (See diagram below.) Figure 44 As shown, in the first pattern matching, two motion vectors (MV0, MV1) are derived by estimating the best-matching pair between pairs in two blocks included in two different reference images (Ref0, Ref1), and distributing these two motion vectors along the motion trajectory of the current block (Cur block). More specifically, for the current block, the difference between the reconstructed image at a specified position in the first coded reference image (Ref0) specified by the MV candidate and the reconstructed image at a specified position in the second coded reference image (Ref1) specified by the symmetric MV obtained by scaling the candidate MV at display time intervals is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value and is likely to produce good results among multiple MV candidates can be selected as the final MV.

[0439] Under the assumption of continuous motion trajectories, the motion vectors (MV0, MV1) of two reference blocks are proportional to the temporal distances (TD0, TD1) between the current image (Cur Pic) and the two reference images (Ref0, Ref1). For example, when the current image is located temporally between two reference images and the temporal distances from the current image to the corresponding two reference images are equal, a mirror-symmetric bidirectional motion vector is derived in the first pattern matching.

[0440] (MV derivation > FRUC > Template matching)

[0441] In the second pattern matching (template matching), pattern matching is performed between a block in the reference image and a template in the current image (the template is a block in the current image that is adjacent to the current block (e.g., the adjacent block is the upper adjacent block and / or the left adjacent block)). Therefore, in the second pattern matching, the blocks in the current image that are adjacent to the current block are used as the defined regions for calculating the evaluation values ​​of the aforementioned MV candidates.

[0442] Figure 45 This is a conceptual diagram used to illustrate an example of pattern matching (template matching) between a template in the current image and a block in a reference image. For example... Figure 45 As shown, in the second pattern matching, the motion vector of the current block (Cur block) is derived by estimating the best-matching block in the reference image (Ref0) that is adjacent to the current block in the current image (Cur Pic). More specifically, the difference between the reconstructed image in the left-adjacent and top-adjacent or left-adjacent and top-adjacent coded regions and the reconstructed image in the corresponding region in the coded reference image (Ref0) and specified by the MV candidate is derived, and the obtained difference is used to calculate the evaluation value. The MV candidate that produces the best evaluation value among multiple MV candidates can be selected as the best MV candidate.

[0443] Information indicating whether an FRUC mode is applied can be signaled at the CU level (e.g., referred to as the FRUC flag). Additionally, when an FRUC mode is applied (e.g., when the FRUC flag is true), information indicating the applicable pattern matching method (e.g., first pattern matching or second pattern matching) can be signaled at the CU level. It should be noted that such signaling does not necessarily need to be performed at the CU level; it can also be performed at other levels (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0444] (MV derivation > Affine mode)

[0445] Affine patterns are patterns that use affine transformations to generate motion vectors (MVs). For example, an MV can be derived on a sub-block basis based on the motion vectors of multiple neighboring blocks. This pattern is also known as an affine motion compensation prediction pattern.

[0446] Figure 46A This is a conceptual diagram used to illustrate an example of MV derivation based on motion vectors of multiple adjacent blocks, on a sub-block basis. Figure 46AIn this context, the current block comprises, for example, sixteen 4×4 sub-blocks. Here, the motion vector V0 of the top-left control point of the current block is derived based on the motion vectors of adjacent blocks, and similarly, the motion vector V1 of the top-right control point of the current block is derived based on the motion vectors of adjacent sub-blocks. The two motion vectors v0 and v1 can be projected according to the expression (1A) indicated below, and the motion vectors (v1, v2, v3, v4, v1) of each sub-block in the current block can be derived. x v y ).

[0447] [Mathematics.1]

[0448]

[0449] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w represents a predetermined weighting coefficient. This weighting coefficient can be pre-determined.

[0450] This information indicating the affine pattern can be signaled at the CU level (e.g., referred to as an affine flag). It should be noted that the signaling indicating the affine pattern does not necessarily need to be executed at the CU level; it can also be executed at other levels (e.g., at the sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0451] Furthermore, affine modes can include several modes for different methods of deriving motion vectors at the top-left and top-right control points. For example, affine modes include two modes: affine inter-frame mode (also known as affine normal inter-frame mode) and affine merge mode.

[0452] (MV derivation > Affine mode)

[0453] Figure 46B This is a conceptual diagram used to illustrate an example of MV derivation in a sub-block manner within an affine pattern that uses three control points. Figure 46B In this context, the current block comprises, for example, sixteen 4×4 blocks. Here, the motion vector V0 of the top-left control point in the current block is derived based on the motion vectors of adjacent blocks. Similarly, the motion vector V1 of the top-right control point in the current block is derived based on the motion vectors of adjacent blocks, and the motion vector V2 of the bottom-left control point in the current block is also derived based on the motion vectors of adjacent blocks. The three motion vectors v0, v1, and v2 can be projected according to the expression (1B) indicated below, and the motion vectors (v1, v2, v2) of each sub-block in the current block can be derived. x v y ).

[0454] [Mathematics.2]

[0455]

[0456] Here, x and y represent the horizontal and vertical positions of the sub-block, respectively, and w and h can be weighting coefficients, which can be predetermined weighting coefficients. In one embodiment, w can represent the width of the current block, and h can represent the height of the current block.

[0457] Affine modes using different numbers of control points (e.g., two and three control points) can be switched and signaled at the CU level. It should be noted that information indicating the number of control points in the affine mode used at the CU level can be signaled at another level (e.g., sequence level, picture level, slice level, tile level, CTU level, or sub-block level).

[0458] Furthermore, this affine mode using three control points can include different methods for deriving motion vectors at the top-left, top-right, and bottom-left control points. For example, similar to the affine mode using two control points, the affine mode using three control points can include both affine inter-frame mode and affine merge mode.

[0459] It should be noted that in affine mode, the size of each sub-block contained in the current block is not limited to 4×4 pixels; it can also be other sizes. For example, the size of each sub-block can be 8×8 pixels.

[0460] (MV derivation > Affine pattern > Control point)

[0461] Figure 47A , Figure 47B and Figure 47C This is a conceptual diagram used to illustrate an example of MV derivation at control points in an affine mode.

[0462] like Figure 47A As shown, in affine mode, for example, motion vector predictors at each control point of the current block are calculated based on multiple motion vectors corresponding to blocks encoded according to the affine mode among the coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) adjacent to the current block. More specifically, coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) are examined in the listed order, and the first valid block encoded according to the affine mode is identified. The motion vector prediction values ​​at the control points of the current block are calculated based on multiple motion vectors corresponding to the identified blocks.

[0463] For example, such as Figure 47BAs shown, when block A, which is adjacent to the left of the current block, has been encoded according to an affine pattern using two control points, motion vectors v3 and v4 projected onto the upper left and upper right corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the upper left control point of the current block and motion vector v1 at the upper right control point of the current block are calculated.

[0464] For example, such as Figure 47C As shown, when block A, which is adjacent to the left of the current block, has been encoded according to an affine pattern using three control points, motion vectors v3, v4, and v5 projected onto the upper left, upper right, and lower left corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, motion vector v0 at the upper left control point of the current block, motion vector v1 at the upper right control point of the current block, and motion vector v2 at the lower left control point of the current block are calculated.

[0465] Figures 47A to 47C The MV derivation method shown can be used for Figure 50 The MV derivation for each control point of the current block in step Sk_1 shown can also be used in the description below. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.

[0466] Figure 48A and 48B This is a conceptual diagram used to illustrate an example of MV derivation at control points in an affine mode.

[0467] Figure 48A This is a conceptual diagram used to illustrate an example affine pattern in which two control points are used.

[0468] In affine mode, such as Figure 48A As shown, the MV selected from the MVs of coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block.

[0469] Figure 48B This is a conceptual diagram illustrating an example affine pattern using three control points.

[0470] In affine mode, such as Figure 48BAs shown, the MV selected from the MVs of coded blocks A, B, and C adjacent to the current block is used as the motion vector v0 at the top-left control point of the current block. Similarly, the MV selected from the MVs of coded blocks D and E adjacent to the current block is used as the motion vector v1 at the top-right control point of the current block. Furthermore, the MV selected from the MVs of coded blocks F and G adjacent to the current block is used as the motion vector v2 at the bottom-left control point of the current block.

[0471] It should be pointed out that, Figure 48A and Figure 48B The MV derivation method shown can be used in the following descriptions. Figure 50 The MV derivation for each control point of the current block in step Sk_1 shown will be described later, or may be used in a later description. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.

[0472] Here, when affine modes with different numbers of control points (e.g., two and three control points) can be switched at the CU level and signaled, the number of control points in the coded block and the number of control points in the current block can be different from each other.

[0473] Figure 49A and Figure 49B This is a conceptual diagram illustrating an example of a method for deriving the MV at a control point when the number of control points used for the encoded block and the number of control points used for the current block are different from each other.

[0474] For example, such as Figure 49A As shown, the current block has three control points at its top left, top right, and bottom left corners. The bottom left corner and block A, which is adjacent to the left of the current block, have been encoded according to an affine pattern, using two control points. In this case, motion vectors v3 and v4 projected onto the top left and top right corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3 and v4, motion vector v0 at the top left control point and motion vector v1 at the top right control point of the current block are calculated. Furthermore, motion vector v2 at the bottom left control point is calculated based on the derived motion vectors v0 and v1.

[0475] For example, such as Figure 49BAs shown, the current block has two control points at its top left and top right corners, and the top right corner and the block A adjacent to the left of the current block have been encoded according to an affine pattern, which uses two control points. In this case, motion vectors v3, v4, and v5 projected onto the top left, top right, and bottom left corners of the encoded block including block A are derived. Then, based on the derived motion vectors v3, v4, and v5, the motion vector v0 at the top left control point of the current block and the motion vector v1 at the top right control point of the current block are calculated.

[0476] It should be pointed out that, Figure 49A and Figure 49B The MV derivation method shown can be used in the following descriptions. Figure 50 The MV derivation for each control point of the current block in step Sk_1 shown will be described later, or may be used in a later description. Figure 51 The MV predictor derivation at each control point of the current block in step Sj_1 is shown.

[0477] (MV Derivation > Affine Pattern > Affine Merging Pattern)

[0478] Figure 50 This is a flowchart illustrating an example of the process in the affine merge pattern.

[0479] In the affine merging mode shown in the figure, firstly, the inter-frame predictor 126 derives the MV at each control point of the current block (step Sk_1). The control points are the top-left corner and the top-right corner of the current block, such as... Figure 46A As shown, or the top-left corner, top-right corner, and bottom-left corner of the current block, as shown. Figure 46B As shown. The inter-frame predictor 126 can encode MV selection information to identify two or three derived MVs in the stream.

[0480] For example, when using Figures 47A to 47C When using the MV derivation method shown, such as Figure 47A As shown, the inter-frame predictor 126 examines coded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) and identifies the first valid block encoded according to the affine pattern.

[0481] Inter-frame predictor 126 uses the identified first valid block encoded according to the identified affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, such as... Figure 47BAs shown, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the coded block, including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block by projecting the motion vectors v3 and v4 at the top-left and top-right control points of the coded block onto the current block.

[0482] Alternatively, when block A is identified and block A has three control points, such as Figure 47C As shown, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block, the motion vector v1 at the top-right control point of the current block, and the motion vector v2 at the bottom-left control point of the current block based on the motion vectors v3, v4, and v5 located at the top-left, top-right, and bottom-left corners of the coding block, including block A. For example, the inter-frame predictor 126 calculates the motion vector v0 at the top-left control point of the current block, the motion vector v1 at the top-right control point of the current block, and the motion vector v2 at the bottom-left control point of the current block by projecting the motion vectors v3, v4, and v5 located at the top-left, top-right, and bottom-left corners of the coding block onto the current block.

[0483] It should be pointed out that, as mentioned above Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, as described above. Figure 49B As shown, when block A is identified and block A has three control points, the MV at two control points can be calculated.

[0484] Next, the inter-frame predictor 126 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 126 calculates the MV of each of the multiple sub-blocks as an affine MV, for example, using two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) (step Sk_2). The inter-frame predictor 126 then uses these affine MVs and the encoded reference image to perform motion compensation for the sub-blocks (step Sk_3). While performing the processes in steps Sk_2 and Sk_3 for each of all the sub-blocks included in the current block, the process of generating a predicted image for the current block using an affine merging pattern ends. In other words, motion compensation for the current block is performed to generate a predicted image for the current block.

[0485] It should be noted that the aforementioned MV candidate list can be generated in step Sk_1. The MV candidate list can, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods can, for example, be... Figures 47A to 47C The MV derivation method shown Figure 48A and 48B The MV derivation method shown Figure 49A and 49B The MV derivation method shown can be any combination of other MV derivation methods.

[0486] It should be noted that, in addition to affine patterns, the MV candidate list may include MV candidates from patterns that perform prediction on a sub-block basis.

[0487] It should be noted that, for example, an MV candidate list including MV candidates can be generated in both an affine merge mode using two control points and an affine merge mode using three control points. Alternatively, an MV candidate list including MV candidates in the affine merge mode using two control points and an MV candidate list including MV candidates in the affine merge mode using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates can be generated in either an affine merge mode using two control points or an affine merge mode using three control points. MV candidates can be, for example, MVs used to encode block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left), or MVs of valid blocks among these blocks.

[0488] It should be noted that the index of one of the MVs in the MV candidate list can be sent as MV selection information.

[0489] (MV Derivation > Affine Mode > Affine Inter-Frame Mode)

[0490] Figure 51 This is a flowchart illustrating an example of a process in affine inter-frame mode.

[0491] In affine inter-frame mode, firstly, the inter-frame predictor 126 derives the MV prediction values ​​(v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_1). For example... Figure 46A or Figure 46B As shown, control points can be, for example, the top-left corner of the current block, the top-right corner of the current block, and the bottom-left corner of the current block.

[0492] For example, when using Figure 48A and 48B When the MV derivation method shown is used, the inter-frame predictor 126, through... Figure 48A or Figure 48BThe MV of any block is selected from the coded blocks near each control point of the current block, and the MV predictor (v0, v1) or (v0, v1, v2) is derived at the corresponding two or three control points of the current block. At this time, the inter-frame predictor 126 encodes the MV predictor selection information in the stream to identify the two or three selected MV predictors.

[0493] For example, the inter-frame predictor 126 can use cost evaluation or the like to determine from the coded blocks adjacent to the current block the block from which an MV predictor has been selected as a control point, and can write a flag in the bitstream indicating which MV predictor has been selected. In other words, the inter-frame predictor 126 outputs MV predictor selection information (e.g., flags) as prediction parameters to the entropy encoder 110 via the prediction parameter generator 130.

[0494] Next, the inter-frame predictor 126 performs motion estimation (steps Sj_3 and Sj_4) while updating the MV predictor selected or derived in step Sj_1 (step Sj_2). In other words, the inter-frame predictor 126 uses the above expression (1A) or expression (1B) to calculate the MV of each sub-block corresponding to the updated MV predictor as an affine MV (step Sj_3). The inter-frame predictor 126 then uses these affine MVs and the encoded reference picture to perform motion compensation for the sub-blocks (step Sj_4). When the MV predictor is updated in step Sj_2, the process in steps Sj_3 and Sj_4 is performed for all blocks in the current block. As a result, for example, the inter-frame predictor 126 determines the MV predictor that produces the minimum cost as the MV at the control point in the motion estimation loop (step Sj_5). At this time, the inter-frame predictor 126 also encodes the difference between the determined MV and the MV predictor as an MV difference in the stream. In other words, the inter-frame predictor 126 outputs the MV difference as a prediction parameter to the entropy encoder 110 through the prediction parameter generator 130.

[0495] Finally, the inter-frame predictor 126 generates a predicted image of the current block by performing motion compensation for the current block using the determined MV and coded reference image (step Sj_6).

[0496] It should be noted that the aforementioned MV candidate list can be generated in step Sj_1. The MV candidate list may, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods may, for example, be... Figures 47A to 47C The MV derivation method shown Figure 48A and 48B The MV derivation method shown Figure 49A and 49B The MV derivation method shown can be any combination of other MV derivation methods.

[0497] It should be noted that, in addition to affine patterns, the MV candidate list may include MV candidates from patterns that perform prediction on a sub-block basis.

[0498] It should be noted that, for example, an MV candidate list including MV candidates can be generated in both an affine inter-frame mode using two control points and an affine inter-frame mode using three control points. Alternatively, an MV candidate list including MV candidates in the affine inter-frame mode using two control points and an MV candidate list including MV candidates in the affine inter-frame mode using three control points can be generated respectively. Alternatively, an MV candidate list including MV candidates can be generated in either an affine inter-frame mode using two control points or an affine inter-frame mode using three control points. MV candidates can be, for example, MVs used to encode block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left), or MVs of valid blocks among these blocks.

[0499] It should be noted that the index indicating one of the MV candidates in the MV candidate list can be sent as MV predictor selection information.

[0500] (MV Derivation > Triangle Pattern)

[0501] In the example above, the inter-frame predictor 126 generates a rectangular prediction image for the current rectangular block. However, the inter-frame predictor 126 can generate multiple prediction images, each with a shape different from the rectangle of the current rectangular block, and can combine multiple prediction images to generate a final rectangular prediction image. The shape different from the rectangle can be, for example, a triangle.

[0502] Figure 52A This is a conceptual diagram used to illustrate the generation of two triangular predicted images.

[0503] Inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a first partition with a triangular shape in the current block using a first MV of a first partition. Similarly, inter-frame predictor 126 generates a triangle prediction image by performing motion compensation on a second partition with a triangular shape in the current block using a second MV of a second partition. Inter-frame predictor 126 then combines these prediction images to generate a prediction image with a rectangular shape identical to that of the current block.

[0504] It should be noted that a first prediction image with a rectangular shape corresponding to the current block can be generated using a first prediction image (MV) as the prediction image for the first partition. Furthermore, a second prediction image with a rectangular shape corresponding to the current block can be generated using a second prediction image (MV) as the prediction image for the second partition. The prediction image for the current block can be generated by performing a weighted sum of the first and second prediction images. It should be noted that the weighted sum may be a region that crosses the boundary between the first and second partitions.

[0505] Figure 52B This is a conceptual diagram illustrating an example of a first portion of a first partition that overlaps with a second partition, and first and second sets of samples that can be weighted as part of a correction process. The first portion may be, for example, one-quarter of the width or height of the first partition. In another example, the first portion may have a width corresponding to N samples adjacent to the edges of the first partition, where N is a positive integer, for example, N could be an integer 2. As shown in the figure, Figure 52B The example on the left shows a rectangular partition whose width is one-quarter of the width of the first partition, where the first group of samples includes samples outside the first part and samples inside the first part, and the second group of samples includes samples inside the first part. Figure 52B The central example shows a rectangular partition whose height is one-quarter of the height of the first partition, wherein the first group of samples includes samples outside the first part and samples inside the first part, and the second group of samples includes samples inside the first part. Figure 52B The example on the right is a triangular partition except for the polygonal part, whose height corresponds to two samples, where the first set of samples includes samples outside the first part and samples inside the first part, and the second set of samples includes samples inside the first part.

[0506] The first part can be the portion of the first partition that overlaps with the adjacent partition. Figure 52C This is a conceptual diagram illustrating a first portion of a first partition, which is the part of the first partition that overlaps with a portion of an adjacent partition. For ease of illustration, a rectangular partition with an overlapping portion that is spatially adjacent to a rectangular partition is shown. Partitions with other shapes can be used, such as triangular partitions, and the overlapping portion can overlap with spatially or temporally adjacent partitions.

[0507] Furthermore, although an example of generating a predicted image for each of two partitions using inter-frame prediction is given, a predicted image for at least one partition can be generated using intra-frame prediction.

[0508] Figure 53 This is a flowchart illustrating an example of a process in a triangle pattern.

[0509] In the triangle mode, firstly, the inter-frame predictor 126 splits the current block into a first partition and a second partition (step Sx_1). At this time, the inter-frame predictor 126 can encode the partition information, which is related to the splitting into partitions, into prediction parameters in the stream. In other words, the inter-frame predictor 126 can output the partition information as prediction parameters to the entropy encoder 110 through the prediction parameter generator 130.

[0510] First, the inter-frame predictor 126 obtains multiple MV candidates for the current block based on information such as the MVs of multiple coded blocks around the current block in time or space (step Sx_2). In other words, the inter-frame predictor 126 generates a list of MV candidates.

[0511] Inter-frame predictor 126 then selects the MV candidate of the first partition and the MV candidate of the second partition from the plurality of MV candidates obtained in step Sx_1 as the first MV and the second MV, respectively (step Sx_3). At this time, inter-frame predictor 126 encodes the MV selection information used to identify the selected MV candidate as prediction parameters in the stream. In other words, inter-frame predictor 126 outputs the MV selection information as prediction parameters to entropy encoder 110 through prediction parameter generator 130.

[0512] Next, the inter-frame predictor 126 performs motion compensation using the selected first MV and coded reference image to generate a first predicted image (step Sx_4). Similarly, the inter-frame predictor 126 performs motion compensation using the selected second MV and coded reference image to generate a second predicted image (step Sx_5).

[0513] Finally, the inter-frame predictor 126 generates the prediction image for the current block by performing a weighted sum of the first and second prediction images (step Sx_6).

[0514] It should be pointed out that, although in Figure 52A In the example shown, the first and second partitions are triangles, but the first and second partitions could be trapezoids, or other shapes that are different from each other. Furthermore, although in Figure 52A and 52C The example shown contains two partitions, but the current block can contain three or more partitions.

[0515] Furthermore, the first and second partitions can overlap. In other words, the first and second partitions can include the same pixel region. In this case, the predicted image of the current block can be generated using the predicted image from the first and second partitions.

[0516] Furthermore, although an example of generating a predicted image for each of two partitions using inter-frame prediction has been shown, a predicted image for at least one partition can be generated using intra-frame prediction.

[0517] It should be noted that the MV candidate list used to select the first MV and the MV candidate list used to select the second MV can be different from each other, or the MV candidate list used to select the first MV can also be used as the MV candidate list used to select the second MV.

[0518] It should be noted that partition information may include indexes indicating the split direction in which at least the current block is split into multiple partitions. MV selection information may include indexes indicating the first selected MV and indexes indicating the second selected MV. One index may indicate multiple pieces of information. For example, an index that jointly indicates part or all of the partition information and part or all of the MV selection information may be encoded.

[0519] (MV Derivation > ATMVP Pattern)

[0520] Figure 54 This is a conceptual diagram used to illustrate an example of an advanced temporal motion vector prediction (ATMVP) pattern in which the MV is derived on a sub-block basis.

[0521] The ATMVP pattern is a pattern categorized as a merge pattern. For example, in the ATMVP pattern, the MV candidate for each sub-block is registered in the MV candidate list for use in the normal merge pattern.

[0522] More specifically, in the ATMVP pattern, firstly, as Figure 54 As shown, a temporal MV reference block associated with the current block is identified in the encoded reference image specified by the MV (MV0) of the adjacent block located at the lower left position relative to the current block. Next, in each sub-block within the current block, an MV is identified for encoding the corresponding region of the sub-block in the temporal MV reference block. MVs identified in this manner are included in the MV candidate list as MV candidates for the sub-blocks in the current block. When selecting an MV candidate for each sub-block from the MV candidate list, motion compensation is performed on the sub-block, where the MV candidate is used as the MV of the sub-block. This generates a predicted image for each sub-block.

[0523] Despite Figure 54 In the example shown, the block located at the bottom left relative to the current block is used as the surrounding MV reference block. It should be noted that another block can be used. Furthermore, the size of the child block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the child block can be toggled in units of slices, blocks, images, etc.

[0524] (MV derivation > DMVR)

[0525] Figure 55 This is a flowchart illustrating the relationship between the merging mode and the decoded motion vector refinement DMVR.

[0526] Inter-frame predictor 126 derives the motion vector of the current block based on the merging mode (step S1_1). Next, inter-frame predictor 126 determines whether to perform motion vector estimation, i.e., motion estimation (step S1_2). Here, when it is determined that motion estimation should not be performed (no in step S1_2), inter-frame predictor 126 determines the motion vector derived in step S1_1 as the final motion vector of the current block (step S1_4). In other words, in this case, the motion vector of the current block is determined according to the merging mode.

[0527] When motion estimation is determined to be performed in step Sl_1 (Yes in step Sl_2), the inter-frame predictor 126 derives the final motion vector of the current block by estimating the region surrounding the reference image specified by the motion vector derived in step Sl_1 (step Sl_3). In other words, in this case, the motion vector of the current block is determined according to the DMVR.

[0528] Figure 56 This is a conceptual diagram illustrating an example of the DMVR process used to determine the MV.

[0529] First, for example in merge mode, MV candidates (L0 and L1) are selected for the current block. Reference pixels are identified from the first reference image (L0) (which is an coded image in the L0 list) based on the MV candidates (L0). Similarly, reference pixels are identified from the first reference image (L1) (which is an coded image in the L1 list) based on the MV candidates (L1). A template is generated by calculating the average of these reference pixels.

[0530] Next, a template is used to estimate the surrounding regions of each MV candidate for the first reference image (L0) and the second reference image (L1), and the MV that produces the minimum cost is determined as the final MV. It should be noted that the cost can be calculated, for example, using the difference between each pixel value in the template and the corresponding pixel value in the estimated region, the value of the MV candidate, etc.

[0531] It is not always necessary to perform the exact same procedure described here. Other procedures can be used to derive the final MV by estimating the region surrounding the MV candidate.

[0532] Figure 57 This is a conceptual diagram used to illustrate another example of DMVR for determining MV. Figure 56 The DMVR example shown is different, in Figure 57 In the example shown, the cost is calculated without generating a template.

[0533] First, the inter-frame predictor 126 estimates the surrounding region of the reference block in each reference image included in the L0 and L1 lists based on the initial MV (which is the MV candidate obtained from each MV candidate list). For example, as Figure 57 As shown, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1. In motion estimation, the inter-frame predictor 126 first sets the search position for the reference images in the L0 list. Based on the position indicated by the vector difference indicating the search position, specifically the initial MV (i.e., InitMV_L0, with a vector difference of MVd_L0 from the search position). The inter-frame predictor 126 then determines the estimated position in the reference images in the L1 list. This search position is represented by the vector difference from the search position indicated by the initial MV (i.e., InitMV_L1). More specifically, the inter-frame predictor 126 determines the vector difference as MVd_L1 by mirroring MVd_L0. In other words, the inter-frame predictor 126 determines the search position in each reference image in the L0 and L1 lists as the position symmetric about the position indicated by the initial MV. The inter-frame predictor 126 calculates the sum of the absolute differences (SAD) between pixel values ​​at the search location in the block as the cost for each search location, and finds the search location that produces the minimum cost.

[0534] Figure 58A This is a conceptual diagram used to illustrate an example of motion estimation in DMVR, and Figure 58B This is a flowchart illustrating an example of the motion estimation process.

[0535] First, in step 1, the inter-frame predictor 126 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and eight surrounding search positions. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at a search position other than the starting point is determined to be minimized, the inter-frame predictor 126 changes its objective to obtaining the search position with the minimum cost and executes the process in step 2. When the cost at the starting point is minimized, the inter-frame predictor 126 skips the process in step 2 and executes the process in step 3.

[0536] In step 2, the inter-frame predictor 126 performs a search similar to that in step 1, using the changed search position as the new starting point based on the result of the process in step 1. The inter-frame predictor 126 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at each search position other than the starting point is minimized, the inter-frame predictor 126 performs the process in step 4. When the cost at the starting point is minimized, the inter-frame predictor 126 performs the process in step 3.

[0537] In step 4, the inter-frame predictor 126 considers the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0538] In step 3, the inter-frame predictor 126 determines the pixel position based on the cost at four points (up, down, left, and right) relative to the starting point in step 1 or step 2, to obtain the sub-pixel precision with the minimum cost, and treats the pixel position as the final search position. Each of the four vectors ((0,1),(0,-1),(-1,0), and (1,0)) is weighted and summed, using the cost at the corresponding position in the four search positions as weights, to determine the pixel position with sub-pixel precision. The inter-frame predictor 126 then determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0539] (Motion compensation > BIO / OBMC / LIC)

[0540] Motion compensation involves modes used to generate and correct predicted images. Modes such as bidirectional optical flow (BIO), overlapping block motion compensation (OBMC), and local illumination compensation (LIC) will be described later.

[0541] Figure 59 This is a flowchart illustrating an example of the process of generating a predicted image.

[0542] Inter-frame predictor 126 generates a predicted image (step Sm_1), and, for example, corrects the predicted image according to any of the above-described modes (step Sm_2).

[0543] Figure 60 This is a flowchart illustrating another example of the process of generating a predicted image.

[0544] Inter-frame predictor 126 determines the motion vector of the current block (step Sn_1). Next, inter-frame predictor 126 generates a predicted image using the motion vector (step Sn_2) and determines whether to perform a correction process (step Sn_3). Here, when it is determined that a correction process should be performed (Yes in step Sn_3), inter-frame predictor 126 generates the final predicted image by correcting the predicted image (step Sn_4). It should be noted that in the LIC described later, luminance and chrominance can be corrected in step Sn_4. When it is determined that a correction process should not be performed (No in step Sn_3), inter-frame predictor 126 outputs the predicted image as the final predicted image without correcting the predicted image (step Sn_5).

[0545] (Motion compensation > OBMC)

[0546] It should be noted that, in addition to the motion information of the current block obtained through motion estimation, motion information from neighboring blocks can also be used to generate inter-frame prediction images. More specifically, by weighted summing of the prediction image based on motion information obtained through motion estimation (in the reference image) and the prediction image based on motion information from neighboring blocks (in the current image), an inter-frame prediction image can be generated for each sub-block in the current block. This inter-frame prediction (motion compensation) is also known as Overlapping Block Motion Compensation (OBMC) or OBMC mode.

[0547] In OBMC mode, information indicating the sub-block size used for OBMC can be signaled at the sequence level (e.g., referred to as the OBMC block size). Additionally, information indicating whether OBMC mode is applied can be signaled at the CU level (e.g., referred to as the OBMC flag). It should be noted that signaling for such information does not necessarily need to be performed at the sequence and CU levels; it can also be performed at other levels (e.g., at the picture, slice, tile, CTU, or sub-block level).

[0548] The OBMC model will be described in more detail. Figure 61 and Figure 62 These are flowcharts and conceptual diagrams used to illustrate an outline of the predictive image correction process performed by OBMC.

[0549] First, such as Figure 62 As shown, a normal motion-compensated predicted image (Pred) is obtained using the MV allocated to the current block. Figure 62 In the image, the arrow "MV" points to the reference image and indicates what the current block of the current image references in order to obtain the predicted image.

[0550] Next, a predicted image (Pred_L) is obtained by applying the motion vector (MV_L) derived for the coded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The motion vector (MV_L) is indicated by the arrow "MV_L", which points to the reference image from the current block. The first correction of the predicted image is performed by overlapping the two predicted images, Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.

[0551] Similarly, a predicted image (Pred_U) is obtained by applying the MV (MV_U) already derived for the coded block adjacent to the current block (reusing the MV of the current block) to the current block. MV (MV_U) is indicated by the arrow "MV_U", which points to the reference image from the current block. A second correction to the predicted image is performed by overlaying the predicted image Pred_U onto a predicted image (e.g., Pred and Pred_L) to which the first correction has already been performed. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained from the second correction is an image where the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.

[0552] Although the above example uses a two-path correction method with the left and top adjacent blocks, it should be noted that the correction method can be a three-path or more-path correction method that uses the right and / or bottom adjacent blocks simultaneously.

[0553] It should be noted that the area where this overlap occurs can be only a part of the area near the block boundary, rather than the entire pixel area of ​​the block.

[0554] It should be noted that the above has described the prediction image correction process based on OBMC for obtaining a prediction image Pred from a reference image by overlaying additional prediction images Pred_L and Pred_U. However, when correcting prediction images based on multiple reference images, a similar process can be applied to each of the multiple reference images. In this case, after obtaining corrected prediction images from each reference image by performing OBMC image correction based on multiple reference images, the obtained corrected prediction images are further overlaid to obtain the final prediction image.

[0555] It should be noted that in OBMC, the current block unit can be a PU or a sub-block unit obtained by further subdividing a PU.

[0556] One example of a method for determining whether to apply OBMC is using the `obmc_flag` method, where `obmc_flag` is a signal indicating whether OBMC is applied. As a concrete example, encoder 100 can determine whether the current block belongs to a region with complex motion. Encoder 100 sets `obmc_flag` to a value of "1" and applies OBMC during encoding when the block belongs to a region with complex motion; when the block does not belong to a region with complex motion, it sets `obmc_flag` to a value of "0" and encodes the block without applying OBMC. Decoder 200 switches between applying and not applying OBMC by decoding the `obmc_flag` in the write stream.

[0557] (Motion compensation > BIO)

[0558] Next, the derivation method for MV is described. First, the mode used to derive MV based on a model assuming uniform linear motion is described. This mode is also known as the bidirectional optical flow (BIO) mode. Furthermore, this bidirectional optical flow can be written as BDOF instead of BIO.

[0559] Figure 63 This is a conceptual diagram used to illustrate a model that assumes uniform linear motion. Figure 63 In the middle, (v x v y Let τ0 represent the velocity vector, and τ1 and τ0 represent the time distance between the current image (Cur Pic) and the two reference images (Ref0, Ref1). x0 MV y0 () represents the MV corresponding to the reference image Ref0, (MV x1 MV y1 ) indicates the MV corresponding to the reference image Ref1.

[0560] Here, it is assumed that the velocity vector (v) x ,v y It exhibits uniform linear motion, (MV) x0 MV y0 ) and (MV x1 MV y1 ) are respectively represented as (v xτ0 ,v yτ0 ) and (-v xτ1 ,-v yτ1 ), and the following optical flow equation (2) is given.

[0561] [Mathematics.3]

[0562]

[0563] Here, I(k) represents the motion-compensated luminance value of the motion-compensated reference image k (k = 0, 1). The optical flow equation indicates that (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image are all zero. Based on the combination of the optical flow equation and Hermite interpolation, the motion vector of each block obtained from, for example, an MV candidate list can be corrected pixel by pixel.

[0564] It should be noted that motion vectors can be derived on the decoder side using methods different from those used to derive motion vectors based on the assumption of uniform linear motion. For example, motion vectors can be derived on a sub-block basis based on the motion vectors of multiple adjacent blocks.

[0565] Figure 64 This is a flowchart illustrating an example of the inter-frame prediction process based on BIO. Figure 65 This is a functional block diagram illustrating an example of the functional configuration of an inter-frame predictor 126 that can perform inter-frame prediction based on BIO.

[0566] like Figure 65 As shown, the inter-frame predictor 126 includes, for example, a memory 126a, a difference image deducer 126b, a gradient image deducer 126c, an optical flow deducer 126d, a correction value deducer 126e, and a predicted image corrector 126f. It should be noted that the memory 126a may be a frame memory 122.

[0567] Inter-frame predictor 126 derives two motion vectors (M0, M1) using two reference images (Ref0, Ref1) that are different from the image (Cur Pic) including the current block. Inter-frame predictor 126 then uses the two motion vectors (M0, M1) to derive the predicted image for the current block (step Sy_1). It should be noted that motion vector M0 is the motion vector (MV) corresponding to the reference image Ref0. x0 MV y0 The motion vector M1 is the motion vector (MV) corresponding to the reference image Ref1. x1 MV y1 ).

[0568] Next, the difference image derivator 126b uses the motion vector M0 and the reference image L0 through the reference memory 126a to derive the difference image I of the current block. 0 Next, the difference image derivator 126b uses the motion vector M1 and the reference image L1 from the reference memory 126a to derive the difference image I of the current block. 1 (Step Sy_2). Here, the difference image I 0 It is the image included in the reference image Ref0 and to be derived for the current block, and the difference image I 1 It is the image included in reference image Ref1 and to be derived for the current block. Difference image I 0 Sum and difference image I 1 Each of the elements in the image can be the same size as the current block. Alternatively, the difference image I... 0 Sum and difference image I 1 Each of these can be an image larger than the current block. Furthermore, the difference image I... 0 Sum and difference image I 1 This can include a predicted image obtained by using motion vectors (M0, M1) and a reference image (L0, L1) and applying a motion compensation filter.

[0569] Additionally, the gradient image deriver 126c derives from the difference image I 0 Sum and difference image I 1 Derive the gradient image (Ix) of the current block. 0 、Ix 1 、Iy 0 、Iy 1 (Step Sy_3). It should be noted that the gradient image in the horizontal direction is (Ix 0 , Ix 1 The gradient image in the vertical direction is (Iy). 0 ,Iy 1 The gradient image derivator 126c can derive each gradient image by, for example, applying a gradient filter to a difference image. The gradient image can indicate the amount of spatial variation of pixel values ​​along the horizontal direction, along the vertical direction, or along both directions.

[0570] Next, the optical flow deriver 126d uses an interpolated image (I 0 I 1 ) and gradient image (Ix 0 、Ix 1 、Iy 0 、Iy 1 For each sub-block of the current block, the optical flow (vx, vy) is derived as a velocity vector (step Sy_4). The optical flow indicates the coefficients used to correct the spatial pixel movement and can be referred to as a local motion estimate, a corrected motion vector, or a corrected weighted vector. As an example, a sub-block can be a 4×4 pixel sub-CU. It should be noted that the optical flow derivation can be performed on a per-pixel unit, etc., rather than on a per-sub-block basis.

[0571] Next, the inter-frame predictor 126 uses optical flow (vx, vy) to correct the predicted image of the current block. For example, the correction value derivator 126e uses optical flow (vx, vy) to derive correction values ​​for the pixel values ​​included in the current block (step Sy_5). The predicted image corrector 126f can then use the correction values ​​to correct the predicted image of the current block (step Sy_6). It should be noted that the correction values ​​can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis.

[0572] It should be noted that the BIO process is not limited to Figure 64 The process is publicly disclosed. For example, it may be possible to execute only the... Figure 64 The processes disclosed in the document may be modified, or different processes may be added or used as alternatives, or these processes may be executed in different processing orders, and so on.

[0573] (Motion compensation > LIC)

[0574] Next, an example of a pattern used to generate a predicted image (prediction) using the Local Illumination Compensation (LIC) process is described.

[0575] Figure 66A This is a conceptual diagram illustrating an example of a predictive image generation method that uses a brightness correction process performed by a LIC. Figure 66B This is a flowchart illustrating an example of a process for generating a predicted image using LIC.

[0576] First, the inter-frame predictor 126 derives the MV from the coded reference image and obtains the reference image corresponding to the current block (step Sz_1).

[0577] Next, the inter-frame predictor 126 extracts information indicating how the luminance values ​​change between the current block and the reference image (step Sz_2). This extraction is performed based on the luminance pixel values ​​of the encoded left adjacent reference region (surrounding reference region) and the encoded upper adjacent reference region (surrounding reference region) in the current image, as well as the luminance pixel values ​​at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 126 uses the information indicating how the luminance values ​​change to calculate the luminance correction parameters (step Sz_3).

[0578] The inter-frame predictor 126 generates a predicted image for the current block by performing a luminance correction process in which luminance correction parameters are applied to a reference image in a reference picture specified by the MV (step Sz_4). In other words, the predicted image (which is a reference image in a reference picture specified by the MV) is corrected based on the luminance correction parameters. In this correction, luminance can be corrected, chrominance can be corrected, or both can be corrected. In other words, chrominance correction parameters can be calculated using information indicating how chrominance changes, and a chrominance correction process can be performed.

[0579] It should be pointed out that, Figure 66A The shape of the surrounding reference area shown is an example; another shape can be used.

[0580] Furthermore, although the process of generating a prediction image from a single reference image is described herein, the case of generating a prediction image from multiple reference images can be described in the same manner. The prediction image can be generated after performing a brightness correction process on the reference image obtained from the reference image in the same manner as described above.

[0581] An example of a method for determining whether to apply a LIC is using a `lic_flag`, which is a signal indicating whether the LIC is applied. As a specific example, encoder 100 determines whether the current block belongs to a region with brightness variations. Encoder 100 sets `lic_flag` to a value of "1" and applies the LIC during encoding when the block belongs to a region with brightness variations; when the block does not belong to a region with brightness variations, it sets `lic_flag` to a value of "0" and performs encoding without applying the LIC. Decoder 200 decodes the `lic_flag` in the write stream and decodes the current block by switching between applying and not applying the LIC based on the flag value.

[0582] One example of a different method for determining whether to apply the LIC procedure is based on whether the LIC procedure has already been applied to surrounding blocks. As a concrete example, when the current block has already been processed in merge mode, the inter-frame predictor 126 determines whether the encoded surrounding block selected in the MV derivation in merge mode has already been encoded using LIC. The inter-frame predictor 126 performs encoding by switching between applying and not applying LIC based on the result. It should be noted that the same procedure is also applied on the decoder 200 side in this example.

[0583] Already referenced Figure 66A and Figure 66B The luminance correction (LIC) process is described and further described below.

[0584] First, the inter-frame predictor 126 derives the MV from the reference picture (which is the coded picture) to obtain the reference picture corresponding to the current block to be encoded.

[0585] Next, the inter-frame predictor 126 uses the luminance pixel values ​​of the coded surrounding reference regions adjacent to the left and top of the current block, as well as the luminance values ​​at the corresponding positions in the reference image specified by MV, to extract information indicating how the luminance values ​​of the reference image change to the luminance values ​​of the current image, and calculates luminance correction parameters. For example, suppose the luminance pixel value of a given pixel in the surrounding reference region of the current image is p0, and the luminance pixel value of the pixel corresponding to the given pixel in the surrounding reference region of the reference image is p1. The inter-frame predictor 126 calculates coefficients A and B for optimizing A×p1+B=p0 as luminance correction parameters for multiple pixels in the surrounding reference region.

[0586] Next, the inter-frame predictor 126 performs a brightness correction process using the brightness correction parameters of the reference image in the reference picture specified by MV to generate a predicted image for the current block. For example, suppose the brightness pixel value in the reference image is p2, and the brightness-corrected brightness pixel value in the predicted image is p3. The inter-frame predictor 126 generates the predicted image after undergoing the brightness correction process by calculating A×p2+B=p3 for each pixel in the reference image.

[0587] For example, a region having a defined number of pixels extracted from each of its upper and left adjacent pixels can be used as a surrounding reference region. Furthermore, the surrounding reference region is not limited to regions adjacent to the current block; it can also be regions not adjacent to the current block. Figure 66A In the example shown, the surrounding reference region in the reference image can be a region in the current image specified by another MV, derived from the surrounding reference region in the current image. For example, the other MV can be an MV within the surrounding reference region of the current image.

[0588] Although the operations performed by encoder 100 are described here, it should be noted that decoder 200 performs similar operations.

[0589] It should be noted that LIC can be applied not only to luminance but also to chrominance. In this case, correction parameters can be derived individually for each of Y, Cb, and Cr, or a common correction parameter can be used for any of Y, Cb, and Cr.

[0590] Furthermore, the LIC procedure can be applied on a sub-block basis. For example, correction parameters can be derived using the surrounding reference regions in the current sub-block and the surrounding reference regions in the reference sub-block of the reference image specified by the MV of the current sub-block.

[0591] (Predictive Controller)

[0592] Prediction controller 128 selects one of the intra-frame prediction signal (the image or signal output from intra-frame predictor 124) and inter-frame prediction signal (the image or signal output from inter-frame predictor 126), and outputs the selected prediction image to subtractor 104 and adder 116 as the prediction signal.

[0593] (Prediction parameter generator)

[0594] Prediction parameter generator 130 can output information related to intra-frame prediction, inter-frame prediction, and the selection of prediction images in prediction controller 128 as prediction parameters to entropy encoder 110. Entropy encoder 110 can generate a stream based on the prediction parameters input from prediction parameter generator 130 and quantized coefficients input from quantizer 108. Prediction parameters can be used in decoder 200. Decoder 200 can receive and decode the stream and perform the same process as the prediction process performed by intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128. Prediction parameters can include, for example, (i) selecting a prediction signal (e.g., the MV, prediction type, or prediction mode used by intra-frame predictor 124 or inter-frame predictor 126), or (ii) optional indices, tags, or values ​​based on the prediction process performed in each of intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128, or values ​​indicating the prediction process.

[0595] (Decoder)

[0596] Next, the decoder 200, which is capable of decoding the stream output from the encoder 100 described above, will be described. Figure 67 This is a block diagram illustrating the functional structure of the decoder 200 in this embodiment. The decoder 200 is an apparatus for decoding a stream of encoded images on a block-by-block basis.

[0597] like Figure 67 As shown, decoder 200 includes entropy decoder 202, inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, cyclic filter 212, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, prediction parameter generator 222, and split determiner 224. It should be noted that intra-frame predictor 216 and inter-frame predictor 218 are configured as part of the prediction executor.

[0598] (Decoder installation example)

[0599] Figure 68 This is a functional block diagram illustrating an example installation of decoder 200. Decoder 200 includes a processor b1 and a memory b2. For example, Figure 67 The decoder 200 shown has multiple constituent elements mounted on it. Figure 68 The processor b1 and memory b2 are shown.

[0600] Processor b1 is a circuit that performs information processing and is coupled to memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. Processor b1 can be a processor, such as a CPU. Furthermore, processor b1 can be an aggregation of multiple electronic circuits. Additionally, for example, processor b1 can function as... Figure 67The functions of two or more constituent elements other than the constituent element used for storing information in the plurality of constituent elements of the decoder 200 shown, etc.

[0601] Memory b2 is a dedicated or general-purpose memory used by processor b1 to decode the stream. Memory b2 can be an electronic circuit and can be connected to processor b1. Furthermore, memory b2 can be included within processor b1. Additionally, memory b2 can be an aggregation of multiple electronic circuits. Furthermore, memory b2 can be a disk, optical disk, etc., or can be represented as a storage or recording medium. Furthermore, memory b2 can be non-volatile memory or volatile memory.

[0602] For example, memory b2 can store images or streams. Furthermore, memory b2 can store programs used by processor b1 to decode the streams.

[0603] Furthermore, for example, memory b2 can serve as Figure 67 The decoder 200 shown represents two or more of the constituent elements used for storing information. More specifically, memory b2 can serve as... Figure 67 The block memory 210 and frame memory 214 shown herein serve a specific purpose. More specifically, memory B2 can store reconstructed images (specifically, reconstructed blocks, reconstructed pictures, etc.).

[0604] It should be pointed out that in decoder 200, it is not Figure 67 All of the multiple constituent elements shown in the document must be implemented, and not all of the processes described herein must be performed. Figure 67 A portion of the constituent elements shown may be contained in another device, or some of the processes described herein may be performed by another device.

[0605] The following describes the overall flow of the process performed by decoder 200, followed by a description of each constituent element included in decoder 200. It should be noted that some constituent elements included in decoder 200 perform the same processes as some constituent elements in encoder 100, and therefore these same processes will not be described in detail again. For example, the inverse quantizer 204, inverse transformer 206, adder 208, block memory 210, frame memory 214, intra-frame predictor 216, inter-frame predictor 218, prediction controller 220, and recurrent filter 212 included in decoder 200 perform processes similar to those performed by the inverse quantizer 112, inverse transformer 114, adder 116, block memory 118, frame memory 122, intra-frame predictor 124, inter-frame predictor 126, prediction controller 128, and recurrent filter 120 included in decoder 200.

[0606] (Overall flow of the decoding process)

[0607] Figure 69 This is a flowchart illustrating an example of the overall decoding process performed by decoder 200.

[0608] First, the split determiner 224 in decoder 200 determines the splitting pattern (step Sp_1) for each of the multiple fixed-size blocks (128×128 pixels) included in the image, based on parameters input from entropy decoder 202. This splitting pattern is selected by encoder 100. Decoder 200 then performs steps Sp_2 to Sp_6 for each of the multiple blocks in the splitting pattern.

[0609] The entropy decoder 202 decodes the encoded quantized coefficients and the prediction parameters of the current block (specifically, entropy decoding) (step Sp_2).

[0610] Next, the inverse quantizer 204 inverse quantizes multiple quantized coefficients, and the inverse transformer 206 inverse transforms the result to recover the prediction residual (i.e., the difference block) (step Sp_3).

[0611] Next, all or some of the prediction executors, including the intra-frame predictor 216, the inter-frame predictor 218, and the prediction controller 220, generate the prediction signal for the current block (step Sp_4).

[0612] Next, adder 208 adds the predicted image to the predicted residual to generate the reconstructed image of the current block (also known as the decoded image block) (step Sp_5).

[0613] When the reconstructed image is generated, the cyclic filter 212 performs filtering of the reconstructed image (step Sp_6).

[0614] Decoder 200 then determines whether the decoding of the entire image has been completed (step Sp_7). If it is determined that the decoding has not been completed (No in step Sp_7), decoder 200 repeats the process that started from step Sp_1.

[0615] It should be noted that these steps Sp_1 to Sp_7 can be executed sequentially by the decoder 200, or two or more of these steps can be executed in parallel. The processing order of two or more of these steps can be modified.

[0616] (Split Determiner)

[0617] Figure 70 This is a conceptual diagram illustrating the relationship between the splitting determiner 224 and other constituent elements in the embodiment. As an example, the splitting determiner 224 may perform the following process.

[0618] For example, the split determiner 224 collects block information from block memory 210 or frame memory 214 and further obtains parameters from entropy decoder 202. The split determiner 224 can then determine the splitting pattern for fixed-size blocks based on the block information and parameters. The split determiner 224 can then output information indicating the determined splitting pattern to inverse transformer 206, intra-frame predictor 216, and inter-frame predictor 218. Inverse transformer 206 can perform an inverse transform of the transform coefficients based on the splitting pattern indicated by the information from split determiner 224. Intra-frame predictor 216 and inter-frame predictor 218 can generate predicted images based on the splitting pattern indicated by the information from split determiner 224.

[0619] (Entropy Decoder)

[0620] Figure 71 This is a block diagram illustrating an example of the functional configuration of the entropy decoder 202.

[0621] Entropy decoder 202 generates quantized coefficients, prediction parameters, and parameters related to the splitting mode by entropy decoding the stream. For example, CABAC is used for entropy decoding. More specifically, entropy decoder 202 includes, for example, a binary arithmetic decoder 202a, a context controller 202b, and a debinarizer 202c. Binary arithmetic decoder 202a uses context values ​​derived by context controller 202b to arithmetic decode the stream into a binary signal. Context controller 202b derives context values ​​based on the characteristics of the syntax elements or the surrounding state (i.e., the probability of occurrence of the binary signal) in the same manner as the context controller 110b of encoder 100. Debinarizer 202c performs debinarization to transform the binary signal output from binary arithmetic decoder 202a into a multi-level signal representing the quantized coefficients as described above. This binarization can be performed according to the binarization method described above.

[0622] Thus, the entropy decoder 202 outputs the quantized coefficients of each block to the inverse quantizer 204. The entropy decoder 202 can then convert the prediction parameters included in the stream (see [link to documentation]). Figure 1 The output is sent to intra-frame predictor 216, inter-frame predictor 218, and prediction controller 220. Intra-frame predictor 216, inter-frame predictor 218, and prediction controller 220 are capable of performing the same prediction process as those performed by intra-frame predictor 124, inter-frame predictor 126, and prediction controller 128 on the encoder 100 side.

[0623] Figure 72 This is a conceptual diagram illustrating the flow of an example CABAC procedure in the entropy decoder 202.

[0624] First, initialization is performed in the CABAC within the entropy decoder 202. During initialization, initialization and setting of the initial context value are performed in the binary arithmetic decoder 202a. The binary arithmetic decoder 202a and the debinarizer 202c then perform arithmetic decoding and debinarization of, for example, the encoded data of a CTU. Meanwhile, the context controller 202b updates the context value each time arithmetic decoding is performed. The context controller 202b then saves the context value for post-processing. For example, the saved context value is used to initialize the context value for the next CTU.

[0625] (Inverse quantizer)

[0626] Inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block, which are inputs from entropy decoder 202. More specifically, inverse quantizer 204 inverse-quantizes the quantized coefficients of the current block based on the quantization parameters corresponding to the quantized coefficients. Inverse quantizer 204 then outputs the inverse-quantized transform coefficients (i.e., transform coefficients) of the current block to inverse transform 206.

[0627] Figure 73 This is a block diagram illustrating an example of the functional configuration of the inverse quantizer 204.

[0628] The inverse quantizer 204 includes, for example, a quantization parameter generator 204a, a predictive quantization parameter generator 204b, a quantization parameter storage 204d, and an inverse quantization executor 204e.

[0629] Figure 74 This is a flowchart illustrating an example of the inverse quantization process performed by inverse quantizer 204.

[0630] Inverse quantizer 204 can be based on Figure 74 The illustrated process performs an inverse quantization procedure for each CU as an example. More specifically, the quantization parameter generator 204a determines whether to perform inverse quantization (step Sv_11). Here, when it is determined that inverse quantization should be performed (yes in step Sv_11), the quantization parameter generator 204a obtains the differential quantization parameters for the current block from the entropy decoder 202 (step Sv_12).

[0631] Next, the predicted quantization parameter generation 204b then obtains the quantization parameters of the processing unit different from the current block from the quantization parameter storage 204d (step Sv_13). The predicted quantization parameter generation 204b generates the predicted quantization parameters for the current block based on the obtained quantization parameters (step Sv_14).

[0632] The quantization parameter generator 204a then generates the quantization parameters for the current block based on the differential quantization parameters of the current block obtained from the entropy decoder 202 and the predicted quantization parameters of the current block generated by the predicted quantization parameter generator 204b (step Sv_15). For example, the differential quantization parameters of the current block obtained from the entropy decoder 202 and the predicted quantization parameters of the current block generated by the predicted quantization parameter generator 204b can be added together to generate the quantization parameters for the current block. Furthermore, the quantization parameter generator 204a stores the quantization parameters of the current block in the quantization parameter storage 204d (step Sv_16).

[0633] Next, the inverse quantization executor 204e uses the quantization parameters generated in step Sv_15 to inverse quantize the quantized coefficients of the current block into transform coefficients (step Sv_17).

[0634] It should be noted that different quantization parameters can be decoded at the bit sequence level, image level, slice level, block level, or CTU level. Furthermore, the initial values ​​of the quantization parameters can be decoded at the sequence level, image level, slice level, block level, or CTU level. In this case, the initial values ​​of the quantization parameters and the difference quantization parameters can be used to generate the quantization parameters.

[0635] It should be noted that the inverse quantizer 204 may include multiple inverse quantizers, and the quantized coefficients may be inverse quantized using an inverse quantization method selected from a variety of inverse quantization methods.

[0636] (Inverse Transformer)

[0637] The inverse transformer 206 recovers the prediction residual by performing an inverse transformation on the transform coefficients from the input of the inverse quantizer 204.

[0638] For example, when the information parsed from the stream indicates that EMT or AMT will be applied (e.g., when the AMT flag is true), the inverse transformer 206 performs an inverse transform on the transform coefficients of the current block based on the information indicating the type of transform to be parsed.

[0639] Furthermore, for example, when the information parsed from the stream indicates that NSST should be applied, the inverse transformer 206 applies a second inverse transform to the transform coefficients.

[0640] Figure 75 This is a flowchart illustrating an example of the process performed by the inverse converter 206.

[0641] For example, inverse transformer 206 determines whether there is information in the stream indicating that no orthogonal transformation was performed (step St_11). Here, when it is determined that such information does not exist (no in step St_11) (e.g., there is no indication of whether an orthogonal transformation was performed; there is an indication to perform an orthogonal transformation); inverse transformer 206 obtains information indicating the transformation type decoded by entropy decoder 202 (step St_12). Next, based on this information, inverse transformer 206 determines the transformation type for the orthogonal transformation in encoder 100 (step St_13). Inverse transformer 206 then uses the determined transformation type to perform an inverse orthogonal transformation (step St_14). Figure 75 As shown, when it is determined that there is information indicating that an orthogonal transformation has not been performed (Yes in step St_11) (for example, there is no explicit instruction to perform an orthogonal transformation; there is no instruction to perform an orthogonal transformation), the orthogonal transformation is not performed.

[0642] Figure 76 This is a flowchart illustrating an example of the process performed by the inverse converter 206.

[0643] For example, the inverse transformer 206 determines whether the transform size is less than or equal to a predetermined value (step Su_11). The determined value can be predetermined. Here, when it is determined that the transform size is less than or equal to the predetermined value (yes in step Su_11), the inverse transformer 206 obtains information from the entropy decoder 202 indicating which transform type was used by the encoder 100 in at least one transform type included in the first transform type group (step Su_12). It should be noted that this information is decoded by the entropy decoder 202 and output to the inverse transformer 206.

[0644] Based on this information, the inverse transformer 206 determines the transformation type for the orthogonal transformation in the encoder 100 (step Su_13). The inverse transformer 206 then performs an inverse orthogonal transformation on the transformation coefficients of the current block using the determined transformation type (step Su_14). When it is determined that the transformation size is not less than or equal to the determined value (No in step Su_11), the inverse transformer 206 performs an inverse transformation on the transformation coefficients of the current block using the second transformation type group (step Su_15).

[0645] It should be pointed out that, as an example, it can be based on Figure 75 or Figure 76The illustrated process performs an inverse orthogonal transform of inverse transformer 206 for each TU. Alternatively, the inverse orthogonal transform can be performed using a defined transform type without decoding information indicating the transform type used for the orthogonal transform. The defined transform type can be a predefined transform type or a default transform type. Specifically, the transform type can be DST7, DCT8, etc. In the inverse orthogonal transform, the inverse transform basis functions corresponding to the transform type are used.

[0646] (Adder)

[0647] Adder 208 reconstructs the current block by adding the prediction residual input from inverse transformer 206 and the prediction image input from prediction controller 220. In other words, it generates a reconstructed image of the current block. Adder 208 then outputs the reconstructed image of the current block to block memory 210 and cyclic filter 212.

[0648] (Block memory)

[0649] Block memory 210 is a memory used to store blocks included in the current image and that can be referenced in intra-frame prediction. More specifically, block memory 210 stores the reconstructed image output from adder 208.

[0650] (Loop Filter)

[0651] The cyclic filter 212 applies the cyclic filter to the reconstructed image generated by the adder 208, outputs the filtered reconstructed image to the frame memory 214, and provides the output of the decoder 200, for example, to a display device, etc.

[0652] When the information parsed from the stream indicating whether the ALF is on or off indicates that the ALF is on, a filter can be selected from multiple filters, for example, based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.

[0653] Figure 77 This is a block diagram illustrating an example of the functional configuration of the loop filter 212. It should be noted that the configuration of the loop filter 212 is similar to the configuration of the loop filter 120 of the encoder 100.

[0654] For example, such as Figure 77 As shown, the recurrent filter 212 includes a deblocking filter executor 212a, a SAO executor 212b, and an ALF executor 212c. The deblocking filter executor 212a performs deblocking filter processing on the reconstructed image. The SAO executor 212b performs the SAO process on the reconstructed image after the deblocking filter process. The ALF executor 212c performs the ALF process on the reconstructed image after the SAO process. It should be noted that the recurrent filter 212 does not always need to include... Figure 77 All constituent elements disclosed herein may be included, but only a portion of the constituent elements may be included. Furthermore, the cyclic filter 212 may be configured to operate differently from... Figure 77 The above process is executed according to the publicly disclosed processing order; it may be omitted. Figure 77 All processes shown, etc.

[0655] (Frame Memory)

[0656] Frame memory 214 is, for example, a memory used to store reference images used in inter-frame prediction, and may also be referred to as a frame buffer. More specifically, frame memory 214 stores the reconstructed image filtered by cyclic filter 212.

[0657] (Predictor (intra-frame predictor, inter-frame predictor, prediction controller))

[0658] Figure 78 This is a flowchart illustrating an example of the process performed by the predictor of decoder 200. It should be noted that the prediction executor may include all or part of the following constituent elements: intra-frame predictor 216; predictor 218; and prediction controller 220. The prediction executor includes, for example, intra-frame predictor 216 and inter-frame predictor 218.

[0659] The predictor generates a prediction image for the current block (step Sq_1). This prediction image can also be referred to as a prediction signal or a prediction block. It should be noted that the prediction signal is, for example, an intra-frame prediction signal or an inter-frame prediction signal. More specifically, the predictor generates the prediction image for the current block by generating the prediction image, recovering the prediction residual, and adding the prediction images to use a reconstructed image, which has already been obtained for another block. The predictor of decoder 200 generates the same prediction image as the predictor of encoder 100. In other words, the prediction image is generated according to a common method or a mutually corresponding method between the predictors.

[0660] The reconstructed image can be, for example, an image in a reference image, or an image of a decoded block (i.e., one of the other blocks mentioned above) in the current image (which is an image that includes the current block). A decoded block in the current image can be, for example, a neighboring block of the current block.

[0661] Figure 79 This is a flowchart illustrating another example of the process performed by the prediction of the decoder 200.

[0662] The predictor determines the method or pattern used to generate the predicted image (step Sr_1). For example, the method or pattern can be determined based on, for example, prediction parameters.

[0663] When the first method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the first method (step Sr_2a). When the second method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the second method (step Sr_2b). When the third method is determined as the mode for generating the predicted image, the predictor generates the predicted image according to the third method (step Sr_2c).

[0664] The first, second, and third methods can be different from each other for generating the predicted image. Each of the first to third methods can be an inter-frame prediction method, an intra-frame prediction method, or another prediction method. The reconstructed image described above can be used in these prediction methods.

[0665] Figure 80 This is a flowchart illustrating another example of the process performed by the prediction of the decoder 200.

[0666] As an example, the predictor can be based on Figure 80 The flowchart shown is used to perform the prediction process. Note that... Figure 80 The intra-block copying shown is a mode of inter-frame prediction, where the block included in the current image is called the reference image or reference block. In other words, intra-block copying does not reference images different from the current image. Furthermore, Figure 80 The PCM mode shown is a mode that is an intra-frame prediction mode that does not perform transformation and quantization.

[0667] (Intra-frame predictor)

[0668] Intra-predictor 216 performs intra-prediction based on the intra-prediction pattern parsed from the stream, using blocks in the current image stored in reference block memory 210, to generate a predicted image of the current block (i.e., the intra-prediction block). More specifically, intra-predictor 216 performs intra-prediction by referencing pixel values ​​(e.g., luminance and / or chrominance values) of one or more blocks adjacent to the current block to generate an intra-prediction image, and then outputs the intra-prediction image to prediction controller 220.

[0669] It should be noted that when the intra-prediction mode of referencing the luma block in the intra-prediction of the chroma block is selected, the intra-predictor 216 can predict the chroma component of the current block based on the luma component of the current block.

[0670] Furthermore, when the information parsed from the stream indicates that PDPC will be applied, the intra-predictor 216 corrects the intra-predicted pixel values ​​based on the horizontal / vertical reference pixel gradients.

[0671] Figure 81 This is a diagram illustrating an example of the process performed by the predictor 216 of the decoder 200.

[0672] The intra-frame predictor 216 first determines whether to use MPM. For example... Figure 81 As shown, the intra predictor 216 determines whether an MPM flag indicating 1 exists in the stream (step Sw_11). Here, when it is determined that an MPM flag indicating 1 exists (Yes in step Sw_11), the intra predictor 216 obtains information from the entropy decoder 202 indicating the intra prediction mode selected in the encoder 100 within the MPM. It should be noted that such information is decoded by the entropy decoder 202 and output to the intra predictor 216. Next, the intra predictor 216 determines the MPM (step Sw_3). The MPM includes, for example, six intra prediction modes. The intra predictor 216 then determines the intra prediction mode included among the multiple intra prediction modes included in the MPM and indicated by the information obtained in step Sw_12 (step Sw_14).

[0673] When it is determined that the MPM flag indicating 1 does not exist (No in step Sw_11), the intra predictor 216 obtains information indicating the intra prediction mode selected in encoder 100 (step Sw_15). In other words, the intra predictor 216 obtains information from entropy decoder 202 indicating an intra prediction mode selected from at least one intra prediction mode in encoder 100 that is never included in the MPM. It should be noted that such information is decoded by entropy decoder 202 and output to intra predictor 216. Intra predictor 216 then determines the intra prediction mode that is not included in the multiple intra prediction modes included in the MPM and is indicated by the information obtained in step Sw_15 (step Sw_17).

[0674] Intra-predictor 216 generates a predicted image based on the intra-prediction mode determined in step Sw_14 or step Sw_17 (step Sw_18).

[0675] (Inter-frame predictor)

[0676] The inter-frame predictor 218 predicts the current block by referring to a reference image stored in the frame memory 214. Prediction is performed on a unit basis: the current block or the current sub-block within the current block. It should be noted that a sub-block is included within a block and is a smaller unit than a block. The size of a sub-block can be 4×4 pixels, 8×8 pixels, or other sizes. The size of the sub-block can be switched on a unit basis, such as slices, blocks, or images.

[0677] For example, the inter-frame predictor 218 generates an inter-frame prediction image of the current block or the current sub-block by performing motion compensation using motion information (e.g., MV) parsed from a stream (e.g., prediction parameters output from the entropy decoder 202) and outputs the inter-frame prediction image to the prediction controller 220.

[0678] When the information parsed from the stream indicates that the OBMC mode should be applied, the inter-frame predictor 218 uses motion information of neighboring blocks in addition to motion information of the current block obtained through motion estimation to generate an inter-frame predicted image.

[0679] Furthermore, when the information parsed from the stream indicates that the FRUC mode should be applied, the inter-frame predictor 218 derives motion information by performing motion estimation based on the mode matching method parsed from the stream (e.g., bilateral matching or template matching). The inter-frame predictor 218 then uses the derived motion information to perform motion compensation (prediction).

[0680] Furthermore, when applying BIO mode, the inter-frame predictor 218 derives the MV based on a model assuming uniform linear motion. Additionally, when information from stream parsing indicates that affine mode should be applied, the inter-frame predictor 218 derives the MV of each sub-block based on the MVs of multiple adjacent blocks.

[0681] (MV Derivation Process)

[0682] Figure 82 This is a flowchart illustrating an example of the MV derivation process in decoder 200.

[0683] For example, inter-frame predictor 218 determines whether to decode motion information (e.g., motion video). For example, inter-frame predictor 218 may make this determination based on prediction modes included in the stream, or based on other information included in the stream. Here, when it is determined that motion information should be decoded, inter-frame predictor 218 derives the motion video of the current block in the mode where motion information is decoded. When it is determined that motion information should not be decoded, inter-frame predictor 218 derives the motion video in the mode where no motion information is decoded.

[0684] Here, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., described later. Modes in which motion information is decoded include the normal inter-frame mode, normal merging mode, and affine mode (specifically, affine inter-frame mode and affine merging mode). It should be noted that the motion information may include not only the MV but also the MV predictor selection information described later. Modes in which no motion information is decoded include the FRUC mode, etc. The inter-frame predictor 218 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.

[0685] Figure 83 This is a flowchart illustrating an example of the MV derivation process in decoder 200.

[0686] For example, the inter-frame predictor 218 can determine whether to decode the MV difference, i.e., based on prediction modes included in the stream, or based on other information included in the stream. Here, when it is determined that the MV difference should be decoded, the inter-frame predictor 218 can derive the MV of the current block in the mode in which the MV difference is decoded. In this case, for example, the MV difference included in the stream is decoded into prediction parameters.

[0687] When it is determined that no MV difference will be decoded, the inter-frame predictor 218 derives the MV in a mode in which no MV difference is decoded. In this case, the encoded MV difference is not included in the stream.

[0688] Here, as described above, the MV derivation modes include the normal inter-frame mode, normal merging mode, FRUC mode, affine mode, etc., as described later. Modes that encode the MV difference include the normal inter-frame mode and the affine mode (specifically, the affine inter-frame mode). Modes that do not encode the MV difference include the FRUC mode, the normal merging mode, and the affine mode (specifically, the affine merging mode). The inter-frame predictor 218 selects a mode from multiple modes for deriving the MV of the current block and uses the selected mode to derive the MV of the current block.

[0689] (MV derivation > Normal inter-frame mode)

[0690] For example, when the information parsed from the stream indicates that a normal inter-frame mode should be applied, the inter-frame predictor 218 derives the MV based on the information parsed from the stream and uses the MV to perform motion compensation (prediction).

[0691] Figure 84 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using normal inter-frame mode.

[0692] The inter-frame predictor 218 of the decoder 200 performs motion compensation for each block. First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks surrounding the current block in time or space (step Sh_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.

[0693] Next, the inter-frame predictor 218 extracts N (N is an integer of 2 or greater) MV candidates from the multiple MV candidates obtained in step Sg_11 as motion vector prediction candidates (also referred to as MV prediction candidates) according to the ranking in the determined priority order (step Sg_12). It should be noted that the ranking in the priority order can be predetermined for the corresponding N MV predictor candidates, and the ranking can be predetermined.

[0694] Next, the inter-frame predictor 218 decodes the MV predictor selection information from the input stream and uses the decoded MV predictor selection information to select one MV predictor candidate from N MV predictor candidates as the MV predictor for the current block (step Sg_13).

[0695] Next, the inter-frame predictor 218 decodes the MV difference from the input stream and derives the MV of the current block by adding the difference of the decoded MV difference to the selected MV predictor (step Sg_14).

[0696] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference image (step Sh_15). The processes in steps Sg_11 to Sg_15 are performed for each block. For example, when the processes in steps Sg_11 to Sg_15 are performed for each block in all blocks of a slice, inter-frame prediction of the slice using normal inter-frame mode is completed. Similarly, when the processes in steps Sg_11 to Sg_15 are performed for each block in all blocks of an image, inter-frame prediction of the image using normal inter-frame mode is completed. It should be noted that not all blocks included in a slice may undergo these processes in steps Sh_11 to Sh_15; when only some blocks undergo these processes, inter-frame prediction of the slice using normal inter-frame mode can be completed. This also applies to the images in steps Sh_11 to Sh_15. When these processes are performed on only some blocks in an image, inter-frame prediction of the image using normal inter-frame mode can be completed.

[0697] (MV derivation > Normal merge mode)

[0698] For example, when the information parsed from the stream indicates that a normal merging mode should be applied, the inter-frame predictor 218 derives the MV and uses the MV to perform motion compensation (prediction).

[0699] Figure 85 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using normal merging mode.

[0700] First, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks around the current block in time or space (step Sh_11). In other words, the inter-frame predictor 218 generates a list of MV candidates.

[0701] Next, the inter-frame predictor 218 selects an MV candidate from the multiple MV candidates obtained in step Sh_11, thereby deriving the MV of the current block (step Sh_12). More specifically, the inter-frame predictor 218 obtains MV selection information included in the stream as prediction parameters, and selects the MV candidate identified by the MV selection information as the MV of the current block.

[0702] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference image (step Sh_13). For example, the processes in steps Sh_11 to Sh_13 are performed for each block. For example, when the processes in steps Sh_11 to Sh_13 are performed for each block in all blocks of a slice, inter-frame prediction of the slice using normal merging mode is completed. Furthermore, when the processes in steps Sh_11 to Sh_13 are performed for each block in all blocks of an image, inter-frame prediction of the image using normal merging mode is completed. It should be noted that not all blocks included in a slice undergo these processes in steps Sh_11 to Sh_13; when only some blocks undergo these processes, inter-frame prediction of the slice using normal merging mode can be completed. This also applies to the images in steps Sh_11 to Sh_13. When these processes are performed on only some blocks in an image, inter-frame prediction of the image using normal merging mode can be completed.

[0703] (MV Derivation > FRUC Pattern)

[0704] For example, when information parsed from the stream indicates that the FRUC mode should be applied, the inter-frame predictor 218 derives the MV in the FRUC mode and uses the MV to perform motion compensation (prediction). In this case, the motion information is derived on the decoder 200 side, without being signaled from the encoder 100 side. For example, the decoder 200 can derive motion information by performing motion estimation. In this case, the decoder 200 performs motion estimation without using any pixel values ​​in the current block.

[0705] Figure 86 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using FRUC mode.

[0706] First, the inter-frame predictor 218 indicates a list of MVs of decoded blocks spatially or temporally adjacent to the current block by using MVs as MV candidates (this list is an MV candidate list, and can also be used, for example, as an MV candidate list for the normal merge mode) (step Si_11). Next, the best MV candidate is selected from the multiple MV candidates registered in the MV candidate list (step Si_12). For example, the inter-frame predictor 218 calculates an evaluation value for each MV candidate included in the MV candidate list and selects one of the MV candidates as the best MV candidate based on the evaluation value. Based on the selected best MV candidate, the inter-frame predictor 218 then derives the MV of the current block (step Si_14). More specifically, for example, the selected best MV candidate is directly derived as the MV of the current block. Alternatively, for example, the MV of the current block can be derived using pattern matching in a region around a certain location, which is included in the reference picture and corresponds to the selected best MV candidate. In other words, pattern matching and evaluation of the reference image can be performed in the region surrounding the best MV candidate, and when an MV with a better evaluation exists, the best MV candidate can be updated to the MV with the better evaluation, and the updated MV can be determined as the final MV of the current block. In an embodiment, updating the MV with the better evaluation may not be performed.

[0707] Finally, the inter-frame predictor 218 generates a predicted image for the current block by performing motion compensation for the current block using the derived MV and the decoded reference image (step Si_15). For example, the processes in steps Si_11 to Si_15 are performed for each block. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of a slice, inter-frame prediction of the slice using FRUC mode is completed. For example, when the processes in steps Si_11 to Si_15 are performed for each block in all blocks of an image, inter-frame prediction of the image using FRUC mode is completed. Each sub-block can be processed in a similar manner to the case of each block.

[0708] (MV Derivation > FRUC Pattern)

[0709] For example, when the information parsed from the stream indicates that an affine merge mode should be applied, the inter-frame predictor 218 derives the MV in the affine merge mode and uses the MV to perform motion compensation (prediction).

[0710] Figure 87 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using affine merging mode.

[0711] In affine merging mode, firstly, the inter-frame predictor 218 derives the MV at each control point of the current block (step Sk_11). The control points are the top-left corner and the top-right corner of the current block, such as... Figure 46A As shown, or the top-left corner, top-right corner, and bottom-left corner of the current block, as shown. Figure 46B As shown.

[0712] For example, when using Figures 47A to 47C When using the MV derivation method shown, such as Figure 47A As shown, the inter-frame predictor 218 examines the decoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left) in this order, and identifies the first valid block decoded according to the affine pattern. The inter-frame predictor 218 uses the identified first valid block decoded according to the affine pattern to derive the MV at the control point. For example, when block A is identified and block A has two control points, as... Figure 47B As shown, the inter-frame predictor 218 calculates the motion vector v0 at the top-left control point of the current block and the motion vector v1 at the top-right control point of the current block based on the motion vectors v3 and v4 at the top-left and top-right control points of the decoded block, including block A. The MV at each control point is derived in this way.

[0713] It should be pointed out that, such as Figure 49A As shown, when block A is identified and block A has two control points, the MV at the three control points can be calculated, and as follows: Figure 49B As shown, when block A is identified and when block A has three control points, the MV at two control points can be calculated.

[0714] Furthermore, when MV selection information is included in the stream as a prediction parameter, the inter-frame predictor 218 can use the MV selection information to derive the MV at each control point of the current block.

[0715] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sk_12). The inter-frame predictor 218 then uses these affine MVs and the decoded reference image to perform motion compensation for the sub-blocks (step Sk_13). While performing the processes in steps Sk_12 and Sk_13 for each sub-block included in the current block, the inter-frame prediction using the affine merging mode for the current block ends. In other words, motion compensation for the current block is performed to generate the predicted image for the current block.

[0716] It should be noted that the aforementioned MV candidate list can be generated in step Sk_11. The MV candidate list may, for example, include a list of MV candidates derived using multiple MV derivation methods for each control point. Multiple MV derivation methods may, for example, be... Figures 47A to 47C The MV derivation method shown Figure 48A and 48B The MV derivation method shown Figure 49A and 49B The MV derivation method shown can be any combination of other MV derivation methods.

[0717] It should be noted that, in addition to affine patterns, the MV candidate list may include MV candidates from patterns that perform prediction on a sub-block basis.

[0718] It should be noted that, for example, an MV candidate list including MV candidates can be generated in both an affine merge pattern using two control points and an affine merge pattern using three control points. Alternatively, an MV candidate list including MV candidates in the affine merge pattern using two control points and an MV candidate list including MV candidates in the affine merge pattern using three control points can be generated separately. Alternatively, an MV candidate list including MV candidates can be generated in either an affine merge pattern using two control points or an affine merge pattern using three control points.

[0719] (MV Derivation > Affine Inter-Frame Mode)

[0720] For example, when the information parsed from the stream indicates that an affine inter-frame mode should be applied, the inter-frame predictor 218 derives the MV in the affine inter-frame mode and uses the MV to perform motion compensation (prediction).

[0721] Figure 88 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using affine inter-frame modes.

[0722] In affine inter-frame mode, firstly, the inter-frame predictor 218 derives the MV prediction values ​​(v0, v1) or (v0, v1, v2) for the corresponding two or three control points of the current block (step Sj_11). For example... Figure 46A or Figure 46B As shown, the control points are the top left corner, the top right corner, and the bottom left corner of the current block.

[0723] Inter-frame predictor 218 obtains MV predictor selection information included in the stream as prediction parameters, and uses the MV identified by the MV predictor selection information to derive the MV predictor at each control point of the current block. For example, when using Figure 48A and Figure 48B When the MV derivation method is shown, the inter-frame predictor 218 uses... Figure 48A or Figure 48BAmong the decoded blocks near the control points of the current block, the MV of the block selected by the MV predictor selection information is derived at the control point of the current block as either (v0, v1) or (v0, v1, v2).

[0724] Next, the inter-frame predictor 218 obtains each MV difference included in the stream as a prediction parameter, and adds the MV predictor at each control point of the current block to the MV difference corresponding to the MV predictor (step Sj_12). This derives the MV of the current block at each control point.

[0725] Next, the inter-frame predictor 218 performs motion compensation for each of the multiple sub-blocks included in the current block. In other words, the inter-frame predictor 218 uses two motion vectors v0 and v1 and the above expression (1A) or three motion vectors v0, v1, and v2 and the above expression (1B) to calculate the MV of each of the multiple sub-blocks as an affine MV (step Sj_13). The inter-frame predictor 218 then uses these affine MVs and the decoded reference image to perform motion compensation for the sub-blocks (step Sj_14). While performing the processes in steps Sj_13 and Sj_14 for each sub-block included in the current block, the inter-frame prediction using the affine merging mode for the current block ends. In other words, motion compensation for the current block is performed to generate the predicted image for the current block.

[0726] It should be noted that the above MV candidate list can be generated in step Sj_11 in the same way as in step Sk_11.

[0727] (MV Derivation > Triangle Pattern)

[0728] For example, when the information parsed from the stream indicates that a triangle pattern should be applied, the inter-frame predictor 218 derives the MV in the triangle pattern and uses the MV to perform motion compensation (prediction).

[0729] Figure 89 This is a flowchart illustrating an example of the inter-frame prediction process in decoder 200 using a pseudo-triangle pattern.

[0730] In the triangle mode, firstly, the inter-frame predictor 218 splits the current block into a first partition and a second partition (step Sx_11). For example, the inter-frame predictor 218 can obtain partition information from the stream as prediction parameters, which is information related to the splitting. The inter-frame predictor 218 can then split the current block into a first partition and a second partition based on the partition information.

[0731] Next, the inter-frame predictor 218 obtains multiple MV candidates for the current block based on information such as the MVs of multiple decoded blocks around the current block in time or space (step Sx_12). In other words, the inter-frame predictor 218 generates a list of MV candidates.

[0732] Inter-frame predictor 218 then selects the MV candidate of the first partition and the MV candidate of the second partition from the plurality of MV candidates obtained in step Sx_11 as the first MV and the second MV, respectively (step Sx_13). At this time, inter-frame predictor 218 can obtain MV selection information from the stream to identify each selected MV candidate as a prediction parameter. Inter-frame predictor 218 can then select the first MV and the second MV based on the MV selection information.

[0733] Next, the inter-frame predictor 218 performs motion compensation using the selected first MV and the decoded reference image to generate a first predicted image (step Sx_14). Similarly, the inter-frame predictor 218 performs motion compensation using the selected second MV and the decoded reference image to generate a second predicted image (step Sx_15).

[0734] Finally, the inter-frame predictor 218 generates the prediction image for the current block by performing a weighted sum of the first and second prediction images (step Sx_16).

[0735] (MV estimate > DMVR)

[0736] For example, information from stream parsing indicates that DMVR will be applied, and inter-frame predictor 218 uses DMVR to perform motion estimation.

[0737] Figure 90 This is a flowchart illustrating an example of the motion estimation process performed by the DMVR in decoder 200.

[0738] Inter-frame predictor 218 derives the MV of the current block based on the merging mode (step S1_11). Next, inter-frame predictor 218 derives the final MV of the current block by searching the region around the reference image indicated by the MV derived in S1_11 (step S1_12). In other words, in this case, the MV of the current block is determined according to DMVR.

[0739] Figure 91 This is a flowchart illustrating an example of the motion estimation process performed by the DMVR in decoder 200, and... Figure 58B same.

[0740] First of all, Figure 58AIn step 1 shown, the inter-frame predictor 218 calculates the cost between the search position indicated by the initial MV (also referred to as the starting point) and eight surrounding search positions. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at one of the search positions other than the starting point is determined to be minimized, the inter-frame predictor 218 changes its objective to obtaining the search position with the minimum cost and executes... Figure 58A The process in step 2 is shown. When the cost at the starting point is minimized, the inter-frame predictor 218 skips... Figure 58A The process in step 2 is shown, and the process in step 3 is executed.

[0741] exist Figure 58A In step 2, as shown, the inter-frame predictor 218 performs a search similar to that in step 1, using the search position after the target change as the new starting point based on the result of the process in step 1. The inter-frame predictor 218 then determines whether the cost at each search position other than the starting point is minimized. Here, when the cost at one of the search positions other than the starting point is determined to be minimized, the inter-frame predictor 218 performs the process in step 4. When the cost at the starting point is minimized, the inter-frame predictor 218 performs the process in step 3.

[0742] In step 4, the inter-frame predictor 218 treats the search position at the starting point as the final search position and determines the difference between the position indicated by the initial MV and the final search position as the vector difference.

[0743] exist Figure 58A In step 3 shown, the inter-frame predictor 218 determines the pixel position based on the cost at four points at the top, bottom, left, and right positions relative to the starting point in step 1 or step 2 to obtain the sub-pixel accuracy with the minimum cost, and regards the pixel position as the final search position.

[0744] The pixel position at sub-pixel precision is determined by weighted summation of each of the four vectors ((0,1),(0,-1),(-1,0), and(1,0)) among the top, bottom, left, and right vectors ((0,1),(0,-1),(-1,0), and(1,0)). The cost at the corresponding position in the four search positions is used as the weight. The inter-frame predictor 218 then determines the vector difference as the difference between the position indicated by the initial MV and the final search position.

[0745] (Motion compensation > BIO / OBMC / LIC)

[0746] For example, when the information parsed from the stream indicates that correction of the predicted image will be performed, the inter-frame predictor 218 corrects the predicted image based on a correction mode when generating the predicted image. This mode is, for example, one of the BIO, OBMC, and LIC mentioned above.

[0747] Figure 92This is a flowchart illustrating an example of the process of generating a predicted image in decoder 200.

[0748] Inter-frame predictor 218 generates a predicted image (step Sm_11) and corrects the predicted image according to any of the above modes (step Sm_12).

[0749] Figure 93 This is a flowchart illustrating another example of the process of generating a predicted image in decoder 200.

[0750] Inter-frame predictor 218 derives the MV of the current block (step Sn_11). Next, inter-frame predictor 218 generates a predicted image using the MV (step Sn_12) and determines whether to perform a correction process (step Sn_13). For example, inter-frame predictor 218 obtains prediction parameters included in the stream and determines whether to perform a correction process based on these prediction parameters. For example, these prediction parameters are flags indicating whether to apply one or more of the above-described modes. Here, when it is determined that a correction process should be performed (Yes in step Sn_13), inter-frame predictor 218 generates a final predicted image by correcting the predicted image (step Sn_14). It should be noted that in LIC, luminance and chrominance can be corrected in step Sn_14. When it is determined that a correction process should not be performed (No in step Sn_13), inter-frame predictor 218 outputs the final predicted image without correcting the predicted image (step Sn_15).

[0751] (Motion compensation > OBMC)

[0752] For example, when the information parsed from the stream indicates that OBMC should be performed, the inter-frame predictor 218 corrects the predicted image based on OBMC when generating the predicted image.

[0753] Figure 94 This is a flowchart illustrating an example of the OBMC correction process for the predicted image in decoder 200. It should be noted that... Figure 94 The flowchart in the diagram represents the use of Figure 62 The correction process for the predicted image of the current image and the reference image is shown.

[0754] First, such as Figure 62 As shown, the inter-frame predictor 218 uses the MV allocated to the current block to obtain the predicted image (Pred) through normal motion compensation.

[0755] Next, the inter-frame predictor 218 obtains a predicted image (Pred_L) by applying the motion vector (MV_L) already derived for the decoded block adjacent to the left of the current block to the current block (reusing the motion vector of the current block). The inter-frame predictor 218 then performs the first correction of the predicted image by overlapping the two predicted images Pred and Pred_L. This provides the effect of blending the boundaries between adjacent blocks.

[0756] Similarly, the inter-frame predictor 218 obtains a predicted image (Pred_U) by applying the MV (MV_U) already derived for the decoded block adjacent to the current block (reusing the motion vector of the current block). The inter-frame predictor 218 then performs a second correction to the predicted image by overlaying the predicted image Pred_U onto a predicted image (e.g., Pred and Pred_L) that has already undergone the first correction. This provides the effect of blending the boundaries between adjacent blocks. The predicted image obtained from the second correction is an image where the boundaries between adjacent blocks have been blended (smoothed), and is therefore the final predicted image for the current block.

[0757] (Motion compensation > BIO)

[0758] For example, when the information parsed from the stream indicates that a BIO should be performed, the inter-frame predictor 218 corrects the predicted image based on the BIO when generating the predicted image.

[0759] Figure 95 This is a flowchart illustrating an example of the BIO correction process for the predicted image in decoder 200.

[0760] like Figure 63 As shown, the inter-frame predictor 218 derives two motion vectors (M0, M1) using two reference images (Ref0, Ref1) that are different from the image (Cur Pic) including the current block. The inter-frame predictor 218 then derives the predicted image for the current block using the two motion vectors (M0, M1) (step Sy_11). It should be noted that motion vector M0 is the motion vector (MVx0, MVy0) corresponding to reference image Ref0, and motion vector M1 is the motion vector (MVx1, MVy1) corresponding to reference image Ref1.

[0761] Next, the inter-frame predictor 218 derives the difference image I0 for the current block using motion vector M0 and reference image L0. Furthermore, the inter-frame predictor 218 derives the difference image I1 for the current block using motion vector M1 and reference image L1 (step Sy_12). Here, difference image I0 is the image included in reference image Ref0 and to be derived for the current block, and difference image I1 is the image included in reference image Ref1 and to be derived for the current block. Each of difference image I0 and difference image I1 can be the same size as the current block. Alternatively, each of difference image I0 and difference image I1 can be a larger image than the current block. Furthermore, difference image I0 and difference image I1 can include predicted images obtained by using motion vectors (M0, M1) and reference images (L0, L1) and applying a motion compensation filter.

[0762] Furthermore, the inter-frame predictor 218 derives the gradient image (Ix0, Ix1, Iy0, Iy1) of the current block from the difference image I0 and the difference image I1 (step Sy_13). It should be noted that the gradient image in the horizontal direction is (Ix0, Ix1), and the gradient image in the vertical direction is (Iy0, Iy1). The inter-frame predictor 218 can derive the gradient images, for example, by applying a gradient filter to the difference image. The gradient image can be an image in which each pixel indicates the spatial change of pixel values ​​along the horizontal direction or the spatial change of pixel values ​​along the vertical direction.

[0763] Next, the inter-frame predictor 218 uses the interpolated image (I0, I1) and the gradient image (Ix0, Ix1, Iy0, Iy1) to derive the optical flow (vx, vy) as a velocity vector for each sub-block of the current block (step Sy_14). As an example, the sub-block can be a 4×4 pixel sub-CU.

[0764] Next, the inter-frame predictor 218 uses optical flow (vx, vy) to correct the predicted image of the current block. For example, the inter-frame predictor 218 uses optical flow (vx, vy) to derive correction values ​​for the pixel values ​​included in the current block (step Sy_15). The inter-frame predictor 218 can then use the correction values ​​to correct the predicted image of the current block (step Sy_16). It should be noted that the correction values ​​can be derived on a pixel-by-pixel basis, or on a multi-pixel basis, or on a sub-block basis, etc.

[0765] It should be noted that the BIO process is not limited to Figure 95 The publicly disclosed process. It can be executed only. Figure 95 The processes disclosed in the document may be modified, or different processes may be added or used as alternatives, or these processes may be executed in a different order of processing.

[0766] (Motion compensation > LIC)

[0767] For example, when the information parsed from the stream indicates that LIC should be performed, the inter-frame predictor 218 corrects the predicted image based on LIC when generating the predicted image.

[0768] Figure 96 This is a flowchart illustrating an example of the correction process of the LIC for the predicted image in decoder 200.

[0769] First, the inter-frame predictor 218 uses MV to obtain a reference image corresponding to the current block from the decoded reference image (step Sz_11).

[0770] Next, the inter-frame predictor 218 extracts information indicating how the luminance values ​​change between the current image and the reference image for the current block (step Sz_12). This extraction can be performed based on the luminance pixel values ​​of the decoded left adjacent reference region (surrounding reference region) and the decoded upper adjacent reference region (surrounding reference region), as well as the luminance pixel values ​​at the corresponding positions in the reference image specified by the derived MV. The inter-frame predictor 218 uses the information indicating how the luminance values ​​change to calculate the luminance correction parameters (step Sz_13).

[0771] Inter-frame predictor 218 generates a predicted image for the current block by performing a brightness correction process, in which brightness correction parameters are applied to a reference image in the reference picture specified by MV (step Sz_14). In other words, the predicted image (which is a reference image in the reference picture specified by MV) is corrected based on the brightness correction parameters. In this correction, either brightness or chromaticity can be corrected.

[0772] (Predictive Controller)

[0773] Prediction controller 220 selects an intra-frame predicted image or an inter-frame predicted image and outputs the selected image to adder 208. In general, the configuration, function, and procedure of prediction controller 220, intra-frame predictor 216, and inter-frame predictor 218 on the decoder 200 side can correspond to the configuration, function, and procedure of prediction controller 128, intra-frame predictor 124, and inter-frame predictor 218 on the encoder 100 side.

[0774] (Decoding using predicted chroma samples)

[0775] In the first aspect, it is determined whether the luminance samples can be used to predict the chrominance samples of the current block, wherein the predicted chrominance samples are used to decode the block.

[0776] Figure 97 This is a flowchart illustrating an example of a process 1000 for decoding a block using predicted chroma samples, which can be performed by, for example... Figure 7 encoder 100 or Figure 67 The decoder 200 is executed. For convenience, refer to... Figure 67 Decoder 200 to describe Figure 97 .

[0777] At S1001, decoder 200 determines whether the current chroma block is within an MxN non-overlapping region aligned with the chroma samples of the MxN grid. Figure 98 and Figure 99 This is a conceptual diagram used to illustrate an example of determining whether the current chroma block is within an MxN non-overlapping region aligned with the chroma samples of the MxN grid. In some formats, such as YUV420, a 16x16 pixel chroma region corresponds to a 32x32 pixel luma region. Figure 98 and Figure 99 As shown, chromaticity blocks within a 32×32 luminance region aligned with a 16×16 chromaticity grid are identified as being within an M×N non-overlapping region aligned with an M×N chromaticity sample grid. Chromaticity blocks not within the 32x32 luminance region are not identified as being within an M×N non-overlapping region aligned with chromaticity samples of the M×N grid. Figure 99 As shown, the luminance sample of the chrominance block can be used to predict the chrominance sample of the chrominance block, since the chrominance block is included in the grid (as shown in the figure, a 16x16 grid), and the co-located luminance block is also inside the co-located 32x32 area.

[0778] In some embodiments, luminance samples may not be used to predict the chromaticity samples of blocks that are not identified as being within an MxN non-overlapping region aligned with the chromaticity samples of the MxN grid. However, when other conditions are met (as discussed below with reference to S1002), luminance samples may be used, for example by default, to predict the chromaticity samples of blocks identified as being within an MxN non-overlapping region aligned with the chromaticity samples of the MxN grid.

[0779] like Figure 97 As shown, when it is not determined at S1001 that the current chroma block is within an MxN non-overlapping region aligned with the chroma samples of the MxN grid, process 1000 proceeds from S1001 to S1004, where decoder 200 predicts the chroma samples of the block without using luminance samples. Process 1000 proceeds from S1004 to S1005, where decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1001 that the current chroma block is within an MxN non-overlapping region, process 1000 proceeds from S1001 to S1002.

[0780] At S1002, decoder 200 determines whether to split the current luminance VPDU into smaller blocks. There are various ways to determine whether to split the current luminance VPDU into smaller blocks; see below for reference. Figure 102 and Figure 103 Some examples are discussed in more detail. When it is not determined at S1002 that the current luminance VPDU will be split into smaller blocks, process 1000 proceeds from S1002 to S1004, where decoder 200 predicts the chroma samples of the block without using luminance samples. Process 1000 proceeds from S1004 to S1005, where decoder 200 decodes the block using the predicted chroma samples. When it is determined at S1002 that the current luminance VPDU will be split into smaller blocks, process 1000 proceeds from S1002 to S1003, where decoder 200 predicts the chroma samples of the block using luminance samples. Process 1000 proceeds from S1003 to S1005, where decoder 200 decodes the block using the predicted chroma samples. In some embodiments, additional considerations may be taken into account to determine whether to use luminance samples to decode the chroma samples of the block, for example, as referenced below. Figure 103 The subject of discussion.

[0781] Figure 100 This is a conceptual diagram used to illustrate a VPDU. A VPDU is non-overlapping and represents the buffer size of a pipeline stage. Figure 100 The left side (labeled a) shows an example of a 128x128 CTU with four 64x64 VPDUs. Figure 100 The right side (labeled b) shows an example of a 128x128 CTU with 16 32x32VPDUs.

[0782] Figure 101 This is a conceptual diagram illustrating an example of how to determine whether a chromaticity sample can be predicted using a luminance sample based on whether a luminance VPDU is split into blocks. The left side shows the luminance CTU, and the right side shows the corresponding chromaticity CTU. As shown, luminance VPDU0 will be split into blocks, while luminance VPDU1 will not. Therefore, luminance samples can be used to predict the chromaticity sample of VPDU0, and luminance samples can be used to predict the chromaticity sample of VPDU1 without using luminance samples.

[0783] Figure 102 This is a conceptual diagram illustrating two example ways to determine whether a luminance VPDU will be split into smaller blocks. Figure 102In the first example shown on the left (labeled a), whether the luminance VPDU should be split can be determined based on the split flag associated with it. As shown, when the split flag value is 1, the VPDU will be split (and luminance samples can be used to predict the chroma samples of the block). When the split flag value is 0, the VPDU is not split (and luminance samples can not be used to predict the chroma samples of the block). Other split flag values ​​can be used to determine whether the luminance VPDU can be split.

[0784] exist Figure 102 In the second example shown on the right (labeled b), the decision to split the luminance VPDU can be based on the quadtree split depth of the luminance block. As shown, the quadtree split depth of the luminance block of VPDU0 is greater than 1, therefore, when decoding the block of VPDU0, luminance samples can be used to predict chrominance samples. Conversely, the quadtree split depth of the block of VPDU1 is less than or equal to 1, therefore, when decoding the block of VPDU0, luminance samples can be used to predict chrominance samples. Other split depth values ​​can be used to determine whether a luminance VPDU can be split.

[0785] Figure 103 This is a conceptual diagram used to illustrate additional considerations that can be taken into account to determine whether to use luminance samples to predict the chrominance samples of a block. As shown in the diagram, whether the current block size is equal to or less than a threshold block size can be used as an additional consideration in determining whether to use luminance samples to predict the chrominance samples of a block.

[0786] The threshold block size can be a default block size, a block size notified by a signal, or a determined block size, and can be either a luminance or chrominance block size. For example, if the threshold block size is a 16x16 luminance block size, then the luminance block size of VPDU0 is greater than 16x16, so it can be determined that luminance samples will not be used to determine the chrominance samples of the block. The threshold block size can be used in S1002 to determine whether the current luminance VPDU will be split into smaller blocks.

[0787] It can be modified in various ways Figure 97 The process 1000 can be modified in several ways. For example, process 1000 can be modified to perform more actions than shown, to perform fewer actions than shown, to perform actions in various orders, or to combine or split actions. For example, prior to S1001 or S1002, process 1000 can be modified based on other considerations (e.g., references). Figure 103 The size of the current block under discussion is used to determine whether to use a luminance sample to predict the chrominance sample of the block. In another example, process 1000 can be modified to omit S1001.

[0788] The blocks described in each aspect can be replaced with rectangular or non-rectangular partitions. Figure 104 Examples of non-rectangular partitions are shown, such as triangular partitions, L-shaped partitions, pentagonal partitions, hexagonal partitions, and polygonal partitions. Other non-rectangular partitions can be used, and various combinations of shapes can be employed. The term "partition" described in each aspect can be replaced with the term "prediction unit." The term "partition" described in each aspect can also be replaced with the term "sub-prediction unit." The term "partition" described in each aspect can also be replaced with the term "encoding unit."

[0789] Among other advantages, determining whether luminance samples can be used to predict chrominance samples for decoding the current block helps reduce reconstruction latency and increases hardware implementation flexibility.

[0790] One or more aspects disclosed herein may be performed in combination with at least a portion of other aspects of this disclosure. Furthermore, one or more aspects disclosed herein may be performed by combining a portion of a process indicated in any flowchart according to these aspects, a portion of a device configuration, a portion of a syntax, etc., with other aspects. The aspects described in the constituent elements of a reference encoder may be similarly performed by the corresponding constituent elements of a decoder.

[0791] (Implementation and Application)

[0792] As described in each of the above embodiments, for example, each functional block or operation block can typically be implemented as an MPU (microprocessor unit) and memory. Furthermore, the process executed by each functional block can be implemented as a program execution unit, such as a processor that reads and executes software (programs) recorded on a recording medium such as ROM. The software can be distributed. The software can be recorded on various recording media, such as semiconductor memory. It should be noted that each functional block can also be implemented as hardware (dedicated circuitry). Various combinations of hardware and software can be employed.

[0793] The processes described in each embodiment can be implemented through integrated processing using a single device (system), or through distributed processing using multiple devices. Furthermore, the processor executing the above programs can be a single processor or multiple processors. In other words, integrated processing can be performed, or distributed processing can be performed.

[0794] The embodiments of this disclosure are not limited to the exemplary embodiments described above; various modifications can be made to the exemplary embodiments, and the results are also included within the scope of the embodiments of this disclosure.

[0795] Next, application examples of the motion picture encoding method (image encoding method) and motion picture decoding method (image decoding method) described in each of the above embodiments will be described, along with various systems for implementing these application examples. Such systems may include an image encoder employing the image encoding method, an image decoder employing the image decoding method, or an image encoder-decoder system that includes both an image encoder and an image decoder. Other configurations of such systems may be modified as needed.

[0796] (Usage example)

[0797] Figure 105 The overall configuration of a content delivery system ex100 suitable for implementing content distribution services is shown. The area in which communication services are provided is divided into cells of desired size, and base stations ex106, ex107, ex108, ex109, and ex110 (which are fixed radio stations in the illustrated example) are located in the respective cells.

[0798] In the content providing system ex100, devices including a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content providing system ex100 can combine and connect any combination of the aforementioned devices. In various embodiments, the devices can be connected directly or indirectly via a telephone network or near-field communication without using base stations ex106 to ex110. Furthermore, a streaming server ex103 can be connected to the devices including the computer ex111, game console ex112, camera ex113, home appliance ex114, and smartphone ex115 via, for example, the Internet ex101. The streaming server ex103 can also be connected to terminals in a hotspot, such as an airplane ex117, via a satellite ex116.

[0799] Note that instead of base stations ex106 to ex110, wireless access points or hotspots can be used. Streaming server ex103 can connect directly to communication network ex104 without going through the Internet ex101 or Internet service provider ex102, and can connect directly to aircraft ex117 without going through satellite ex116.

[0800] The camera ex113 can be a device capable of capturing still images and videos, such as a digital camera. The smartphone ex115 can be a smartphone device, a cellular phone, or a Personal Handheld Phone System (PHS) phone, capable of operating under 2G, 3G, 3.9G, and 4G systems, as well as next-generation 5G mobile communication system standards.

[0801] Household appliances, such as refrigerators or equipment included in household fuel cell cogeneration systems, are examples of such appliances.

[0802] In the content delivery system ex100, a terminal including image and / or video capture capabilities can perform real-time streaming, for example, by connecting to a streaming server ex103 via a base station ex106. During real-time streaming, the terminal (e.g., a computer ex111, a gaming device ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal in an airplane ex117) can perform the encoding processing described in the above embodiments on still images or video content captured by the user via the terminal. The encoded video data and audio data obtained by encoding the audio corresponding to the video can be multiplexed, and the obtained data can be sent to the streaming server ex103. In other words, the terminal acts as an image encoder according to one aspect of this disclosure.

[0803] The streaming server ex103 streams content data to the client requesting the stream. Examples of clients include a computer ex111, a gaming device ex112, a camera ex113, a home appliance ex114, a smartphone ex115, and a terminal ex117 inside an aircraft, all capable of decoding the encoded data described above. The device receiving the streaming data can decode and reproduce the received data. In other words, according to one aspect of this disclosure, each device can function as an image decoder.

[0804] (Decentralized processing)

[0805] The ex103 streaming server can be implemented as multiple servers or computers, with tasks such as data processing, logging, and streaming divided among them. For example, the ex103 streaming server can be implemented as a Content Delivery Network (CDN) that streams content via a network connecting multiple edge servers located around the world. In a CDN, edge servers physically close to the client can be dynamically assigned to the client. Content is cached and streamed to the edge servers to reduce loading time. For example, in the event of some type of error or connection change, such as due to a surge in business, data can be streamed at high speed and stably because, for example, processing can be divided among multiple edge servers, or streaming tasks can be switched to different edge servers and streaming can continue, thus avoiding any impact on the network.

[0806] Distributed processing is not limited to the division of processing for streaming; the encoding of captured data can be divided and performed between terminals, on the server side, or both. In one example, in typical encoding, processing is performed in two loops. The first loop is used to detect the complexity of the image frame-by-frame or scene-by-scene, or to detect the encoding load. The second loop is used for processing to maintain image quality and improve encoding efficiency. For example, by having the terminal perform the first encoding loop and the server receiving the content perform the second encoding loop, the processing load on the terminal can be reduced and the quality of the content and encoding efficiency can be improved. In this case, upon receiving a decoding request, the encoded data generated by the first loop performed by one terminal can be received and reproduced on another terminal in near real-time. This enables smooth real-time streaming.

[0807] In another example, a camera such as ex113 extracts features (or characteristic quantities) from an image, compresses the data associated with these features into metadata, and sends the compressed metadata to a server. For instance, the server determines the importance of objects based on the features and adjusts the quantization precision accordingly to perform compression appropriate to the meaning (or content importance) of the image. In this second compression process performed by the server, the feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction. Furthermore, encodings with relatively low processing loads (e.g., variable-length coding (VLC)) can be processed by the terminal, while encodings with relatively high processing loads (e.g., context-adaptive binary arithmetic coding (CABAC)) can be processed by the server.

[0808] In another example, there exists a situation where multiple terminals in a stadium, shopping mall, or factory capture multiple videos of roughly the same scene. In this case, for example, encoding can be distributed by dividing the processing task into units among the multiple terminals capturing video and among other terminals and servers that are not capturing video when necessary. A unit can be, for example, a group of pictures (GOP), a picture, or a tile generated from dividing pictures. This allows for reduced loading time and enables streaming that is closer to real-time.

[0809] Because the videos share largely similar scenes, they can be managed and / or instructed by the server so that videos captured by the terminal can reference each other. Furthermore, the server can receive encoded data from the terminal, change the reference relationships between data items, or manually correct or replace images before performing encoding. This makes it possible to generate higher quality and more efficient streams for individual data items.

[0810] Furthermore, the server can stream video data after performing transcoding to convert the video data's encoding format. For example, the server can convert the encoding format from MPEG to VP (e.g., VP9), or H.264 to H.265, etc.

[0811] In this way, encoding can be performed by a terminal or one or more servers. Therefore, although the device performing the encoding is referred to as a "server" or "terminal" in the following description, some or all of the process performed by a server can be performed by a terminal, and similarly, some or all of the process performed by a terminal can be performed by a server. The same applies to the decoding process.

[0812] (3D, multi-angle)

[0813] The use of images or videos captured simultaneously from different scenes by multiple terminals such as the EX113 camera and / or the EX115 smartphone, or images or videos of the same scene captured from different angles, is increasing. Videos captured by the terminals can be combined based on, for example, the relative positional relationships between the terminals obtained individually, or regions in the video that have matching feature points.

[0814] In addition to encoding 2D moving images, the server can also encode still images based on scene analysis of the moving images (e.g., automatically or at user-specified time points) and send the encoded still images to the receiving terminal. Furthermore, when the server can obtain the relative positional relationships between video capture terminals, it can generate 3D geometry of the scene based on video of the same scene captured from different angles, in addition to 2D moving images. The server can individually encode 3D data generated from, for example, point clouds, and based on the results of identifying or tracking people or objects using the 3D data, it can select or reconstruct from videos captured by multiple terminals and generate video to be sent to the receiving terminal.

[0815] This allows users to enjoy a scene by freely selecting the video corresponding to the video capture terminal, and also allows users to enjoy content by extracting the video from the selected viewpoint from 3D data reconstructed from multiple images or videos. Furthermore, for video, sound can be recorded from relatively different angles; the server can multiplex audio from a specific angle or space with the corresponding video and send the multiplexed video and audio.

[0816] In recent years, content that combines the real and virtual worlds, such as virtual reality (VR) and augmented reality (AR) content, has also become popular. In the case of VR images, a server can create images from two viewpoints, one for the left eye and one for the right eye, and perform encoding that allows reference between the two viewpoint images, such as multi-view encoding (MVC). Alternatively, the images can be encoded as separate streams without reference. When the images are decoded into separate streams, the streams can be synchronized during playback, thereby reconstructing a virtual 3D space based on the user's viewpoint.

[0817] In the case of AR images, a server can overlay information about virtual objects existing in virtual space onto camera information representing real-world space, such as based on 3D position or motion from the user's perspective. The decoder can acquire or store the virtual object information and 3D data, along with the motion from the user's perspective, to generate a 2D image, and then generate overlay data through seamless image stitching. Alternatively, in addition to requesting virtual object information, the decoder can send the motion from the user's perspective to the server. The server can generate overlay data based on the received motion and the 3D data stored on the server, encode the generated overlay data, and stream it to the decoder. Note that the overlay data typically includes an α value representing transparency in addition to RGB values, and the server sets the α value for parts other than the objects generated from the 3D data to, for example, 0, and encoding can be performed when these parts are transparent. Alternatively, the server can set the background to a predetermined RGB value, such as chroma key, and generate data that sets the area outside the object as the background. The predetermined RGB values ​​can be pre-determined.

[0818] Decoding of similar streaming data can be performed by the client (e.g., a terminal) on the server side, or it can be partitioned between them. In one example, a terminal can send a receive request to the server, the requested content can be received and decoded by another terminal, and the decoded signal can be sent to a device with a display. High-quality image data can be reproduced by distributing processing regardless of the processing capabilities of the terminals themselves and appropriately selecting the content. In yet another example, while a television is receiving large-size image data, regions of the image can be decoded and displayed on a personal terminal or the terminals of one or more viewers on the television, for example, by dividing the image into tiles. This allows viewers to share a large view and each viewer can examine a region he or she specifies, or examine a region more closely in detail.

[0819] In situations where multiple wireless connections can be established at short, medium, and long distances, both indoors and outdoors, streaming system standards such as MPEG-DASH can be used to seamlessly receive content. Users can switch data in real time while freely selecting decoders or display devices (including user terminals, displays placed indoors or outdoors, etc.). Furthermore, using information such as about the user's location, decoding can be performed simultaneously by switching which terminal handles decoding and which terminal handles content display. This allows information to be mapped and displayed while the user is en route to their destination, on the walls of nearby buildings with embedded devices capable of displaying content, or on a portion of the ground. Additionally, the bitrate of the received data can be switched based on the accessibility of the encoded data over the network (e.g., when the encoded data is cached on a server that can be quickly accessed from the receiving terminal, or when the encoded data is copied to an edge server in a content delivery service).

[0820] (Webpage optimization)

[0821] For example, Figure 106 An example of a webpage display screen on a computer ex111 is shown. For example, Figure 107 This example shows a webpage display screen on a smartphone ex115. Figure 106 and Figure 107 As shown, a webpage may include multiple image links as links to image content, and the appearance of the webpage may vary depending on the device used to view it. When multiple image links are visible on the screen, the display device (decoder) may display still images or I-images included in the content as image links until the user explicitly selects an image link, or until the image link is located approximately in the center of the screen or the entire image link fits the screen; multiple still images or I-images may be used to display videos such as animated GIFs; or only the base layer may be received, decoded, and displayed.

[0822] When a user selects an image link, the display device performs decoding, for example, by giving the base layer the highest priority. Note that if the webpage's HTML code contains information indicating that the content is scalable, the display device can decode down to the enhancement layer. Furthermore, to facilitate real-time playback, before a selection is made or when bandwidth is severely limited, the display device can reduce the delay between the point at which the foreground image is decoded and the point at which the decoded image is displayed (i.e., the delay from the start of content dec...

Claims

1. An encoder, comprising: Circuit; as well as A memory coupled to the circuit; The circuit performs the following operations during operation: Determine whether to split the current Luminance Virtual Pipeline Decoding Unit (VPDU) into smaller blocks; In response to determining that the current luminance virtual pipeline decoding unit should not be split into smaller blocks, blocks of chrominance samples are predicted without using luminance samples; In response to determining that the luminance virtual pipeline decoding unit is split into smaller blocks, blocks of luminance samples are used to predict chrominance samples; and The block is encoded using predicted chroma samples.

2. A decoder, comprising: Circuit; A memory coupled to the circuit; The circuit performs the following operations during operation: Determine whether to split the current Luminance Virtual Pipeline Decoding Unit (VPDU) into smaller blocks; In response to determining that the current luminance virtual pipeline decoding unit should not be split into smaller blocks, blocks of chrominance samples are predicted without using luminance samples; In response to determining that the luminance virtual pipeline decoding unit is split into smaller blocks, blocks of luminance samples are used to predict chrominance samples; and The block is decoded using predicted chroma samples.

3. A non-transitory medium that stores a bit stream and can be read by a computer. The bitstream includes syntax for instructing the computer to perform a decoding process. The decoding process includes: Determine whether to split the current Luminance Virtual Pipeline Decoding Unit (VPDU) into smaller blocks; In response to determining that the current luminance virtual pipeline decoding unit should not be split into smaller blocks, blocks of chrominance samples are predicted without using luminance samples; In response to determining that the luminance virtual pipeline decoding unit is split into smaller blocks, blocks of luminance samples are used to predict chrominance samples; as well as The block is decoded using predicted chroma samples.