Decoding device, encoding device, decoding method, and encoding method

By utilizing a neural network to decode computing power information in video coding, the solution addresses inefficiencies in existing technologies, enhancing encoding efficiency, image quality, and reducing processing load and circuit size.

WO2026094988A1PCT designated stage Publication Date: 2026-05-07PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-10-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in improving encoding efficiency, image quality, processing load, and circuit size, as well as selecting optimal elements and operations such as filters, blocks, motion vectors, and reference pictures.

Method used

Incorporating a neural network into decoding devices to decode computing power information from a bitstream, allowing for efficient decoding processes based on the neural network's capabilities, and integrating this information into the encoding process to optimize encoding decisions.

Benefits of technology

Enhances encoding efficiency, improves image quality, reduces processing load and circuit size, and allows for appropriate selection of encoding/decoding components, while maintaining processing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025038118_07052026_PF_FP_ABST
    Figure JP2025038118_07052026_PF_FP_ABST
Patent Text Reader

Abstract

This decoding device (200) is provided with a circuit (b1) and a memory (b2) connected to the circuit. During operation, the circuit (b1) decodes calculation capability information associated with a neural network from a bitstream (S701), and decodes an image from the bitstream using the neural network on the basis of the calculation capability information (S702). For example, during operation, the circuit (b1) may further determine whether or not the calculation capability of the decoding device (200) satisfies the calculation capability indicated by the calculation capability information, and if it is determined that the calculation capability of the decoding device (200) satisfies the calculation capability indicated by the calculation capability information, the circuit (b1) may generate a first image from the bitstream using the neural network and decode a second image from the bitstream using the first image as a reference.
Need to check novelty before this filing date? Find Prior Art

Description

Decoding device, encoding device, decoding method, and encoding method

[0001] This disclosure relates to encoding devices, etc.

[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), MPEG-LA, H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding). With this advancement, there is a constant need to provide improvements and optimizations to video coding technology to handle the ever-increasing volume of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video coding.

[0003] Non-patent document 1 relates to an example of a conventional standard concerning the video coding technology described above.

[0004] H. 266 (ISO / IEC 23090-3) / VVC (Versatile Video Coding)

[0005] Regarding the encoding methods described above, there is a need for proposals for new methods to improve encoding efficiency, image quality, processing load, circuit size, or to appropriately select elements or actions such as filters, blocks, size, motion vectors, reference pictures, or reference blocks.

[0006] This disclosure provides a configuration or method that can contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, improved processing applied to the decoded image, or provision of new processing. This disclosure may also include configurations or methods that can contribute to other benefits not mentioned above.

[0007] For example, a decoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, decodes computing power information associated with a neural network from a bitstream, and uses the neural network to decode an image from the bitstream based on the computing power information.

[0008] Each embodiment, or any part thereof, of the present disclosure enables at least one of the following: improved encoding efficiency, improved image quality, reduced encoding / decoding processing load, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment, or any part thereof, of the present disclosure enables appropriate selection of components / operations such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. The present disclosure also includes disclosures of configurations or methods that may provide benefits other than those mentioned above, such as configurations or methods that improve encoding efficiency while suppressing an increase in processing load.

[0009] Further advantages and effects of one aspect of this disclosure will be made apparent from the specification and drawings. Such advantages and / or effects may be obtained by several embodiments and features described in the specification and drawings, but not all of them are necessarily provided to obtain one or more advantages and / or effects.

[0010] These general or specific embodiments may be implemented as a system, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of system, method, integrated circuit, computer program, and recording medium.

[0011] A configuration or method relating to one aspect of this disclosure may contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. A configuration or method relating to one aspect of this disclosure may also contribute to other benefits not mentioned above.

[0012] This is a schematic diagram showing an example of the configuration of a transmission system according to an embodiment. This is a block diagram showing an example of the implementation of an encoding device according to an embodiment. This is a block diagram showing an example of the configuration of an encoding device according to an embodiment. This is a flowchart showing an example of the overall encoding process by the encoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the encoding device according to an embodiment. This is a block diagram showing an example of the implementation of a decoding device according to an embodiment. This is a block diagram showing an example of the configuration of a decoding device according to an embodiment. This is a flowchart showing an example of the overall decoding process by the decoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the decoding device according to an embodiment. This is a schematic diagram showing another example of the configuration of a transmission system according to an embodiment. This is a schematic diagram showing an example of the configuration of a transmitting device according to an embodiment. This is a block diagram showing an example of the implementation of a transmitting device according to an embodiment. This is a schematic diagram showing an example of the configuration of a receiving device according to an embodiment. This is a block diagram showing an example of the implementation of a receiving device according to an embodiment. This is a schematic diagram showing an example of the configuration of a bitstream generation device according to an embodiment. This is a block diagram showing an example of the implementation of a bitstream generation device according to an embodiment. This is a schematic diagram showing an example of the configuration of a storage medium and a computer according to an embodiment. This is a schematic diagram showing an example of the configuration of a computer according to an embodiment. This is a diagram showing an example of the hierarchical structure of data in a stream. This is a diagram showing an example of the configuration of a slice. This is a diagram showing an example of the configuration of a tile. This is a diagram showing an example of the configuration of a time-scalable stream. This is a diagram showing an example of the configuration of a stream using multilayer encoding. This is a diagram showing an example of block partitioning. This is a diagram showing an example of the configuration of the split section. This is a diagram showing an example of a split pattern. This is a diagram showing an example of a syntax tree for a split pattern. This is a diagram showing another example of a syntax tree for a split pattern. This is a block diagram showing an example of the configuration of the loop filter section. This is a diagram showing an example of the shape of the filter used in ALF (adaptive loop filter). This is a diagram showing another example of the shape of the filter used in ALF. This is a diagram showing another example of the shape of the filter used in ALF. This is a diagram showing an example of CCALF. This is a diagram showing the filter shape of CCALF.This is a block diagram showing an example of a detailed configuration of the loop filter section that functions as a DBF. This is a flowchart showing an example of the processing of the DBF processing section. This is a flowchart showing another example of the processing of the DBF processing section. This is a diagram showing an example of a DBF having filter characteristics symmetric with respect to the block boundary. This is a diagram to explain an example of a block boundary on which DBF processing is performed. This is a diagram showing an example of the block boundary strength Bs value. This is a block diagram showing an example of the configuration of the loop filter section. This is a flowchart showing an example of the processing performed in the prediction section of the encoding device. This is a flowchart showing another example of the processing performed in the prediction section of the encoding device. This is a flowchart showing another example of the processing performed in the prediction section of the encoding device. This is a diagram showing an example of each reference picture. This is a conceptual diagram showing an example of a reference picture list. This is a conceptual diagram showing another example of a reference picture. This is a conceptual diagram showing another example of generating a generated reference picture. This is a conceptual diagram showing another example of generating a generated reference picture. This is a flowchart showing the basic processing flow of inter prediction. This is a flowchart showing an example of MV derivation. This is a flowchart showing another example of MV derivation. This is a diagram showing an example of the classification of each mode of MV derivation. This is a diagram showing an example of the classification of each mode of MV derivation. This is a flowchart showing an example of inter prediction using normal inter mode. This is a flowchart showing an example of inter prediction using normal merge mode. This is a diagram illustrating an example of MV derivation using normal merge mode. This is a diagram illustrating an example of MV derivation using HMVP (History-based Motion Vector Prediction / Predictor) mode. This is a flowchart illustrating an example of FRUC (frame rate up conversion). This is a diagram illustrating an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. This is a diagram illustrating an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture. This is a diagram illustrating an example of MV derivation at the subblock level in affine mode using two control points.This is a diagram illustrating an example of deriving the MV per subblock in affine mode using three control points. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a diagram illustrating an affine mode with two control points. This is a diagram illustrating an affine mode with three control points. This is a conceptual diagram illustrating an example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block. This is a conceptual diagram illustrating another example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block. This is a flowchart illustrating an example of processing in affine merge mode. This is a flowchart illustrating an example of processing in affine inter mode. This is a diagram illustrating an example of two regions where a predicted image is generated by GPM. This is a conceptual diagram showing the pattern of two regions defined by GPM. This is a conceptual diagram showing the weighted average of pixel values ​​at region boundaries. This is a flowchart illustrating an example of GPM mode. This figure shows an example of the ATMVP (Advanced Temporal Motion Vector Prediction / Predictor) mode in which MV is derived for each subblock. This figure shows the relationship between merge mode and DMVR (decoder motion vector refinement). This is a conceptual diagram to explain an example of DMVR. This is a conceptual diagram to explain another example of DMVR for determining MV. This figure shows an example of motion search in DMVR. This is a flowchart showing an example of motion search in DMVR. This is a flowchart showing an example of predictive image generation. This is a flowchart showing another example of predictive image generation. This is a flowchart to explain an example of predictive image correction processing by OBMC (overlapped block motion compensation). This is a conceptual diagram to explain an example of predictive image correction processing by OBMC. This figure shows a model that assumes uniform linear motion.This is a flowchart showing an example of inter prediction according to BIO. This is a diagram showing an example of the configuration of an inter prediction unit that performs inter prediction according to BIO. This is a diagram to explain an example of a prediction image generation method using brightness correction processing by LIC (local illumination compensation). This is a flowchart showing an example of a prediction image generation method using brightness correction processing by LIC. This is a conceptual diagram showing an example of a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing pipeline processing. This is a block diagram showing an example of a configuration used in NN processing. This is a block diagram showing another example of a configuration used in NN processing. This is a conceptual diagram showing a patch used in NN processing. This is a conceptual diagram showing an extended patch that includes an extended region. This is a conceptual diagram showing an image containing two patches with overlapping regions. This is a conceptual diagram showing a patch defined in each picture. This is a conceptual diagram showing an example of a patch generation method. This is a conceptual diagram showing another example of a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a flowchart of the decoding process in the first embodiment. This is a flowchart of the encoding process in the first embodiment. This is a diagram showing examples of the display order and encoding order of multiple images. This is a diagram showing examples of images stored in a buffer. This is a diagram showing an example of processing in period T1. This is a diagram showing an example of processing in period T2. This is a diagram showing an example of processing in period T3. This is a diagram showing an example of processing in period T4. This is a diagram showing examples of the first, second, and third images. This is a diagram showing examples of the display order and encoding order of multiple images. This is a diagram showing examples of images stored in a buffer. This is a diagram showing an example of processing in period T1. This is a diagram showing an example of processing in period T2. This is a diagram showing an example of processing in period T3. This is a diagram showing an example of processing in period T4. This is a diagram showing examples of the first, second, and third images. This is a diagram showing examples of the display order and encoding order of multiple images. This is a diagram showing examples of images stored in a buffer.This figure shows an example of processing during period T1. This figure shows an example of processing during period T2. This figure shows an example of processing during period T3. This figure shows an example of processing during period T4. This figure shows examples of the first, second, and third images. This figure shows examples of the first, second, and third images. This figure shows examples of the display order and encoding order of multiple images. This figure shows the first example of processing. This figure shows the second example of processing. This figure shows examples of the display order and encoding order of multiple images. This figure shows an example of an image stored in a buffer. This figure shows an example of processing during period T1. This figure shows an example of processing during period T2. This figure shows an example of processing during period T3. This figure shows an example of processing during period T4. This figure shows examples of the first, second, and third images. This figure shows examples of the first, second, and third images. This figure shows the first example of a reference relationship. This figure shows the second example of a reference relationship. This figure shows the first example of syntax. This figure shows another example of the first example of syntax. This figure shows an example of a unit represented by nn_processed_delay. This figure shows another example of the first example of syntax. This figure shows an example of syntax when two images are used in NN processing in the second example. This figure shows a modified version of the syntax when two images are used in NN processing in the second example. This figure shows an example of syntax when NN processing is applied directly to the target image to generate a processed image in the second example. This figure shows a modified version of the syntax when NN processing is applied directly to the target image to generate a processed image in the second example. This figure shows a third example of syntax. This figure shows an example of is_nn_processed_flag associated with multiple images stored in a buffer. This is a flowchart of the decoding process by the decoding device. This is a flowchart of the encoding process by the encoding device. This is a flowchart of the decoding process according to the second embodiment. This is a flowchart of the encoding process according to the second embodiment. This figure shows an example of a GPU profile. This figure shows another example of a GPU profile. This figure shows an example of the syntax of a GPU profile. This figure shows an example of level limits associated with a GPU profile. This figure shows another example of level limits associated with a GPU profile.This figure shows an example of the syntax for GPU profile levels. This figure shows an example of defining tiers associated with a GPU profile. This figure shows an example of classifying GPU computing power into tiers. This figure shows an example of the syntax for GPU profile tiers. This figure shows an example of determining whether the computing power of a decoder satisfies the computing power information. This figure shows an example of the display order and encoding order of multiple images. This figure shows an example of processing in period T1. This figure shows an example of processing in period T2. This figure shows an example of the syntax for operation count information. This figure shows an example of processing in period T1. This figure shows an example of processing in period T1+p. This figure shows an example of processing in period T1+2p. This figure shows an example of processing in period T1. This figure shows an example of processing in period T2. This figure shows an example of the syntax for timing information. This figure shows an example of the syntax for timing information. This figure shows an example of the storage location of computing power information in a bitstream. This figure shows another example of the storage location of computing power information in a bitstream. This figure shows an example of the storage location of timing information in a bitstream. This figure shows another example of the storage location of timing information in a bitstream. This is a flowchart of the decoding process by the decoder. This is a flowchart of the encoding process by the encoder. This is a diagram showing an example of the overall configuration of a content supply system that realizes a content distribution service. This is a diagram showing an example of the configuration of a content distribution system. This is a diagram showing an example of a web page display screen. This is a diagram showing an example of a smartphone. This is a block diagram showing an example of a smartphone configuration.

[0013] [Introduction] A decoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, decodes computational power information associated with a neural network from a bitstream, and decodes an image from the bitstream using the neural network based on the computational power information. According to this, the decoding device can appropriately perform decoding processing using a neural network based on the computational power information.

[0014] For example, the circuit may, in the operation described above, further determine whether the computing power of the decoder satisfies the computing power indicated by the computing power information, and if it is determined that the computing power of the decoder satisfies the computing power indicated by the computing power information, it may generate a first image from the bitstream using the neural network and decode a second image from the bitstream by referring to the first image. In this way, the decoder can determine, based on the computing power information, whether it is capable of performing decoding processing using a neural network.

[0015] For example, the bitstream may further include timing information, and the circuit may generate the first image from the bitstream using the neural network at the timing indicated by the timing information if it determines that the computing power of the decoding device satisfies the computing power indicated by the computing power information. In this case, the decoding device can perform processing using the neural network at an appropriate timing based on the timing information.

[0016] For example, the bitstream further includes calculation count information indicating the total number of calculations performed on the input image input to the neural network, and the circuit may, in the determination, derive the GPU processing time of the decoder from the total number of calculations and the GPU (Graphics Processing Unit) inference speed of the decoder, and determine whether the GPU processing time of the decoder satisfies the computing power indicated by the computing power information, thereby determining whether the computing power of the decoder satisfies the computing power indicated by the computing power information. According to this, the decoder can use the calculation count information to determine whether the decoder is capable of performing decoding using a neural network.

[0017] For example, the computing power information may represent a GPU profile, and the GPU profile may define a set of decoding tools equipped with a neural network. In this case, the GPU profile can inform the decoding device of the computing power required for processing using the neural network.

[0018] For example, the computing power information may indicate one of several levels of the GPU profile, and each of these levels may define an upper limit on the computing power of the GPU. In this case, the level of the GPU profile can be used to notify the decoder of the computing power required for processing using a neural network.

[0019] For example, the computing power information may indicate one of several tiers of the GPU profile, and each of the tiers may define an upper limit on the computing power of the GPU. In this case, the computing power required for processing using a neural network can be notified to the decoding device by the tier of the GPU profile.

[0020] For example, the computing power information may include at least one of the following: the maximum resolution of the input image input to the neural network, the maximum inference speed of the GPU, the maximum memory of the GPU, the accuracy of the neural network, and the maximum GPU processing time. Based on this, the decoding device can determine whether it is capable of performing decoding using a neural network, based on the computing power information.

[0021] For example, the computing power information may be included in VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), picture header, slice header, APS (Adaptation Parameter Set), SEI (Supplemental Enhancement Information), VUI (Video Usability Information), tile header, or metadata.

[0022] For example, the timing information may be included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

[0023] For example, the calculation count information may be included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

[0024] An encoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit generates encoded data by encoding an image using a neural network during operation, and generates a bitstream including the encoded data and computing power information associated with the neural network.Accordingly, a decoding device that decodes the bitstream can appropriately perform decoding processing using a neural network based on the computing power information.

[0025] For example, the circuit may, in the operation described above, further determine whether the computing power of the encoding device satisfies the computing power indicated by the computing power information, and if it is determined that the computing power of the encoding device satisfies the computing power indicated by the computing power information, it may generate a first image using the neural network and encode a second image by referring to the first image. In this case, the encoding device can determine, based on the computing power information, whether it is capable of performing encoding processing using a neural network.

[0026] For example, the bitstream may further include timing information indicating the timing at which the decoding device generates a first image from the bitstream using the neural network. This allows the decoding device decoding the bitstream to perform processing using the neural network at an appropriate timing based on the timing information.

[0027] For example, the bitstream may further include calculation count information indicating the total number of calculations performed on the input image input to the neural network. This allows a decoding device that decodes the bitstream to use the calculation count information to determine whether the encoding device is capable of performing decoding using the neural network.

[0028] For example, the computing power information may represent a GPU profile, and the GPU profile may define a set of coding tools equipped with a neural network. In this case, the GPU profile can inform the decoding device of the computing power required for processing using the neural network.

[0029] For example, the computing power information may indicate one of several levels of the GPU profile, and each of these levels may define an upper limit on the computing power of the GPU. In this case, the level of the GPU profile can be used to notify the decoder of the computing power required for processing using a neural network.

[0030] For example, the computing power information may indicate one of several tiers of the GPU profile, and each of the tiers may define an upper limit on the computing power of the GPU. In this case, the computing power required for processing using a neural network can be notified to the decoding device by the tier of the GPU profile.

[0031] For example, the computing power information may include at least one of the following: the maximum resolution of the input image input to the neural network, the maximum inference speed of the GPU, the maximum memory of the GPU, the accuracy of the neural network, and the maximum GPU processing time. Based on this, a decoding device that decodes the bitstream can determine whether it is capable of performing decoding using a neural network, based on the computing power information.

[0032] For example, the computing power information may be included in VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), picture header, slice header, APS (Adaptation Parameter Set), SEI (Supplemental Enhancement Information), VUI (Video Usability Information), tile header, or metadata.

[0033] For example, the timing information may be included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

[0034] For example, the calculation count information may be included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

[0035] A decoding method according to one aspect of this disclosure decodes computing power information associated with a neural network from a bitstream, and decodes an image from the bitstream using the neural network based on the computing power information. According to this, the decoding method can appropriately perform the decoding process using the neural network based on the computing power information.

[0036] An encoding method according to one aspect of this disclosure generates encoded data by encoding an image using a neural network, and generates a bitstream including the encoded data and computing power information associated with the neural network.According to this, a decoding device that decodes the bitstream can appropriately perform decoding processing using the neural network based on the computing power information.

[0037] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0038] [Definition of Terms] Each term may be defined as follows, for example:

[0039] (1) Image: A unit of data composed of a collection of pixels, consisting of pictures and smaller blocks, and includes both still images and videos.

[0040] (2) Chroma Chroma represents a sample sequence or single sample that represents one of two color difference signals. The two color difference signals may be represented by the symbols Cb and Cr. Chroma is an adjective, for example, represented by the symbols Cb and Cr, or U and V. The term chrominance can also be used instead of chroma.

[0041] (3) Luminance (luma) Luminance refers to a sample sequence or a single sample representing a monochrome signal associated with the primary colors. Luminance is an adjective represented by the symbols Y or L. The term luminance can also be used instead of lumina.

[0042] (4) Component A component represents a single array or sample. A component may be, for example, one of the three arrays (i.e., luminance and two color differences) that make up a color format picture, or a single sample, or one array or a single sample that makes up a monochrome format picture.

[0043] (5) A picture is an image processing unit composed of a set of picture samples, and is sometimes called a frame or field. A picture can be a set of luminance samples, a set of two color difference samples, or a set of both luminance samples and a set of two color difference samples. A set of samples can also be described as a matrix of samples.

[0044] (5-1) I-Picture An I-picture is a picture encoded using only intra-prediction. I-pictures can be decoded independently. I-pictures are also called intra-pictures, intra-frames, keypictures, or keyframes.

[0045] (5-2) P-Picture A P-Picture is a picture encoded using single prediction. In other words, only one picture is referenced to encode the P-Picture. Single prediction is also called unidirectional prediction.

[0046] (5-3) B-Picture A B-picture is a picture encoded using bidirectional prediction. Bidirectional prediction is also expressed as dual prediction and means a prediction that references multiple pictures. Bidirectional prediction may refer to multiple pictures in different directions from the picture being processed (for example, a forward picture and a backward picture in different temporal directions), or it may refer to pictures in the same direction from the picture being processed.

[0047] Note that P-pictures and B-pictures are also called inter-pictures or inter-frames.

[0048] (6) A block is a processing unit of a set containing a specific number of pixels, and its name is not restricted, as shown in the following examples. It is also not restricted in shape, and includes not only rectangles made up of M x N pixels and squares made up of M x M pixels, but also other shapes.

[0049] (Examples of blocks) Slice / Tile / Brick CTB (Coding Tree Block) / CTU (Coding Tree Unit) / Superblock / Segment / Basic division unit / CU / Processing block unit / Prediction block unit (PU) / Translation block unit (TU) / Unit / Subblock / VPDU / Hardware processing division unit

[0050] A CTB is a processing unit into which components are divided. A CTB contains samples. A CTB is also called a CTU, superblock, or basic division unit. A CTB may be, for example, an N x N square block.

[0051] A Coding Block (CB) is a processing unit obtained by dividing a Coding Block (CTB). For example, it may be an M x N rectangular block. A CB is also called a Coding Unit (CU).

[0052] (7) A pixel / sample is the smallest unit of a picture, and includes not only pixels at integer positions but also pixels at decimal positions that are generated based on pixels at integer positions. A pixel / sample is a fundamental element that makes up a picture.

[0053] (8) Pixel value / sample value: An intrinsic value of a pixel, which includes not only luminance value, chrominance value, and RGB gradation, but also depth value, or binary values ​​of 0 or 1.

[0054] (9) Flags A flag is a variable or a single-bit syntax element. For example, a flag can only take one of two values: 0 or 1.

[0055] (10) A symbol or code used to transmit signal information, which includes not only discretized digital signals but also analog signals that take continuous values.

[0056] (11) Stream / Bitstream: Refers to a sequence of digital data or a flow of digital data. A stream / bitstream may consist of a single stream, or it may be divided into multiple layers and composed of multiple streams. It also includes cases where data is transmitted via serial communication over a single transmission path, as well as cases where data is transmitted via packet communication over multiple transmission paths.

[0057] (12) In the case of difference / difference scalar quantities, it is sufficient that the difference operation is included in addition to the simple difference (x - y), and this includes the absolute value of the difference (|x - y|), the squared difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a, b is a constant), and the offset difference (x - y + a: a is the offset). Note that the difference operation also includes cases where bit shifts are performed (x >> 1 - y >> 1).

[0058] (13) In the case of summation scalar quantities, it is sufficient that the sum operation is included in addition to the simple sum (x + y), and this includes the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), weighted sum (ax + by: a, b is a constant), and offset sum (x + y + a: a is the offset). Note that the sum operation also includes cases where bit shifts are performed (x >> 1 + y >> 1).

[0059] (14) Based on: This includes cases where factors other than the subject being based on are taken into consideration. It also includes cases where the result is obtained not only by obtaining a direct result, but also by obtaining an intermediate result.

[0060] (15) Using: This includes cases where elements other than the target of use are taken into consideration. It also includes cases where the result is obtained not only by obtaining a direct result, but also by obtaining an intermediate result.

[0061] (16) Prohibit, forbid. This can be rephrased as not being allowed. Also, not prohibiting or allowing something does not necessarily mean it is an obligation.

[0062] (17) To restrict (limit, restriction / restrict / restricted) This can be rephrased as not being allowed. Also, not being prohibited or being permitted does not necessarily mean being obligated. Furthermore, it is sufficient if it is prohibited in part quantitatively or qualitatively, and it also includes cases where it is prohibited entirely.

[0063] (18) MV (motion vector) A two-dimensional vector used in interpretation, which represents the offset from the coordinates of the decoded image to the coordinates of the reference image.

[0064] (19) Data assigned to a specific column or row in a database, such as an index table or list. For example, a reference index is an index for a reference picture list, where one reference picture is assigned to the reference picture index.

[0065] [Explanation of Descriptions] In the drawings, the same reference number indicates the same or similar component. Also, the size and relative position of components in the drawings are not necessarily depicted to a constant scale.

[0066] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all general or specific examples. The numerical values, shapes, materials, components, arrangement and connection forms of components, steps, relationships and sequences of steps shown in the following embodiments are examples only and are not intended to limit the scope of the claims.

[0067] The following describes embodiments of encoding and decoding devices. These embodiments are examples of encoding and decoding devices to which the processes and / or configurations described in each aspect of this disclosure can be applied. The processes and / or configurations can also be implemented in encoding and decoding devices different from those in the embodiments. For example, with respect to the processes and / or configurations applicable to the embodiments, one of the following may be implemented:

[0068] (1) Any of the multiple components of the encoding or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0069] (2) In the encoding or decoding device of the embodiment, any changes such as addition, replacement, or deletion of functions or processes performed by some of the multiple components of the encoding or decoding device may be made. For example, any of the functions or processes may be replaced or combined with other functions or processes described in any of the embodiments of this disclosure.

[0070] (3) In the methods performed by the encoding or decoding apparatus of the embodiment, any modifications such as additions, replacements, and deletions may be made to some of the processes included in the method. For example, any of the processes in the method may be replaced with or combined with other processes described in any of the embodiments of this disclosure.

[0071] (4) Some of the multiple components constituting the encoding or decoding device of the embodiment may be combined with components described in any of the embodiments of this disclosure, or with components that have some of the functions described in any of the embodiments of this disclosure, or with components that perform some of the processing performed by the components described in any of the embodiments of this disclosure.

[0072] (5) Components that provide some of the functions of the encoding or decoding device of the embodiment, or components that perform some of the processing of the encoding or decoding device of the embodiment, may be combined with or replaced with components described in any of the embodiments of this disclosure, components that provide some of the functions described in any of the embodiments of this disclosure, or components that perform some of the processing described in any of the embodiments of this disclosure.

[0073] (6) In a method performed by an encoding or decoding device of an embodiment, any of the processes included in the method may be replaced or combined with a process described in any of the embodiments of the present disclosure, or any of the similar processes.

[0074] (7) Some of the processes included in the methods performed by the encoding or decoding device of the embodiment may be combined with the processes described in any of the embodiments of this disclosure.

[0075] (8) The methods of carrying out the processes and / or configurations described in each aspect of the present disclosure are not limited to the encoding or decoding devices of the embodiments. For example, the processes and / or configurations may be carried out in devices used for purposes other than the video encoding or video decoding disclosed in the embodiments.

[0076] [System Configuration] Figure 1 is a schematic diagram showing an example of the configuration of the transmission system according to this embodiment.

[0077] A transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, an encoding device 100, a network Nw, and a decoding device 200, as shown in Figure 1.

[0078] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.

[0079] The original image input to the encoding device 100 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. Furthermore, the image is a higher-level concept than sequences, pictures, and blocks, and is not limited by spatial and temporal domains unless otherwise specified. The image consists of a sequence of pixels or pixel values, and the signal or pixel values ​​representing the image are also called samples.

[0080] Furthermore, the stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal. In addition, the encoding device 100 may be called an image encoding device or a video encoding device, and the encoding method by the encoding device 100 may be called an encoding method, an image encoding method, or a video encoding method.

[0081] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting.

[0082] Furthermore, the network Nw may be replaced by a storage medium that records streams such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).

[0083] The decoding device 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method that corresponds to the encoding method by the encoding device 100.

[0084] The decoding device 200 may also be called an image decoding device or a video decoding device, and the decoding method performed by the decoding device 200 may be called a decoding method, an image decoding method, or a video decoding method.

[0085] [Encoding device] Next, the encoding device 100 according to the embodiment will be described.

[0086] [Implementation Example of Encoding Device] Figure 2 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, the multiple components of the encoding device 100 shown in Figure 3, which will be described later, are implemented by the processor a1 and memory a2 shown.

[0087] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit for encoding images. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits. Furthermore, for example, processor a1 may play the role of multiple components of the encoding device 100 shown in Figure 3, which will be described later, excluding the component for storing information.

[0088] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode an image is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0089] For example, memory a2 may store the image to be encoded, or it may store a stream corresponding to the encoded image. Alternatively, memory a2 may store a program for processor a1 to encode the image.

[0090] Furthermore, for example, memory a2 may play the role of an information storage component among the multiple components of the encoding device 100 shown in Figure 3, which will be described later. Specifically, memory a2 may play the role of the block memory 118 and frame memory 122 shown in Figure 3, which will be described later. More specifically, memory a2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0091] The generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be encoded. Whether or not to encode them is determined appropriately based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.

[0092] Furthermore, in the encoding device 100, not all of the multiple components shown in Figure 3 described later are to be implemented, nor are all of the multiple processes described above to be performed. Some of the multiple components shown in Figure 3 described later may be included in other devices, and some of the multiple processes described later may be performed by other devices.

[0093] The following is a general description of the encoding device 100, followed by a description of the components included in the encoding device 100.

[0094] [Example of Encoding Device Configuration] Figure 3 is a block diagram showing an example of the configuration of an encoding device 100 according to an embodiment. The encoding device 100 encodes images in block units.

[0095] As shown in Figure 3, the encoding device 100 is a device that encodes an image in block units and comprises a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, and a prediction unit. The prediction unit comprises an intra-prediction unit 124, an inter-prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130.

[0096] These components are implemented, for example, by the processor a1 and memory a2 of the encoding device 100 shown in Figure 2 above.

[0097] [Overall Encoding Process Flow] Figure 4A is a flowchart showing an example of the overall encoding process by the encoding device 100. Figure 4B is a flowchart showing an example of the block-by-block processing by the encoding device 100.

[0098] First, the division unit 102 of the encoding device 100 divides each picture contained in the original image into multiple blocks. For example, the division unit 102 first divides the picture into blocks of a fixed size (for example, 128 x 128 pixels) (step Sa_1a). These fixed-size blocks are sometimes called coding tree units (CTUs).

[0099] Then, the division unit 102 selects a division pattern for the fixed-size block and further divides the fixed-size block into multiple blocks to constitute the selected division pattern (step Sa_2a).

[0100] The further divided blocks are of a variable size, for example, 64 x 64 pixels or less. The vertical and horizontal pixel counts of the variable size can be any combination of 4, 8, 16, 32, or 64. These variable-sized blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs).

[0101] In various implementation examples, CU, PU, ​​and TU do not need to be distinguished, and some or all blocks within a picture may be processing units for CU, PU, ​​or TU.

[0102] Then, the encoding device 100 performs the processing shown in steps Sa_3 to Sa_11 in Figure 4B for each of the multiple blocks (step Sa_3a). Here, an example of a CTU size of 128 x 128 pixels is given, but other sizes are also possible. For example, the size of the CTU may be 256 x 256 pixels, or a larger size such as 512 x 512 pixels. In this case, the vertical and horizontal size of the divided blocks, i.e., the number of pixels, may be greater than 64 pixels, such as 128 pixels or 256 pixels.

[0103] Furthermore, the division unit 102 outputs parameters indicating the division pattern to the conversion unit 106, the inverse conversion unit 114, the intra-prediction unit 124, the inter-prediction unit 126, and the entropy coding unit 110. The conversion unit 106 may convert the prediction residuals based on these parameters, and the intra-prediction unit 124 and the inter-prediction unit 126 may generate a prediction image based on these parameters. The entropy coding unit 110 may also perform entropy coding on these parameters.

[0104] Then, the encoding device 100 processes each of the multiple blocks (step Sa_3a). Specifically, the encoding device 100 processes steps Sa_3 to Sa_11 as shown in Figure 4B. Then, the encoding device 100 determines whether or not the encoding of the entire picture is complete (step Sa_4a), and if it determines that it is not complete (No. in step Sa_4a), it repeats the processing from step Sa_2a.

[0105] Next, the processing for each block will be explained (Sa_3 to Sa_11 in Figure 4B). First, the prediction unit generates a predicted image of the current block (step Sa_3). Specifically, the prediction unit generates a predicted image of the current block by referring to a reconstructed image generated by encoding and then decoding other blocks. The predicted image is also called a predicted signal, predicted block, or predicted sample.

[0106] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.

[0107] The predicted image is, for example, an intra-prediction image (intra-prediction signal) generated based on intra-prediction, or an inter-prediction image (inter-prediction signal) generated based on inter-prediction.

[0108] Intra prediction is a method of predicting the current block by referring to blocks within the current picture, and is also called in-screen prediction. Specifically, the intra prediction unit 124 performs intra prediction by referring to the pixel values ​​(e.g., luminance values ​​or chrominance values, etc.) of blocks adjacent to the current block that are included in the current picture stored in the block memory 118. As a result, the intra prediction unit 124 generates an intra prediction image and outputs the intra prediction image to the prediction control unit 128.

[0109] Inter-prediction is a method of predicting the current block by referring to a reference picture different from the current picture, and is also called inter-screen prediction. Specifically, the inter-prediction unit 126 generates an inter-predicted image by referring to a reference picture stored in the frame memory 122 and performing inter-prediction of the current block, and outputs the inter-predicted image to the prediction control unit 128.

[0110] Note that prediction processing using the reconstructed image may not be performed on some blocks, such as the first block of the first image to be encoded. In that case, the original image will be output in the next subtraction process without performing subtraction on the original image. The interpretation unit 126 may also generate a predicted image by performing prediction processing without using a reference image.

[0111] Next, the subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units that are input from the division unit 102 and divided by the division unit 102. In other words, the subtraction unit 104 generates the difference between the current block and the predicted image as the predicted residual (step Sa_4).

[0112] The prediction residual is also called the prediction error. The original image is the input signal to the encoding device 100, and is, for example, a signal representing the image of each picture that makes up the video (e.g., a luminance (luma) signal and two chroma (chroma) signals). If the prediction process is skipped, the prediction residual becomes the value of the original image. For example, the prediction process is skipped for the first block in the processing order.

[0113] Next, the conversion unit 106 applies a conversion process to the predicted residual to generate conversion coefficients and outputs the conversion coefficients to the quantization unit 108 (step Sa_5).

[0114] The transformation process performed by the transformation unit 106 is, for example, an orthogonal transformation such as a discrete cosine transform (DCT) or discrete sine transform (DST) that transforms the predicted residual in the spatial domain into transformation coefficients in the frequency domain. However, other transformation processes such as weblet transforms, non-orthogonal transforms, or transformation processes expressed by matrix operations such as NSST (non-separable secondary transform) as defined in VVC may also be used.

[0115] The conversion unit 106 may perform a conversion process selected from among a plurality of candidate conversion processes. The conversion unit 106 may output conversion coefficients generated by sequentially applying a plurality of conversion processes to the predicted residual, or it may output the predicted residual as is without performing any conversion processes.

[0116] Depending on the processing performed by the conversion unit 106, information indicating whether or not a conversion process is applied to the predicted residuals, and / or information indicating the conversion type, may be encoded in the bitstream. This information is, for example, signaled at the CU level, but is not limited to the CU level and may be encoded at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).

[0117] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106 (step Sa_6). Specifically, the quantization unit 108 quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the conversion coefficients. The quantization unit 108 then outputs a plurality of quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112. In other words, the quantization unit 108 generates quantization coefficients and outputs them to the entropy coding unit 110 and the inverse quantization unit 112.

[0118] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the error in the quantization coefficient (quantization error) increases.

[0119] Furthermore, the quantization unit 108 may quantize the conversion coefficients based on a quantization matrix. In other words, quantization parameters (QP) and / or a quantization matrix may be used for quantization. The quantization parameters (QP) and the quantization matrix may be encoded, for example, at the sequence level, picture level, slice level, brick level, or CTU level.

[0120] The quantization unit 108 performs quantization processing in a predetermined scanning order. This predetermined scanning order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined as ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).

[0121] Next, the entropy coding unit 110 generates a stream (step Sa_7) by coding (specifically, entropy coding) the plurality of quantization coefficients and the prediction parameters related to the generation of the predicted image. For this entropy coding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding), which is used in VVC, may be used. Alternatively, a coding method that is a modified version of CABAC used in VVC may be used.

[0122] Furthermore, the processing of the entropy encoding unit 110 does not necessarily have to be performed after the quantization process; for example, it may be performed outside the loop of processing for each block.

[0123] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore the predicted residuals by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (steps Sa_8 and Sa_9).

[0124] Next, the summing unit 116 reconstructs the current block by adding the predicted image to the recovered predicted residual (step Sa_10). This generates a reconstructed image. The reconstructed image is also called a reconstructed block, and the reconstructed image generated by the encoding device 100 is also called a local decoded block or local decoded image.

[0125] Next, the loop filter unit 120 performs filtering on the reconstructed image as needed (step Sa_11). Specifically, the loop filter unit 120 applies loop filtering to the reconstructed image output from the adder unit 116 and outputs the filtered reconstructed image to the frame memory 122.

[0126] Loop filters are filters used to reduce block noise that occurs at block boundaries, and include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), and sample adaptive offset (SAO). Here, loop filters are described as in-loop filters used within the coding loop, but loop filters may also be out-loop filters used outside the coding loop.

[0127] In the example described above, the encoding device 100 selects one division pattern for a fixed-size block and encodes each block according to that division pattern. However, it may also encode each block according to multiple division patterns. In this case, the encoding device 100 may evaluate the cost of each of the multiple division patterns and, for example, select the stream obtained by encoding according to the division pattern with the smallest cost as the final output stream.

[0128] Furthermore, the processes in steps Sa_1a to Sa_4a and Sa_3 to Sa_11 may be performed sequentially by the encoding device 100, some of these processes may be performed in parallel, and the order may be changed.

[0129] The encoding process performed by such an encoding device 100 is a hybrid encoding using predictive encoding and transformative encoding. Furthermore, predictive encoding is performed by an encoding loop consisting of a subtraction unit 104, a transformer unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transformer unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra-prediction unit 124, an inter-prediction unit 126, and a prediction control unit 128. In other words, the prediction processing unit consisting of the intra-prediction unit 124 and the inter-prediction unit 126 constitutes a part of the encoding loop.

[0130] [Decoding Device] Next, a decoding device 200 capable of decoding the stream output from the encoding device 100 will be described.

[0131] [Implementation Example of Decryption Device] Figure 5 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and memory b2. For example, the multiple components of the decoding device 200 shown in Figure 6, which will be described later, are implemented by the processor b1 and memory b2 shown in Figure 5.

[0132] Processor b1 is a circuit that performs information processing and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit for decoding streams. Processor b1 may be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits. Furthermore, for example, processor b1 may play the role of multiple components of the decoding device 200 shown in Figure 6, etc., described later, excluding the component for storing information.

[0133] Memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode the stream is stored. Memory b2 may be an electronic circuit and may be connected to the processor b1. Memory b2 may also be included in the processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.

[0134] For example, memory b2 may store an image or a stream. Alternatively, memory b2 may store a program for processor b1 to decode the stream.

[0135] Furthermore, for example, memory b2 may play the role of an information storage component among the multiple components of the decoding device 200 shown in Figure 6 below. Specifically, memory b2 may play the role of a block memory 210 and a frame memory 214 shown in Figure 6 below. More specifically, memory b2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0136] The generated parameters disclosed in the embodiments may be stored in memory b2 and referenced in the processing of processor b1. The parameters may be generated by a decoding process or not.

[0137] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 6, etc., described later, nor does it need to perform all of the processes described above. Some of the components shown in Figure 6, etc., described later, may be included in other devices, and some of the processes described above may be performed by other devices.

[0138] The following describes the decoding device 200 in general, followed by a description of its components. Note that detailed explanations may be omitted for components of the decoding device 200 that perform the same processing as those in the encoding device 100.

[0139] [Example of Decoding Device Configuration] Figure 6 is a block diagram showing an example of the configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes a stream, which is an encoded image, in block units.

[0140] As shown in Figure 6, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, a prediction unit, a prediction control unit 220, a prediction parameter generation unit 222, and a division determination unit 224. The prediction unit includes an intra-prediction unit 216 and an inter-prediction unit 218.

[0141] These components are implemented, for example, by the processor b1 and memory b2 of the decoding device 200 shown in Figure 5 above.

[0142] Furthermore, components included in the decoding device 200 may perform the same processing as components included in the encoding device 100, and their explanation may be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, adder unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, adder unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120, respectively.

[0143] [Overall Decryption Process Flow] Figure 7A is a flowchart showing an example of the overall decryption process by the decryption device 200. Figure 7B is a flowchart showing an example of the block-by-block processing by the decryption device 200.

[0144] First, the entropy decoding unit 202 of the decoding device 200 acquires the stream output from the encoding device 100. Then, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the current block contained in the bitstream (step Sp_1a).

[0145] The division determination unit 224 determines the division pattern for each of the multiple fixed-size blocks (128 x 128 pixels) contained in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_2a). This division pattern is the division pattern selected by the encoding device 100. The decoding device 200 then performs block-specific processing for each of the multiple blocks. Specifically, it performs the processing in steps Sp_1 to Sp_5 for each of the multiple blocks (step Sp_3a).

[0146] The decoding device 200 then determines whether or not the decoding of the entire picture is complete (step Sp_4a). If it determines that it is not complete (No. in step Sp_4a), it repeats the process from step Sp_2a.

[0147] Next, we will explain the processing for each block (Sp_1 to Sp_5 in Figure 7B).

[0148] First, the inverse quantization unit 204 inversely quantizes the quantization coefficients of the current block, which are input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inversely quantizes the quantization coefficient based on the quantization parameter corresponding to that quantization coefficient (step Sp_1).

[0149] For example, the inverse quantization unit 204 may acquire a parameter indicating whether or not to perform inverse quantization, and quantization parameters (such as difference quantization parameters and QP index). The inverse quantization unit 204 may then decide whether or not to perform inverse quantization based on the acquired parameters, and may perform the inverse quantization process if it is decided to perform inverse quantization.

[0150] The inverse quantization unit 204 then outputs the inversely quantized quantization coefficients (i.e., transformation coefficients) of the current block to the inverse transformation unit 206.

[0151] Next, the inverse transform unit 206 restores the predicted residual by inversely transforming the transformation coefficients, which are input from the inverse quantization unit 204 (step Sp_2). Here, inverse transform refers to the reverse transformation process of the transformation process described in the encoding device 100. In other words, the inverse transform unit 206 can also be said to be performing a transformation process.

[0152] Next, the prediction unit, consisting of the intra-prediction unit 216, the inter-prediction unit 218, and the prediction control unit 220, generates a predicted image of the current block (step Sp_3). Here, the intra-prediction unit 216 and the inter-prediction unit 218 may perform the same processing as described above on the encoding device 100. Note that the predicted image needs to be generated before the addition process because it is used in the subsequent addition process, but it does not necessarily have to be done after the prediction residual restoration process. For example, the prediction image generation process may be performed in parallel with the prediction residual restoration process.

[0153] Next, the summing unit 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted image to the predicted residual (step Sp_4). In other words, the summing unit 208 generates a reconstructed image of the current block by adding the predicted image to the predicted residual. The summing unit 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.

[0154] The block memory 210 is a block referenced in intra prediction and is a storage unit for storing blocks within the current picture. Specifically, the block memory 210 stores the reconstructed image output from the adder 208.

[0155] The loop filter unit 212 performs filtering on the reconstructed image (step Sp_5). The filtered reconstructed image is output to the frame memory 214 and the display device, etc.

[0156] In Figures 6 and 7B, the loop filter unit 212 processes the reconstructed image within the loop; that is, it outputs the filtered reconstructed image to the frame memory 214. Some or all of the processing in the loop filter unit 212 may be performed outside the loop. That is, filtering may be performed before outputting to a display device or the like.

[0157] Furthermore, the processes in steps Sp_1a to Sp_4a and Sp_1 to Sp_5 may be performed sequentially by the decoding device 200, some of these processes may be performed in parallel, or the order of these processes may be changed.

[0158] [Another Example of System Configuration] Figure 8 is a schematic diagram showing another example of the configuration of the transmission system according to this embodiment. The transmission system Trs is a system that transmits generated streams and receives transmitted streams. Such a transmission system Trs includes, for example, a transmitting device 2000, a network Nw, and a receiving device 3000, as shown in Figure 8. Note that the transmission system Trs does not need to have all of these components and may consist of only some of the devices.

[0159] The network Nw may be the Internet, a wide-area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network transmitting broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, the network Nw may be replaced by a storage medium that records streams, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc®).

[0160] The transmitting device 2000 transmits the stream over the network Nw. The transmitting device 2000 may also take the encoded stream as input and output the acquired stream. Alternatively, the transmitting device 2000 may have the encoding processing configuration disclosed as the encoding device 100, take the original image before encoding as input, encode the acquired image to generate a stream, and transmit it.

[0161] When the transmitting device 2000 generates a stream, the original input image before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. The stream includes, for example, an encoded image and control information for decoding that encoded image. This encoding compresses the image.

[0162] The transmitting device 2000 can reduce the load on stream transmission by reducing the amount of data in the stream, which includes the encoded image.

[0163] The receiving device 3000 receives an encoded stream from the network Nw. The receiving device 3000 may store the received stream in memory, transfer the received stream to another device, or perform any processing on the received stream. For example, the receiving device 3000 may have a decoding process configuration disclosed as the decoding device 200 and decode the received stream.

[0164] The receiving device 3000 can reduce the delay of the stream by reducing the amount of data in the received stream.

[0165] [Transmitting Device] Figure 9 is a schematic diagram showing an example configuration of the transmitting device 2000 according to this embodiment. The transmitting device 2000 is a device that transmits a stream.

[0166] [Implementation Example of Transmitting Device] Next, a transmitting device 2000 according to the embodiment will be described. Figure 10 is a block diagram showing an implementation example of the transmitting device 2000. The transmitting device 2000 includes a processor a1 and a transmitting unit a3 that outputs a stream. The transmitting device 2000 may also further include a memory a2 as shown in Figure 2, or a receiving unit that acquires a video or stream.

[0167] Processor a1 is a circuit that performs information processing. Processor a1 may also be a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that transmits a stream. Processor a1 may also be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.

[0168] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to transmit a stream is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0169] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to send the stream.

[0170] The stream transmitted by the transmitting device 2000 may be the stream (also called a bitstream) described in this disclosure.

[0171] Furthermore, in the transmitting device 2000, the processor a1 may perform some or all of the roles of the transmitting unit a3. Alternatively, the transmitting device 2000 may not have a transmitting unit a3, and the processor a1 may output the stream.

[0172] [Receiving device] Figure 11 is a schematic diagram showing an example of the configuration of a receiving device 3000 according to this embodiment. The receiving device 3000 is a device that receives a stream.

[0173] [Implementation Example of Receiving Device] Next, a receiving device 3000 according to the embodiment will be described. Figure 12 is a block diagram showing an implementation example of the receiving device 3000. The receiving device 3000 includes a processor b1 and a receiving unit b4 that receives a stream. The receiving device 3000 may also further include a memory b2 as shown in Figure 5, or a transmitting unit that transmits a video or stream.

[0174] Processor b1 is a circuit that performs information processing. Processor b1 may also be a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that receives a stream. Processor b1 may also be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits.

[0175] Memory b2 is a dedicated or general-purpose memory that stores information for processor b1 to receive streams. Memory b2 may be an electronic circuit and may be connected to processor b1. Memory b2 may also be included in processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.

[0176] For example, memory b2 may store the input stream or the decoded image. Alternatively, memory b2 may store a program for processor b1 to receive the stream.

[0177] The stream received by the receiving device 3000 may be the stream (also called a bitstream) described in this disclosure.

[0178] Furthermore, in the receiving device 3000, the processor b1 may perform some or all of the functions of the receiving unit b4. Alternatively, the receiving device 3000 may not have a receiving unit b4, and the processor b1 may receive the stream.

[0179] [Bitstream Generation Device] Figure 13 is a schematic diagram showing an example of the configuration of a bitstream generation device according to this embodiment. The bitstream generation device 1000 is a device that generates a bitstream and transmits the generated stream.

[0180] An image is input to the bitstream generator 1000. The bitstream generator 1000 generates a stream by encoding the input image and outputs the stream. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.

[0181] The original image input to the bitstream generator 1000 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image.

[0182] Furthermore, an image is a broader concept than sequences, pictures, and blocks, and is not limited to spatial or temporal domains unless otherwise specified. An image consists of a sequence of pixels or pixel values, and the signal or pixel values ​​representing that image are also called samples. A stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal.

[0183] Furthermore, the bitstream generation device may also be called the encoding device, image encoding device, or video encoding device, and the method for generating a bitstream by the bitstream generation device may be called the encoding method, image encoding method, or video encoding method.

[0184] [Implementation Example of Bitstream Generation Device] Next, a bitstream generation device 1000 according to an embodiment will be described. Figure 14 is a block diagram showing an implementation example of the bitstream generation device 1000. The bitstream generation device 1000 includes a processor a1 and a memory a2. The bitstream generation device 1000 may also include an input unit for inputting moving images, and an output unit for outputting the generated bitstream.

[0185] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that generates a bitstream. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.

[0186] Memory a2 is a dedicated or general-purpose memory in which information for processor a1 to generate a bitstream is stored. Memory a2 may be an electronic circuit and may be connected to processor a1. Memory a2 may also be included in processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.

[0187] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to generate the bitstream.

[0188] The bitstream generator 1000 may have the same configuration as the encoding device 100 in this disclosure. Specifically, the bitstream generator 1000 may implement the multiple components shown in Figure 3. However, not all of the multiple components shown in Figure 3 are to be implemented, nor are all of the multiple processes to be performed. Some of the multiple components shown in Figure 3 may be included in other devices, and some of the multiple processes described above may be performed by other devices.

[0189] Furthermore, for example, processor a1 may perform the role of multiple components of the encoding device 100 shown in Figure 3, excluding the component for storing information. Also, for example, memory a2 may perform the role of the component for storing information among the multiple components of the encoding device 100 shown in Figure 3.

[0190] Specifically, memory a2 may function as the block memory 118 and frame memory 122 shown in Figure 3. More specifically, memory a2 may store reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.).

[0191] Furthermore, during operation, processor a1 uses memory a2 to generate a bitstream that includes parameters for causing the decoding device 200 to execute processing, and / or parameters for switching the processing of the decoding device 200 or other specific processing. For example, processor a1 generates these parameters, includes the generated parameters in the bitstream, and generates the bitstream.

[0192] Here, the "parameters" to be executed by the decoding device 200 may be parameters for executing any decoding process exemplified in this disclosure as a process to be executed. For example, the "parameters" to be executed by the decoding device 200 may be any parameter in the syntax described in this disclosure. Furthermore, the "process of the decoding device 200" or "other specific process" that can be switched may be any decoding process exemplified in this disclosure as a process performed by obtaining the syntax.

[0193] Furthermore, the generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be included in the bitstream. Whether or not to include them in the bitstream is appropriately determined based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.

[0194] [Storage Medium] Figure 15A is a schematic diagram showing an example of the configuration of a storage medium and a computer according to this embodiment. The storage medium 4000 is a medium for storing bitstreams. The storage medium 4000 is connected to a processor b1 included in the computer 5000 and outputs a bitstream, which is an executable instruction for the computer 5000, to the computer 5000. For example, when the computer 5000 reads the bitstream from the storage medium 4000, the storage medium 4000 outputs the bitstream to the computer 5000.

[0195] The encoded stream stored in the storage medium 4000 is an instruction that the computer 5000 can execute. The processor b1 in the computer 5000 generates a reconstructed image by executing a bitstream decoding process based on the acquired bitstream. In this way, the computer 5000, having acquired a stream from the storage medium 4000, can decode the bitstream based on the parameters and other information contained in the encoded stream.

[0196] Furthermore, the storage medium 4000 may be the memory b2 described herein. For example, the memory b2 may store a program for the processor b1 to generate a bitstream. The storage medium 4000 may be a non-transitory computer-readable medium. The storage medium 4000 is also referred to as a recording medium.

[0197] The storage medium 4000 may include an input / output unit for inputting and outputting bitstreams. The storage medium 4000 may also include an input unit for receiving bitstreams and an output unit for outputting bitstreams.

[0198] Furthermore, the computer 5000 may also be the decoding device 200. In other words, the computer 5000 may be read as the decoding device 200.

[0199] Figure 15B is a schematic diagram showing an example of the configuration of a computer according to this embodiment. The schematic diagram shown in Figure 15B differs from the example in Figure 15A in that the storage medium 4000 is included in the computer 5000. The storage medium 4000 and the processor b1 may be the same as those described in Figure 15A.

[0200] Here, the bitstream includes parameters that indicate instructions for the computer 5000 to execute. The "parameters" may include parameters that cause the computer 5000 to perform processing, and / or parameters that switch the processing or other specific processing of the computer 5000.

[0201] The “parameter” may be any parameter that performs any decryption process exemplified in this disclosure as an example of a process to be performed. For example, it may be any parameter in the syntax described in this disclosure. The process to be performed and the process to be switched may be any decryption process exemplified in this disclosure as a process performed by obtaining the syntax.

[0202] [Examples of Parameters] In the above, for the bitstream generator 1000, "parameters that cause the decoding device 200 to execute processing, and / or parameters that switch the processing of the decoding device 200 or other specific processing" are described. Also, for the storage medium 4000, "parameters that cause the computer 5000 to execute processing, and / or parameters that switch the processing of the computer 5000 or other specific processing" are described. These parameters may be, for example, the following parameters.

[0203] The "parameters" may include, for example, parameters indicating the method of dividing the block. The computer 5000 may divide the block based on the parameters indicating the method of dividing the block and perform any of the processes illustrated in this disclosure on the divided blocks.

[0204] The "parameters" may include, for example, parameters indicating the prediction mode to be applied. The computer 5000 may determine the prediction mode to be applied based on the parameters indicating the prediction mode to be applied, and may perform any prediction process exemplified in this disclosure corresponding to the determined prediction mode.

[0205] The "parameters" may include, for example, a parameter indicating whether or not process y can be performed within range x. The computer 5000 may determine whether or not process y can be performed on a target block based on the parameter indicating whether or not process y can be performed within range x.

[0206] A flag (parameter) indicating whether or not process y can be performed within range x is shown, for example, as xxx_yyy_enabled_flag / xxx_yyy_disabled_flag. Here, yyy may represent process y, which is indicated by the flag as either or not. Process y may also be any process exemplified in this disclosure.

[0207] xxx may represent the level to which parameters are assigned. For example, xxx may correspond to sps, pps, ph, sh, cu (Coding Unit), or block, and this level may indicate the range x. More specifically, if it is sps, the range x is a sequence; if it is pps, the range x is a picture; if it is ph, the range x is a picture; and if it is sh, the range x is a slice. Note that the levels to which parameters are assigned are not limited to these.

[0208] If the parameter indicates that process y is not feasible, then process y will not be performed in the blocks included in range x. If the parameter indicates that process y is feasible, then process y is feasible in the blocks included in range x. In other words, if the parameter indicates that process y is feasible, then process y may or may not be performed in the blocks included in range x.

[0209] The flag indicating whether or not process y can be performed may be defined for multiple headers. For example, process y may be permitted for a sequence, but not for a particular picture within that sequence.

[0210] In the above case, the sps_yyy_enabled_flag (or sps_yyy_disabled_flag) for the sequence may have a value indicating that process y can be performed. Furthermore, the ph_yyy_enabled_flag (or ph_yyy_disabled_flag) for a certain picture included in the sequence may have a value indicating that process y cannot be performed.

[0211] As described above, the levels correspond to sps, pps, ph, sh, cu (Coding Unit), or blocks, etc. An example definition for each level is shown below. Note that the definitions for each level can be changed.

[0212] The description of `sps` (Sequence Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more CLVS (Coded Layer Video Sequences). Here, the number of zero or more CLVSs is determined by the content of the syntax elements contained in the PPS referenced by the syntax elements contained in each picture header. In other words, `sps` corresponds to the sequence level.

[0213] The description of pps (Picture Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more encoded pictures. Here, zero or more encoded pictures are determined by the syntax elements contained in each picture header. In other words, pps corresponds to the picture level.

[0214] ph (PH: Picture Header) corresponds to a syntax structure containing syntax elements that apply to all slices of an encoded picture. In other words, ph corresponds to the picture level.

[0215] The entry for sh (SH: Slice Header) corresponds to the portion of the encoded slice that contains all tiles or data elements related to CTU rows within a slice. In other words, sh corresponds to the slice level.

[0216] The notation `cu` (Coding Unit) corresponds to the coding block of the sample and the syntax structure used to encode that sample. In other words, `cu` corresponds to the coding unit level or the block level.

[0217] Here, the above coding blocks correspond to (i) the coding block for the luminance sample and the two corresponding coding blocks for the chrominance sample of a picture having three sample sequences in single-tree mode, (ii) the coding block for the luminance sample of a picture having three sample sequences in dual-tree mode, (iii) the two coding blocks for the chrominance sample of a picture having three sample sequences in dual-tree mode, or (iv) the coding block for the sample of a monochrome picture.

[0218] Although a flag indicating whether or not process y can be performed has been described above, a flag indicating whether or not to perform process y (xxx_yyy_flag) may be used in a similar manner. Furthermore, the flag indicating whether or not process y can be performed and the flag indicating whether or not to perform process y may be used in combination.

[0219] [Data Structure] Figure 16 shows an example of the hierarchical structure of data in a stream. A stream includes, for example, a video sequence. This video sequence includes, for example, a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and multiple pictures, as shown in Figure 16(a).

[0220] VPS includes encoding parameters common to multiple layers in a video composed of multiple layers, and encoding parameters related to the multiple layers included in the video, or to individual layers.

[0221] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the picture. Multiple SPSs may exist.

[0222] The PPS includes parameters used for the picture, i.e., encoding parameters that the decoding device 200 references to decode each picture in the sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. There may be multiple PPSs. Also, the SPS and PPS are sometimes simply referred to as parameter sets.

[0223] The picture may include a picture header and one or more slices, as shown in Figure 16(b). The picture header includes encoding parameters that the decoding device 200 refers to in order to decode the one or more slices.

[0224] A slice includes a slice header and one or more Coding Tree Units (CTUs), as shown in Figure 16(c). The slice header includes coding parameters that the decoding device 200 references to decode the one or more CTUs.

[0225] A picture may not contain slices, but instead contain tile groups. In this case, a tile group may contain one or more tiles, and each tile may contain one or more CTUs. Furthermore, a tile may contain one or more slices, and a slice may contain one or more tiles.

[0226] A CTU is also called a superblock or basic partitioning unit. Such a CTU includes a CTU header and one or more CUs (Coding Units), as shown in Figure 16(d). The CTU header contains coding parameters that the decoding device 200 refers to in order to decode one or more CUs.

[0227] A CU may be divided into multiple smaller CUs. Furthermore, as shown in Figure 16(e), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating the predicted residual, which will be described later.

[0228] Note that a CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but in SBT, for example, as described later, it may include multiple TUs smaller than the CU. Also, a CU may be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes that CU. A VPDU is a fixed unit that can be processed in one stage when performing pipeline processing in hardware, for example.

[0229] Note that a stream does not necessarily have to have some of the hierarchical levels shown in Figure 16. Also, the order of these hierarchical levels may be changed, and any hierarchical level may be replaced by another hierarchical level.

[0230] Furthermore, the picture that is currently being processed by a device such as the encoding device 100 or the decoding device 200 is called the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded; if the processing is decoding, the current picture is synonymous with the picture to be decoded.

[0231] Furthermore, a block, such as CU or CU, that is currently being processed by a device such as an encoding device 100 or a decoding device 200 is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded; if the processing is decoding, the current block is synonymous with the block to be decoded.

[0232] [Picture Composition: Slices / Tiles] To decode pictures in parallel, pictures may be composed of slices or tiles.

[0233] A slice is the basic encoding unit that makes up a picture. A picture is composed of, for example, one or more slices. A slice consists of one or more consecutive CTUs.

[0234] Figure 17 shows an example of a slice configuration. For example, a picture contains 11 x 8 CTUs and is divided into four slices (slices 1-4). Slice 1 consists of, for example, 16 CTUs, slice 2 consists of, for example, 21 CTUs, slice 3 consists of, for example, 29 CTUs, and slice 4 consists of, for example, 22 CTUs. Here, each CTU in the picture belongs to one of the slices.

[0235] The shape of a slice is a horizontal division of the picture. The boundaries of a slice do not have to be at the edges of the screen, but can be anywhere among the CTU boundaries within the screen. The processing order (encoding order or decoding order) of the CTUs within a slice is, for example, the raster scan order. A slice also includes a slice header and encoded data. The slice header may describe the characteristics of the slice, such as the CTU address of the beginning of the slice and the slice type.

[0236] A tile is a rectangular area that makes up a picture. Each tile may be assigned a number called a TileId in the order of the raster scan.

[0237] Figure 18 shows an example of a tile configuration. For example, a picture contains 11 x 8 CTUs and is divided into four rectangular tile regions (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used.

[0238] If tiles are not used, multiple CTUs within a picture are processed, for example, in raster scan order. If tiles are used, at least one CTU in each of the multiple tiles is processed, for example, in raster scan order. For example, as shown in Figure 18, the processing order of the multiple CTUs contained in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0239] Note that one tile may contain one or more slices, and one slice may contain one or more tiles.

[0240] A picture may be composed of tilesets. A tileset may contain one or more tile groups, or one or more tiles. A picture may consist of only one of a tileset, a tile group, or a tile. For example, the order in which multiple tiles are scanned in raster order for each tileset is defined as the basic coding order of the tiles. Within each tileset, a collection of one or more tiles whose basic coding order is consecutive is defined as a tile group. Such a picture may be composed of the division unit 102 (see Figure 3) described later.

[0241] [Temporal Scalable Encoding] Figure 19 shows an example of a temporally scalable stream configuration.

[0242] The encoding device 100 may generate a temporally scalable stream by encoding multiple pictures in multiple layers, as shown in Figure 19. For example, the encoding device 100 achieves scalability by encoding each picture in its own layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called temporally scalable encoding.

[0243] This allows the decoding device 200 to switch the frame rate of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.

[0244] As a result, the decoding device 200 can freely switch between decoding the same content in low-frame-rate and high-frame-rate formats. For example, a user of the stream can watch part of the video using their smartphone while on the go, and then watch the rest of the video using an internet TV or other device after returning home.

[0245] Each of the aforementioned smartphones and devices incorporates a decoding device 200, which may or may not have the same performance. In this case, if the device decodes up to the upper layers of the stream, the user can view high-frame-rate video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different frame rates, thereby reducing the processing load.

[0246] In temporally scalable coding, layers are sometimes referred to as temporal layers or temporal sublayers.

[0247] [Multilayer Encoding] Figure 20 shows an example of a stream configuration using multilayer encoding. Each of AU0 to AU3 corresponds to an AU (Access Unit). Each of Layer0 to Layer3 corresponds to a layer in multilayer encoding. Each PU (Picture Unit) corresponds to a picture. Each of POC_1 to POC_3 corresponds to a POC (Picture Of Count) and indicates the display order.

[0248] As shown in Figure 20, the encoding device 100 may generate a stream in which spatial resolution, image quality, or multiplexed content can be scalably changed by encoding multiple pictures into multiple layers. For example, the encoding device 100 achieves scalability by encoding pictures layer by layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called multi-layer encoding.

[0249] This allows the decoding device 200 to switch the spatial resolution, image quality, or multiplexed content of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.

[0250] As a result, the decoding device 200 can freely switch between decoding low-resolution and high-resolution content, low-quality and high-quality content, and basic and customized content for the same content. For example, a user of the stream can watch part of the video on their smartphone while on the go, and then watch the rest of the video on an internet TV or other device after returning home.

[0251] Furthermore, each of the aforementioned smartphones and devices incorporates a decoding device 200 with identical or different performance characteristics. In this case, if the device decodes up to the upper layers of the stream, the user can view high-definition video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different image quality, thereby reducing the processing load.

[0252] [Metadata] Each layer of the stream may include metadata based on statistical information of the image. The decoding device 200 may generate high-resolution video by super-resolution the pictures of each layer based on the metadata. Super-resolution may be either an improvement in the signal-to-noise ratio (SN) at the same resolution, or an increase in resolution. The metadata may include information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values ​​in the filtering process, machine learning, or least-squares operation used in the super-resolution process.

[0253] Alternatively, the picture may be divided into tiles or the like, depending on the meaning of each object within the picture. In this case, the decoding device 200 may decode only a portion of the picture by selecting the tiles to be decoded.

[0254] Furthermore, object attributes (such as person, car, or ball) and their position within the picture (such as their coordinate position within the same picture) may be stored as metadata. In this case, the decoding device 200 can identify the position of a desired object based on the metadata and determine the tile containing that object. For example, the metadata is stored using a data storage structure different from pixel data, such as SEI in HEVC. This metadata may indicate, for example, the position, size, or color of the main object.

[0255] Furthermore, metadata may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. This allows the decoding device 200 to obtain information such as the time when a specific person appears in the video, and by using that time and the information in the picture units, it can identify the picture in which the object exists and the position of the object within that picture.

[0256] [Dividing section] The dividing section 102 first divides the picture into blocks of a fixed size (e.g., 128 x 128 pixels). Then, the dividing section 102 divides each of the fixed-size blocks into blocks of a variable size (e.g., 64 x 64 pixels or less) based on, for example, a recursive quadtree and / or binary tree block partitioning.

[0257] Figure 21 shows an example of block partitioning in the embodiment. In Figure 21, solid lines represent block boundaries due to quadtree block partitioning, and dashed lines represent block boundaries due to binary tree block partitioning.

[0258] Here, block 10 is a 128x128 pixel square block. This block 10 is first divided into four 64x64 pixel square blocks (quadtree block partitioning).

[0259] The top-left 64x64 pixel square block is further divided vertically into two rectangular blocks, each consisting of 32x64 pixels, and the left 32x64 pixel rectangular block is further divided vertically into two rectangular blocks, each consisting of 16x64 pixels (binary tree block partitioning). As a result, the top-left 64x64 pixel square block is divided into two 16x64 pixel rectangular blocks 11 and 12 and a 32x64 pixel rectangular block 13.

[0260] The 64x64 pixel square block in the upper right is horizontally divided into two rectangular blocks 14 and 15, each consisting of 64x32 pixels (binary tree block partitioning).

[0261] The 64x64 pixel square block in the lower left is divided into four square blocks, each consisting of 32x32 pixels (quadtree block partitioning). Of these four 32x32 pixel square blocks, the upper left and lower right blocks are further divided.

[0262] The 32x32 pixel square block in the upper left is vertically divided into two rectangular blocks, each consisting of 16x32 pixels. The rectangular block on the right, each consisting of 16x32 pixels, is further horizontally divided into two 16x16 pixel square blocks (binary tree block partitioning). The 32x32 pixel square block in the lower right is horizontally divided into two rectangular blocks, each consisting of 32x16 pixels (binary tree block partitioning).

[0263] As a result, the 64x64 pixel square block in the lower left is divided into a 16x32 pixel rectangular block 16, two 16x16 pixel square blocks 17 and 18, two 32x32 pixel square blocks 19 and 20, and two 32x16 pixel rectangular blocks 21 and 22.

[0264] The 64x64 pixel block 23 in the lower right corner will not be divided.

[0265] As described above, in Figure 21, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.

[0266] In Figure 21, one block was divided into four or two blocks (quadrutree or binary tree block partitioning), but partitioning is not limited to these. For example, one block may be divided into three blocks (ternary tree block partitioning). Partitioning that includes such ternary tree block partitioning is sometimes called MTT (Multi-Type Tree) partitioning.

[0267] Furthermore, while this example demonstrates a binary tree block partitioning method where a block is divided into two rectangular blocks of the same size in a 1:1 ratio, binary tree block partitioning is not limited to this. For example, a block may be divided into two rectangular blocks of different sizes in a 1:2 or 1:3 ratio. In such cases, parameters indicating these ratios may be used as block partitioning information.

[0268] Figure 22 shows an example of the configuration of the division unit 102. As shown in Figure 22, the division unit 102 may include a block division determination unit 102a. The block division determination unit 102a may perform the following processing as an example.

[0269] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the division pattern based on that block information. The division unit 102 divides the original image according to the division pattern and outputs one or more blocks obtained by the division to the subtraction unit 104.

[0270] Furthermore, the block division determination unit 102a outputs parameters indicating the division pattern described above to the conversion unit 106, the inverse conversion unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy coding unit 110. The conversion unit 106 may convert the prediction residuals based on these parameters, and the intra prediction unit 124 and the inter prediction unit 126 may generate a prediction image based on these parameters. The entropy coding unit 110 may also perform entropy coding on these parameters.

[0271] The parameters related to the splitting pattern may be written to the stream as follows, for example.

[0272] Figure 23 shows examples of division patterns. Division patterns include, for example, quadripartition (QT), which divides the block into two horizontally and two vertically; tripartition (HT or VT), which divides the block in the same direction in a 1:2:1 ratio; duplicate division (HB or VB), which divides the block in the same direction in a 1:1 ratio; and no division (NS).

[0273] Note that in the case of four divisions and no division, the division pattern does not have a block division direction, while in the case of two divisions and three divisions, the division pattern has division direction information.

[0274] Figures 24A and 24B show examples of syntax trees for split patterns. In the example in Figure 24A, first there is information indicating whether or not to perform a split (S: Split flag), then there is information indicating whether or not to perform a four-way split (QT: QT flag). Next there is information indicating whether to perform a three-way split or a two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag).

[0275] Furthermore, the same division process may be repeatedly applied to each of the one or more blocks obtained by the division using such a division pattern. That is, as an example, the determination of whether or not to divide, whether or not to divide into four, whether or not to divide horizontally or vertically, and whether or not to divide into three or two may be performed recursively, and the results of the determinations performed may be encoded into a stream according to the encoding order disclosed in the syntax tree shown in Figure 24A.

[0276] Furthermore, in the syntax tree shown in Figure 24A, the information is arranged in the order of S, QT, TT, and Ver, but it may also be arranged in the order of S, QT, Ver, and BT. In other words, in the example in Figure 24B, first there is information indicating whether or not to perform a split (S: Split flag), then there is information indicating whether or not to perform a four-way split (QT: QT flag). Next there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag), and finally there is information indicating whether to perform a two-way split or a three-way split (BT: BT flag or TT: TT flag).

[0277] Note that the division patterns described here are just examples; you may use other division patterns, or only a part of the division patterns described.

[0278] [Loop Filtering Section in Encoding Process] The loop filter section 120 in the encoding device 100 applies a filter to the reconstructed image output from the adder 116 and outputs the filtered reconstructed image to the frame memory 122. Here, the filter performed by the loop filter section 120 is a filter used within the encoding loop (also called a loop filter or in-loop filter).

[0279] Filtering is a technique that reduces encoding noise generated by quantization processing within the encoding loop. Filtering can directly reduce visually noticeable block distortion or ringing distortion, improving subjective and / or objective performance. Furthermore, the loop filter unit 120 prevents image quality degradation from propagating between frames by performing filtering within the loop.

[0280] Filters include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), sample adaptive offset (SAO), luminance mapping with chroma scaling (LMCS), or any combination thereof.

[0281] Figure 25 is a block diagram showing an example of the configuration of the loop filter unit 120. The loop filter unit 120 includes, for example, an LMCS processing unit 120d, a DBF processing unit 120a, a SAO processing unit 120b, and an ALF processing unit 120c, as shown in Figure 25.

[0282] The LMCS processing unit 120d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 120a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 120b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 120c applies ALF processing to the reconstructed image after SAO processing. Details of ALF and DBF will be described later.

[0283] SAO processing is a process that improves image quality by reducing ringing (a phenomenon in which pixel values ​​are distorted in a wave-like manner around edges) and correcting pixel value misalignment. Examples of SAO processing include edge offset processing and band offset processing.

[0284] LMCS processing is a process that reassigns codewords assigned to rarely occurring pixel values ​​to codewords assigned to frequently occurring pixel values ​​when there is a large bias in the pixel value distribution of the luminance signal of the original image. Examples of LMCS processing include mapping the luminance signal using a piecewise linear model based on the pixel value distribution of the original image, and scaling the residual color difference signal according to the luminance pixel values.

[0285] Furthermore, the loop filter unit 120 does not necessarily have to include all of the processing units disclosed in Figure 25, and may include only some of them. Also, the loop filter unit 120 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 25.

[0286] [Loop Filter Section > Adaptive Loop Filter] In ALF, a least-squares error filter is applied to remove encoding distortion. For example, for each 2x2 pixel subblock within the current block, one filter selected from several filters is applied based on the direction and activity of the local gradient.

[0287] Specifically, first, subblocks (for example, 2x2 pixel subblocks) are classified into multiple classes (for example, 15 or 25 classes). The classification of subblocks is done, for example, based on the direction of the gradient and the activity level. In a specific example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the subblocks are classified into multiple classes.

[0288] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). The gradient activation value A is derived, for example, by adding the gradients in multiple directions and quantizing the sum.

[0289] Based on the results of this classification, a filter for the subblock is determined from among multiple filters.

[0290] For example, circularly symmetrical shapes are used for filters in ALF. Figures 26A to 26C show several examples of filter shapes used in ALF. Figure 26A shows a 5x5 diamond-shaped filter, Figure 26B shows a 7x7 diamond-shaped filter, and Figure 26C shows a 9x9 diamond-shaped filter.

[0291] Information indicating the shape of the filter is typically signaled at the picture level. However, the signaling of information indicating the shape of the filter is not limited to the picture level; it may be at other levels (e.g., sequence level, slice level, brick level, CTU level, or CU level).

[0292] The on / off status of the ALF may be determined, for example, at the picture level or the CU level. For example, the decision to apply the ALF may be made at the CU level for luminance, and at the picture level for color difference. Information indicating whether the ALF is on or off is usually signaled at the picture level or the CU level.

[0293] Note that the signaling of ALF on / off information is not limited to picture level or CU level, but may be at other levels (e.g., sequence level, slice level, brick level, or CTU level). If the ALF on / off information read from the stream indicates that ALF is on, one filter is selected from several filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.

[0294] Furthermore, as described above, one filter is selected from among several filters and ALF processing is applied to the subblock. For each of these filters (for example, up to 15 or 25 filters), the set of coefficients used in that filter is usually signaled at the picture level. However, the signaling of the coefficient set is not limited to the picture level and may be at other levels (for example, sequence level, slice level, brick level, CTU level, CU level, or subblock level).

[0295] [Loop Filter Section > CCALF] CC-ALF (Cross Component Adaptive Loop Filter) is a type of restoration filter that uses correlations between color components. The filtering results of eight luminance signals corresponding to the positions of the color difference signal being processed are added to the color difference signal. To enable parallel processing with ALF, the luminance signal before the application of ALF is used for filtering.

[0296] Figure 26D shows a configuration diagram of an example of a CCALF in this disclosure. Figure 26E shows an example of the filter shape of a CCALF in this disclosure.

[0297] One example of CC-ALF operates by applying a linear diamond-shaped filter (Figures 26D and 26E) to the luminance channel of each chromatic difference component. For example, the filter coefficients are transmitted via APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and indicated by a context-encoded flag received for each block of samples. The block size and CC-ALF enable flag are received at the slice level for each chromatic difference component.

[0298] The syntax and semantics of CC-ALF are provided in the encoding standard. Block sizes of 16x16, 32x32, 64x64, and 128x128 may also be supported for color difference samples.

[0299] [Loop Filter Section > DBF] In DBF processing, the loop filter section 120 reduces distortion occurring at block boundaries by applying a filter to the block boundaries of the reconstructed image.

[0300] Figure 27A is a block diagram showing an example of the detailed configuration of the DBF processing unit 120a. The DBF processing unit 120a includes, for example, a boundary determination unit 1201, a filter determination unit 1203, a filter processing unit 1205, a processing determination unit 1208, a filter characteristic determination unit 1207, and switches 1202, 1204, and 1206.

[0301] Figure 27B is a flowchart showing an example of the processing performed by the DBF processing unit 120a.

[0302] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0b).

[0303] Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1b). If it determines to perform DBF, it proceeds to the next determination. If the DBF processing unit 120a determines not to perform DBF, it terminates the process without performing DBF (step S_2b).

[0304] Specifically, for example, the DBF processing unit 120a may determine in the boundary determination unit 1201 whether or not the target pixel for DBF processing is located near a block boundary, and if the target pixel is not located near a block boundary, it may decide not to perform DBF. Alternatively, the DBF processing unit 120a may calculate a Bs value and decide not to perform DBF according to the Bs value. For example, if the Bs value is equal to 0, it indicates that DBF will not be performed, and if it is 1 or greater, it indicates that there is a possibility of performing DBF.

[0305] Next, the DBF processing unit 120a determines whether or not to perform DBF on the target pixel in the filter determination unit 1203 (step S_3b). In other words, the DBF processing unit 120a determines the type of filter to be applied. Specifically, for example, the filter determination unit 1203 determines whether or not to perform DBF processing on the target pixel based on the pixel values ​​of at least one surrounding pixel located around the target pixel. The image before filtering is an image consisting of the target pixel and at least one surrounding pixel located around that target pixel.

[0306] The filter determination unit 1203 may compare a value based on the pixel value of at least one peripheral pixel, and / or the difference between multiple pixels, etc., with a threshold value, and decide whether to execute the DBF to be determined if the value is greater than the threshold value, and whether to execute the DBF to be determined if the value is less than or equal to the threshold value. The threshold value is, for example, β or tC determined based on the quantization parameter QP. In other words, in the DBF process, for example, the DBF to be performed is selected based on quantization.

[0307] In one example of the process for determining whether or not to perform the DBF to be judged, it is first determined whether or not to perform DBF(a). If it is determined that DBF(a) should be performed, DBF(a) is performed (step S_4b). If it is determined that DBF(a) should not be performed, the next determination is made. In the next determination, it is determined whether or not to perform DBF(b). If it is determined that DBF(b) should be performed, DBF(b) is performed (step S_5b). If it is determined that DBF(b) should not be performed, DBF is not performed (step S_2b).

[0308] Furthermore, if it is decided not to implement DBF(b), there may be multiple DBF candidates, such as deciding whether or not to implement DBF(c). Here, DBF(a), which has been decided to implement, may be further subdivided, and it may be decided whether or not to implement DBF(a_1) or DBF(a_2).

[0309] Here, DBF(a), DBF(b), DBF(a_1), and DBF(a_2) may be short filter, long filter, weak filter, and strong filter, respectively.

[0310] Figure 27C is a flowchart showing another example of the processing performed by the DBF processing unit 120a.

[0311] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0c). Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1c), and if it determines to perform DBF, it selects the DBF to perform (step S_2c). Then, the DBF processing unit 120a performs the selected DBF (step S_3c). If the DBF processing unit 120a determines in S_1c not to perform DBF, it terminates the process without performing DBF (step S_4c).

[0312] Note that the processing of S_1c may be the same as or different from the processing of S_1b. Also, the processing of S_2c may be the same as or different from the processing of S_3b.

[0313] Figure 28 shows an example of a DBF (Digital Block Filter) with symmetrical filter characteristics with respect to block boundaries. DBF is a process that reduces distortion at block boundaries by correcting the sample values ​​on both sides of the block boundary. The number of pixels to be filtered and the number of pixels used for filtering vary depending on the type of filter. The higher the filter strength, the greater the number of pixels to be filtered and the greater the number of pixels used for filtering.

[0314] When DBF processing is performed on the block boundary between block P and block Q adjacent to block P, the pixels to be processed are pn (n=0, ..., n) in block P and qm (m=0, ..., m) in block Q. Pixels with a value of n or m of 0 are adjacent to the boundary, while pixels with larger values ​​are further from the boundary.

[0315] For example, in a strong filter, as shown in Figure 28, processing is performed on pixels p0 to p2 in block P and pixels q0 to q2 in block Q. The respective pixel values ​​of pixels q0 to q2 are changed to pixel values ​​q'0 to q'2 by performing the calculation shown in the following equation.

[0316] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8 q'1=(p0+q0+q1+q2+2) / 4 q'2=(p0+q0+q1+3×q2+2×q3+4) / 8

[0317] In the above equations, p0 to p2 and q0 to q2 are the pixel values ​​of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the opposite side of the block boundary. Furthermore, the coefficient multiplied by the pixel value of each pixel used in the deblocking filter process on the right-hand side of each of the above equations is the filter coefficient.

[0318] Furthermore, in DBF processing, clipping may be performed to ensure that the pixel value after calculation does not change beyond a threshold. In this clipping process, the pixel value after calculation using the above formula is clipped to "pre-calculation pixel value ± a × threshold (where a is an integer)" using a threshold determined from the quantization parameters. In other words, the amount of change in the pixel value before and after calculation is clipped so that it is within the range of ± a × threshold (where a is an integer). This prevents excessive smoothing. Specifically, for example, the pixel values ​​q'0 to q'2 are expressed by the following formula.

[0319] q'0=Clip3 (q0-3×tC, q0+3×tC, (p1+2×p0+2×q0+2×q1+q2+4)>>3) q'1=Clip3 (q1-2×tC, q1+2×tC, (p0+q0+q1+q2+2)>>2) q'2=Clip3 (q2-1×tC, q2+1×tC, (p0+q0+q1+3×q2+2×q3+4)>>3)

[0320] Furthermore, Clip3 may be defined as follows:

[0321]

[0322] Figure 29 is a diagram illustrating an example of a block boundary where DBF processing is performed. Figure 30 is a diagram illustrating an example of a block boundary strength Bs value.

[0323] The block boundaries on which DBF processing is performed are, for example, the CU, PU, ​​or TU boundaries of an 8x8 pixel block as shown in Figure 29. DBF processing is performed in units of, for example, 4 rows or 4 columns. First, for blocks P and Q shown in Figure 29, the Boundary Strength (Bs) value is determined as shown in Figure 30. The Bs value is determined independently for each color component.

[0324] In Figure 30, conditions listed higher up have higher priority, and evaluation is performed in order from the highest priority condition. If a condition is met, evaluation of lower priority conditions is not performed. The magnitude and range of block distortion differ depending on the boundary, and excessive smoothing can cause problems such as blurring. Therefore, in this example, the Bs value is determined according to the prediction modes of blocks P and Q, the presence or absence of motion vectors and transformation coefficients.

[0325] Furthermore, based on the determined Bs value, it may be decided whether or not to perform DBF processing of different strengths even for block boundaries belonging to the same image. DBF processing for chrominance signals is performed when the Bs value is 2. DBF processing for luminance signals is performed when the Bs value is 1 or greater and predetermined conditions are met. Note that the criteria for determining the Bs value are not limited to the example shown in Figure 30, and may be determined based on other parameters.

[0326] The above method for determining the Bs value and the method for performing DBF processing based on the Bs value are examples only, and the method for determining the Bs value and the method for performing DBF processing based on the Bs value are not limited to the above example.

[0327] In DBF processing, it is determined whether or not to apply a filter to each of the luminance sample and chrominance sample, and the type of filter to be applied. Furthermore, in DBF processing, it is determined whether or not to apply a filter to each of the vertical and horizontal boundaries contained in each of the luminance sample and chrominance sample, and the type of filter to be applied.

[0328] Some or all of the processing may be performed on either the luminance sample or the chrominance sample, or on both. The luminance sample and the chrominance sample may undergo the same processing, or they may undergo different processing.

[0329] [Loop Filtering Unit in Decoding Process] The loop filtering unit 212 in the decoding device 200 applies a loop filter to the reconstructed image generated by the summing unit 208, and outputs the filtered reconstructed image to the frame memory 214 and the display device, etc.

[0330] Figure 31 is a block diagram showing an example of the configuration of the loop filter unit 212. The loop filter unit 212 has a configuration similar to that of the loop filter unit 120 of the encoding device 100. The loop filter unit 212 includes, for example, an LMCS processing unit 212d, a DBF processing unit 212a, a SAO processing unit 212b, and an ALF processing unit 212c, as shown in Figure 31.

[0331] The LMCS processing unit 212d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 212a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 212b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 212c applies ALF processing to the reconstructed image after SAO processing.

[0332] Furthermore, the loop filter unit 212 does not necessarily have to include all of the processing units disclosed in Figure 31, and may include only some of them. Also, the loop filter unit 212 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 31.

[0333] [Prediction Unit (Intra Prediction Unit, Inter Prediction Unit, Prediction Control Unit)] Figure 32 is a flowchart showing an example of processing performed in the prediction unit of the encoding device 100. For example, the prediction unit consists of all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.

[0334] The prediction unit generates a predicted image of the current block (step Sb_1). The predicted image may be, for example, an intra-prediction image (intra-prediction signal) or an inter-prediction image (inter-prediction signal). Specifically, the prediction unit generates a predicted image of the current block using a reconstructed image already obtained by generating predicted images for other blocks, generating prediction residuals, generating quantization coefficients, restoring prediction residuals, and adding the predicted images.

[0335] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.

[0336] Figure 33 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.

[0337] The prediction unit generates a predicted image using a first method (step Sc_1a), a second method (step Sc_1b), and a third method (step Sc_1c). The first, second, and third methods are different methods for generating predicted images, and may be, for example, an interpretation method, an intraprediction method, and other prediction methods. These prediction methods may use the reconstructed images described above.

[0338] Next, the prediction unit evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit calculates a cost C for each of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, and evaluates the predicted images by comparing the costs C of those predicted images.

[0339] The cost C is calculated using the R-D optimization model formula, for example, C = D + λ × R. In this formula, D is the coding distortion of the predicted image, which can be expressed as, for example, the sum of the absolute differences between the pixel values ​​of the current block and the pixel values ​​of the predicted image. R is the bitrate of the stream, and λ is, for example, the Lagrange multiplier.

[0340] Next, the prediction unit selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the prediction unit selects a method or mode for obtaining the final predicted image. For example, the prediction unit selects the predicted image with the smallest cost C based on the cost C calculated for those predicted images. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be based on parameters used in the coding process.

[0341] The encoding device 100 may signal information to identify the selected prediction image, scheme, or mode into a stream. This information may be, for example, a flag. Based on this information, the decoding device 200 can generate a prediction image according to the scheme or mode selected by the encoding device 100.

[0342] In the example shown in Figure 33, the prediction unit generates predicted images using each method and then selects one of the predicted images. However, the prediction unit may also select a method or mode based on the parameters used in the encoding process described above before generating those predicted images, and then generate the predicted images according to that method or mode.

[0343] For example, the first method and the second method are intra-prediction and inter-prediction, respectively, and the prediction unit may select the final predicted image for the current block from the predicted images generated according to these prediction methods.

[0344] Figure 34 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.

[0345] First, the prediction unit generates a predicted image by intra-prediction (step Sd_1a) and then generates a predicted image by inter-prediction (step Sd_1b). The predicted image generated by intra-prediction is also called an intra-prediction image, and the predicted image generated by inter-prediction is also called an inter-prediction image.

[0346] Next, the prediction unit evaluates both the intra-predicted image and the inter-predicted image (step Sd_2). The cost C mentioned above may be used for this evaluation. The prediction unit may then select the prediction image with the smallest cost C from the intra-predicted image and the inter-predicted image as the final prediction image for the current block (step Sd_3). In other words, a prediction method or mode for generating the prediction image for the current block is selected.

[0347] [Prediction Control Unit] The prediction control unit 128 selects either an intra-prediction image (an image or signal output from the intra-prediction unit 124) or an inter-prediction image (an image or signal output from the inter-prediction unit 126), and outputs the selected prediction image to the subtraction unit 104 and the addition unit 116.

[0348] [Prediction Parameter Generation Unit] The prediction parameter generation unit 130 may output information regarding intra-prediction, inter-prediction, and the selection of a predicted image in the prediction control unit 128 as prediction parameters to the entropy coding unit 110. The entropy coding unit 110 may generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by the decoding device 200.

[0349] The decoding device 200 may receive and decode the stream and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.

[0350] The prediction parameters may include a selected prediction signal (e.g., MV, prediction type, or prediction mode used in the intra-prediction unit 124 or the inter-prediction unit 126), or any index, flag, or value that is based on or indicates the prediction processing performed in the intra-prediction unit 124, the inter-prediction unit 126, and the prediction control unit 128.

[0351] [Interpretation Unit] The interpretation unit 126 generates a predicted image (interpretation image) by performing interpretation (also called inter-screen prediction) of the current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture.

[0352] Interpretation is performed on a current block or current subblock within a current block. A subblock is a smaller unit contained within a block. The size of a subblock can be 4x4 pixels, 8x8 pixels, or other sizes. The size of a subblock may be switched to units such as slices, bricks, or pictures.

[0353] For example, the interpretation unit 126 performs motion estimation within the reference picture for the current block or current subblock and finds the reference block or subblock that best matches the current block or current subblock. Then, the interpretation unit 126 obtains motion information (e.g., motion vectors) that compensates for the movement or change from the reference block or subblock to the current block or subblock.

[0354] The inter-prediction unit 126 performs motion compensation (or motion prediction) based on the motion information and generates an inter-prediction image of the current block or sub-block. The inter-prediction unit 126 outputs the generated inter-prediction image to the prediction control unit 128.

[0355] Motion compensation is a process that generates a predicted image using one or more reference pictures and one or more motion vectors. The predicted image is generated by interpolation, which uses sample values ​​from the region specified by the motion vector as the region corresponding to the current block in the reference picture and its surrounding region.

[0356] Furthermore, if the horizontal and vertical elements of the motion vector are represented by integer values, the sample values ​​of the corresponding regions may be used directly without interpolation. Also, in motion compensation, the predicted image may be generated using the reconstructed image contained in the current picture without using one or more reference pictures.

[0357] In this disclosure, the term motion compensation is used, but this term may be replaced with the term interpolation. This replacement may be applied throughout the specification.

[0358] The motion information used for motion compensation may be signaled as an interprediction image in various forms. For example, the motion vector may be signaled. As another example, the difference between the motion vector and the predicted motion vector may be signaled.

[0359] [Reference Picture List] Figure 35 is a diagram showing an example of each reference picture, and Figure 36 is a conceptual diagram showing an example of a reference picture list. The reference picture list is a list showing one or more reference pictures stored in the frame memory 122. In Figure 35, rectangles represent pictures, arrows indicate the reference relationships between pictures, the horizontal axis represents time, I, P, and B in the rectangles represent intra-prediction pictures, single-prediction pictures, and double-prediction pictures, respectively, and the numbers in the rectangles indicate the decoding order.

[0360] As shown in Figure 35, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As shown in Figure 36, the reference picture list is a list representing candidates for reference pictures, and for example, one picture (or slice) may have one or more reference picture lists. For example, if the current picture is a single-prediction picture, one reference picture list is used, and if the current picture is a double-prediction picture, two reference picture lists are used.

[0361] In the examples in Figures 35 and 36, picture B3, which is the current picture currPic, has two reference picture lists, the L0 list and the L1 list. When the current picture currPic is picture B3, the candidate reference pictures for that current picture currPic are I0, P1, and B2, and each reference picture list (i.e., the L0 list and the L1 list) points to these pictures.

[0362] The interpretation unit 126 or the prediction control unit 128 specifies which picture in each reference picture list to actually reference using the reference picture index refIdxLx. In Figure 36, reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.

[0363] Such reference picture lists may be generated on a sequence, picture, slice, brick, CTU, or CU basis. Furthermore, the reference picture index indicating the reference picture referenced in interpretation among the reference pictures shown in the reference picture list may be encoded at the sequence, picture, slice, brick, CTU, or CU level. Additionally, a common reference picture list may be used across multiple interpretation modes.

[0364] [Generated Reference Picture] Figure 37A is a conceptual diagram showing another example of a reference picture. As illustrated in Figure 37A, a new image (PicA') generated based on an image (PicA) may be used as a reference picture. A reference picture generated in this way is defined as a generated reference picture. A generated reference picture may be used as a reference picture in the prediction process.

[0365] Additionally, the generated reference picture may be added to the same reference picture list as the reference picture list to which other reference pictures are added. Alternatively, a separate reference picture list may be generated, and the generated reference picture may be added to that separate reference picture list.

[0366] Figure 37B is a conceptual diagram illustrating another example of generating a generated reference picture. In the example in Figure 37B, a new image (PicAB') is generated by inputting two images (PicA and PicB) into an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit). The generated new image (PicAB') is then used as a generated reference picture for interpretation.

[0367] Interpretation using this generated reference picture may also be called neural network intercoding (NN inter). Note that the NPU or GPU is not limited to two images, but can receive two or more images as input.

[0368] Figure 37C is a conceptual diagram illustrating another example of generating a generated reference picture. As illustrated in Figure 37C, the size of the input image and the size of the generated image may differ. In this example, the size of the generated image (PicA') is smaller than the size of the input image (PicA). The size of the generated image is not limited to this example and may be larger than the size of the input image. Image processing or the size of the generated image per frame may be based on the performance of the NPU or GPU. This size information may also be included in the bitstream.

[0369] Note that in the above example, a new reference picture (also called a generated image or a processed image) is generated from one or more encoded or decoded reference pictures using a neural network. The method for generating the new reference picture is not limited to the above example. For example, a new reference picture may be generated using a process such as image conversion, and the generated new reference picture may be defined as a generated reference picture.

[0370] By adding these generated reference pictures to the reference picture list so that they can be referenced, there is a possibility of improving the encoding efficiency.

[0371] Note that in the above example, images are generated in picture units, but the generation unit is not limited to the above example. The generated reference data may be generated in a predetermined unit such as a block, a slice, or a tile included in the picture.

[0372] [Basic Flow of Inter Prediction] FIG. 38 is a flowchart showing the basic flow of inter prediction.

[0373] The inter prediction unit 126 first generates a prediction image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the prediction image as a prediction residual (step Se_4).

[0374] Here, in generating the prediction image, the inter prediction unit 126 generates the prediction image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). Also, in determining the MV, the inter prediction unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2).

[0375] The selection of a candidate MV is performed, for example, by the interpretation unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Note that previously derived MVs may be added to the candidate MV list as candidate MVs. Furthermore, in the MV derivation process, the interpretation unit 126 may determine the selected at least one candidate MV as the MV for the current block by selecting at least one more candidate MV from the at least one candidate MV.

[0376] Alternatively, the interpretation unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by each of the selected at least one candidate MVs. This search of the region of the reference picture may be called motion estimation.

[0377] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the interpretation unit 126, but processing such as step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.

[0378] Furthermore, a candidate MV list may be created for each process in each interpretation mode, or a common candidate MV list may be used across multiple interpretation modes. Also, the processes in steps Se_3 and Se_4 correspond to the processes in steps Sa_3 and Sa_4 shown in Figure 4B, respectively. In addition, the process in step Se_3 corresponds to the process in step Sd_1b in Figure 34.

[0379] [MV Derivation Flowchart] Figure 39 is a flowchart showing an example of MV derivation.

[0380] The interpretation unit 126 may derive the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information may be encoded as prediction parameters and then signaled. That is, the encoded motion information is included in the stream.

[0381] Alternatively, the interpretation unit 126 may derive MV in a mode that does not encode motion information. In this case, motion information is not included in the stream.

[0382] Here, the modes for MV derivation include the normal intermode, normal merge mode, FRUC mode, and affine mode, which will be described later. Of these modes, the modes that encode motion information include the normal intermode, normal merge mode, and affine mode (specifically, the affine intermode and affine merge mode). Note that motion information may include not only MV but also the predicted MV selection information described later. Modes that do not encode motion information include the FRUC mode, among others.

[0383] The interpretation unit 126 selects a mode from these multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.

[0384] Figure 40 is a flowchart showing another example of MV derivation.

[0385] The interpretation unit 126 may derive the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. This differential MV is the difference between the MV of the current block and its predicted MV. The predicted MV is the predicted motion vector.

[0386] Alternatively, the interpretation unit 126 may derive MV in a mode that does not encode the differential MV. In this case, the encoded differential MV is not included in the stream.

[0387] As mentioned above, the modes for MV derivation include the normal inter, normal merge mode, FRUC mode, and affine mode, which will be described later. Of these modes, the modes that encode differential MV include the normal inter mode and affine mode (specifically, affine inter mode). Modes that do not encode differential MV include the FRUC mode, normal merge mode, and affine mode (specifically, affine merge mode).

[0388] The interpretation unit 126 selects a mode from these multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.

[0389] [Modes of MV Derivation] Figures 41A and 41B show an example of the classification of each mode of MV derivation. For example, as shown in Figure 41A, the modes of MV derivation can be broadly classified into three modes depending on whether motion information is encoded and whether differential MV is encoded. The three modes are intermode, merge mode, and FRUC (frame rate up-conversion) mode. Intermode is a mode that performs motion search and encodes motion information and differential MV.

[0390] For example, as shown in Figure 41B, the intermode includes the affine intermode and the normal intermode. The merge mode is a mode in which motion search is not performed, and the MV is selected from the surrounding encoded blocks, and the MV of the current block is derived using that MV. This merge mode is basically a mode in which motion information is encoded, but the differential MV is not encoded.

[0391] For example, as shown in Figure 41B, the merge modes include normal merge mode (sometimes called regular merge mode), MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined inter merge / intra prediction) mode, GPM mode, ATMVP mode, and affine merge mode. Here, among the modes included in the merge modes, the MMVD mode is an exception in which the differential MV is encoded.

[0392] The aforementioned affine merge mode and affine inter mode are modes included in affine mode. Affine mode is a mode that assumes an affine transformation and derives the MV of each of the multiple subblocks that make up the current block as the MV of the current block. FRUC mode is a mode that derives the MV of the current block by performing a search between encoded regions, and does not encode either motion information or differential MV. Details of each of these modes will be described later.

[0393] Note that the classification of modes shown in Figures 41A and 41B is merely an example and is not limited to this classification. For example, if the differential MV is encoded in CIIP mode, that CIIP mode is classified as an intermode. Also, the MV derivation mode may include GPM (Geometric Partitioning Mode). GPM may also be expressed as geometric shape partitioning prediction mode or GPM mode.

[0394] [MV Derivation > Normal Intermode] Normal intermode is an interpretation mode that derives the MV of the current block by finding blocks similar to the image of the current block from the region of the reference picture indicated by the candidate MV. In this normal intermode, the differential MV is also encoded.

[0395] Figure 42 is a flowchart showing an example of inter-mode prediction.

[0396] The interpretation unit 126 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks surrounding the current block in time or space (step Sg_1). In other words, the interpretation unit 126 creates a candidate MV list.

[0397] Next, the interpretation unit 126 extracts N candidate MVs (where N is an integer greater than or equal to 2) from the multiple candidate MVs obtained in step Sg_1 as predicted MV candidates, according to a predetermined priority order (step Sg_2). The priority order is predetermined for each of the N candidate MVs.

[0398] Next, the interpretation unit 126 selects one predicted MV candidate from the N predicted MV candidates as the predicted MV for the current block (step Sg_3). At this time, the interpretation unit 126 encodes predicted MV selection information into a stream to identify the selected predicted MV. In other words, the interpretation unit 126 outputs the predicted MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0399] Next, the interpretation unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the interpretation unit 126 further encodes the difference between the derived MV and the predicted MV as the difference MV into the stream. In other words, the interpretation unit 126 outputs the difference MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. The encoded reference picture is a picture consisting of multiple blocks that have been reconstructed after encoding.

[0400] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5).

[0401] Steps Sg_1 to Sg_5 are performed for each block. For example, once steps Sg_1 to Sg_5 have been performed for each of the blocks contained in a slice, the inter-mode prediction for that slice is complete. Similarly, once steps Sg_1 to Sg_5 have been performed for each of the blocks contained in a picture, the inter-mode prediction for that picture is complete.

[0402] Note that the processing in steps Sg_1 to Sg_5 is not performed for all blocks included in a slice; if it is performed for some blocks, the inter prediction using the normal inter mode for that slice may be terminated. Similarly, if the processing in steps Sg_1 to Sg_5 is performed for some blocks included in a picture, the inter prediction using the normal inter mode for that picture may be terminated.

[0403] The predicted image is the interprediction signal described above. Furthermore, information indicating the interprediction mode used to generate the predicted image (normal intermode in the example above), which is included in the encoded signal, is encoded, for example, as a prediction parameter.

[0404] The candidate MV list may be the same as the list used in other modes. Furthermore, processing related to the candidate MV list may be applied to processing related to lists used in other modes. This processing related to the candidate MV list may include, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.

[0405] [MV Derivation > Normal Merge Mode] The normal merge mode is an inter prediction mode that derives an MV by selecting a candidate MV from the candidate MV list as the MV of the current block. Note that the normal merge mode is a merge mode in a narrow sense and may simply be called the merge mode. In the present embodiment, the normal merge mode and the merge mode are distinguished, and the merge mode may be used in a broad sense. The candidate MVs of the merge mode are also called merge candidates or merge mode candidates.

[0406] Figure 43 is a flowchart showing an example of inter prediction by the normal merge mode.

[0407] First, the inter prediction unit 126 obtains a plurality of candidate MVs for the current block based on information such as the MVs of a plurality of encoded blocks temporally or spatially around the current block (step Sh_1). That is, the inter prediction unit 126 creates a candidate MV list.

[0408] Next, the inter prediction unit 126 derives the MV of the current block by selecting one candidate MV from the plurality of candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the inter prediction unit 126 encodes the MV selection information for identifying the selected candidate MV into the stream. That is, the inter prediction unit 126 outputs the MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0409] Finally, the inter prediction unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).

[0410] Steps Sh_1 to Sh_3 are executed for each block, for example. For instance, once steps Sh_1 to Sh_3 have been executed for each block in a slice, the inter prediction using normal merge mode for that slice is complete. Similarly, once steps Sh_1 to Sh_3 have been executed for each block in a picture, the inter prediction using normal merge mode for that picture is complete.

[0411] Note that the processing in steps Sh_1 to Sh_3 is not performed on all blocks included in a slice; if it is performed on some blocks, the inter prediction using normal merge mode for that slice may be terminated. Similarly, if the processing in steps Sh_1 to Sh_3 is performed on some blocks included in a picture, the inter prediction using normal merge mode for that picture may be terminated.

[0412] Furthermore, information indicating the inter-prediction mode used to generate the predicted image (normal merge mode in the example above), which is included in the stream, is encoded, for example, as prediction parameters.

[0413] Figure 44 is a diagram illustrating an example of the MV derivation process for the current picture using normal merge mode.

[0414] First, the interpretation unit 126 generates a candidate MV list in which candidate MVs are registered. Candidate MVs include spatially adjacent candidate MVs, which are the MVs of multiple encoded blocks located spatially around the current block; temporally adjacent candidate MVs, which are the MVs of nearby blocks projected onto the current block's position in the encoded reference picture; combined candidate MVs, which are MVs generated by combining the MV values ​​of spatially adjacent candidate MVs and temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero.

[0415] Next, the interpretation unit 126 selects one candidate MV from among the multiple candidate MVs registered in the candidate MV list, thereby determining that one candidate MV as the MV for the current block.

[0416] Furthermore, the entropy coding unit 110 writes and encodes merge_idx, a signal indicating which candidate MV was selected, into the stream.

[0417] Note that the candidate MVs registered in the candidate MV list explained in Figure 44 are just an example, and the number of candidates may differ from the number shown in the figure, the configuration may not include some of the candidate MV types shown in the figure, or it may include candidate MV types other than those shown in the figure.

[0418] The final MV may be determined by performing DMVR (decoder motion vector refinement), described later, using the MV of the current block derived by normal merge mode. Note that in normal merge mode, the differential MV is not encoded, but in MMVD mode, the differential MV is encoded.

[0419] The MMVD mode selects one candidate MV from a list of candidate MVs, similar to the normal merge mode, but encodes a differential MV. Such MMVD may be classified as a merge mode along with the normal merge mode, as shown in Figure 41B. Note that the differential MV in MMVD mode does not have to be the same as the differential MV used in inter-mode; for example, the derivation of the differential MV in MMVD mode may be a less computationally intensive process than the derivation of the differential MV in inter-mode.

[0420] Alternatively, a Combined Inter Merge / Intra Prediction (CIIP) mode may be used, which generates a prediction image for the current block by overlaying the prediction image generated by inter prediction with the prediction image generated by intra prediction.

[0421] The candidate MV list may also be referred to as the candidate list. Furthermore, merge_idx is the MV selection information.

[0422] Furthermore, in the MV derivation process, including the merge mode, a process to correct the MV of the candidate MV may be performed.

[0423] For example, template matching may be performed on multiple candidate MVs by referencing the reconstructed image of the encoded block and the encoded reference picture, and a correction process may be performed on the candidate MV list. Next, the MV of the current block may be derived by selecting one candidate MV from the multiple corrected candidate MVs. Note that the correction process is not limited to this example and may be performed using other methods. Also, the correction process is not limited to being applied to normal merge mode.

[0424] For example, correction processing may be applied when correcting candidate MVs in other modes. Other modes may include template matching (TM) mode, bilateral matching (BM) merge mode, MMVD (Merge with Motion Vector Difference) merge mode, intrablock copy (IBC) merge mode, CIIP (Combined inter merge / intra prediction) merge mode, GPM merge mode, or affine merge mode.

[0425] Here, the intrablock copy (IBC) merge mode is a mode in which a predicted image is obtained by copying prediction blocks from the encoded or decoded peripheral region of the same picture. The CIIP (Combined inter merge / intra prediction) merge mode is a mode in which the inter-prediction image generated in the normal merge mode and the intra-prediction image generated in the planar prediction mode are combined using a weighted average.

[0426] Furthermore, correction processes such as DMVR correction and template matching correction, as described later, may be applied. The correction processes are not limited to these methods. Other correction processes may be used, or multiple correction processes may be used in combination.

[0427] Furthermore, new merge candidates may be generated and added to the merge mode candidate list in Figure 44. For example, the motion vector mvA and reference picture index refIdxA of the encoded or decoded blocks adjacent to the block to be encoded or decoded are added as spatial adjacency candidates MV. In this case, new merge candidates may be generated using mvB and refidxB of the referenced block referenced using mvA and refIdxA.

[0428] This makes it possible to generate a candidate list with highly accurate motion vectors for the reference picture, which can improve encoding efficiency.

[0429] For example, as described above, a reference block is calculated from the motion vector and reference picture index of the merge candidates included in the merge mode candidate list, and a new merge candidate is generated using the motion vector and reference picture index of that reference block. This makes it possible to generate motion vectors for previously encoded or decoded reference pictures with high accuracy while tracking the motion, thereby improving encoding efficiency.

[0430] The merge candidates calculated as described above are sometimes called CMVP (Chained motion vector prediction) merge candidates.

[0431] Furthermore, the method of generating merge candidates using CMVP is not limited to normal merge mode; it may also be applied to other merge modes, or to the generation of predicted motion vectors, etc. For example, CMVP may be applied to affine merge mode. This may improve the prediction accuracy of affine prediction and improve coding efficiency.

[0432] Furthermore, if the reference block belongs to an intrapicture or intraslice, or if the reference block is a block encoded by intraprediction, the reference block may not have a motion vector. In this case, it is not possible to generate CMVP merge candidates, and therefore, it is acceptable that no CMVP merge candidates are generated from that merge candidate.

[0433] Furthermore, if a referenced block has a block vector (bv) which is a motion vector for blocks within the same picture, CMVP merge candidates may be generated using bv.

[0434] Furthermore, the reference blocks within the generated reference picture mentioned above do not contain motion vectors or reference picture information (such as the reference picture index). Therefore, if a reference block within the generated reference picture is selected as the reference target for generating a CMVP merge candidate, it is not possible to generate a CMVP merge candidate. Thus, it is acceptable that no CMVP merge candidates are generated from that merge candidate.

[0435] Alternatively, the motion vector may be tracked and added until the reference block no longer has motion vectors or reference picture information, and a CMVP merge candidate may be generated using that motion vector and reference picture information. This may improve encoding efficiency.

[0436] [MV Derivation > HMVP Mode] Figure 45 is a diagram illustrating an example of the MV derivation process for the current picture using HMVP mode.

[0437] In normal merge mode, the MV of the current block (e.g., CU) is determined by selecting one candidate MV from a list of candidate MVs generated by referencing the encoded block (e.g., CU). Other candidate MVs may be registered in this list. This mode, where other candidate MVs are registered, is called HMVP mode.

[0438] In HMVP mode, candidate MVs are managed using a separate FIFO (First-In First-Out) buffer for HMV, distinct from the candidate MV list used in normal merge mode.

[0439] The FIFO buffer stores motion information such as MV for blocks that have been processed in the past, in reverse chronological order. In the management of this FIFO buffer, each time a block is processed, the MV of the most recent block (i.e., the most recently processed CU) is stored in the FIFO buffer, and in its place, the MV of the oldest CU in the FIFO buffer (i.e., the CU that was processed first) is deleted from the FIFO buffer. In the example shown in Figure 45, HMVP1 is the MV of the most recent block, and HMVP5 is the MV of the oldest block.

[0440] For example, the interpretation unit 126 checks each MV managed in the FIFO buffer, starting with HMVP1, whether that MV is different from all the candidate MVs already registered in the normal merge mode candidate MV list. If the interpretation unit 126 determines that it is different from all the candidate MVs, it may add the MV managed in the FIFO buffer as a candidate MV to the normal merge mode candidate MV list.

[0441] At this time, one or more candidate MVs may be registered from the FIFO buffer.

[0442] By using HMVP mode in this way, it becomes possible to include not only MVs of spatially or temporally adjacent blocks to the current block, but also MVs of previously processed blocks as candidates. As a result, the variety of candidate MVs in normal merge mode is broadened, which increases the likelihood of improving encoding efficiency.

[0443] Furthermore, the aforementioned MV may also be motion information. In other words, the information stored in the candidate MV list and FIFO buffer may include not only the MV value, but also information indicating the referenced picture, the direction and number of referenced pictures, etc. Also, the aforementioned block is, for example, a CU.

[0444] Note that the candidate MV list and FIFO buffer in Figure 45 are just examples, and the candidate MV list and FIFO buffer may be lists or buffers of different sizes than those in Figure 45, or the candidate MVs may be registered in a different order than those in Figure 45. Furthermore, the processing described here is common to both the encoding device 100 and the decoding device 200.

[0445] Furthermore, HMVP mode can be applied to modes other than normal merge mode. For example, movement information such as MV of blocks previously processed in affine mode can be stored in the FIFO buffer in chronological order from newest to oldest and used as candidate MV. A mode in which HMV mode is applied to affine mode may be called history affine mode.

[0446] [MV Derivation > FRUC Mode] Motion information may be derived on the decoding device 200 side without being signaled from the encoding device 100 side. For example, motion information may be derived by performing a motion search on the decoding device 200 side. In this case, the decoding device 200 side performs a motion search without using the pixel values ​​of the current block. Modes for performing such a motion search on the decoding device 200 side include the FRUC (frame rate up-conversion) mode or the PMMVD (pattern matched motion vector derivative) mode.

[0447] An example of FRUC processing is shown in Figure 46. First, a list is generated that refers to the MVs of each encoded block that is spatially or temporally adjacent to the current block and indicates those MVs as candidate MVs (i.e., a candidate MV list, which may be the same as the candidate MV list for normal merge mode) (step Si_1).

[0448] Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and based on that evaluation value, one candidate MV is selected as the best candidate MV.

[0449] Then, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block may be derived by performing pattern matching in the area surrounding the position in the reference picture corresponding to the selected best candidate MV.

[0450] In other words, a search is performed in the area surrounding the best candidate MV using pattern matching and evaluation values ​​in the reference picture. If an MV with a better evaluation value is found, the best candidate MV may be updated to that MV and made the final MV for the current block. It is not necessary to update to an MV with a better evaluation value.

[0451] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).

[0452] Steps Si_1 to Si_5 are performed for each block, for example. For instance, once steps Si_1 to Si_5 have been performed for each block in a slice, the inter prediction using FRUC mode for that slice is complete. Similarly, once steps Si_1 to Si_5 have been performed for each block in a picture, the inter prediction using FRUC mode for that picture is complete.

[0453] Note that the processes in steps Si_1 to Si_5 are not performed on all blocks included in a slice; if they are performed on some blocks, the inter prediction using FRUC mode for that slice may be terminated. Similarly, if the processes in steps Si_1 to Si_5 are performed on some blocks included in a picture, the inter prediction using FRUC mode for that picture may be terminated.

[0454] Subblocks may be processed in the same way as blocks as described above.

[0455] The evaluation value may be calculated by various methods. For example, the reconstructed image of a region in the reference picture corresponding to the MV may be compared with the reconstructed image of a predetermined region (which may be, for example, a region in another reference picture or a region in an adjacent block of the current picture, as shown below). The difference in pixel values ​​between the two reconstructed images may then be calculated and used as the evaluation value for the MV. In addition to the difference value, other information may also be used to calculate the evaluation value.

[0456] Next, we will explain pattern matching in detail. First, one candidate MV included in the candidate MV list (also called the merge list or merge candidate list) is selected as the starting point for the search using pattern matching.

[0457] For pattern matching, either first-order pattern matching or second-order pattern matching may be used. First-order pattern matching and second-order pattern matching are sometimes called bilateral matching and template matching, respectively.

[0458] [MV Derivation > FRUC > Bilateral Matching] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are aligned with the motion trajectory of the current block. Therefore, in the first pattern matching, a region in the other reference picture aligned with the motion trajectory of the current block is used as a predetermined region for calculating the evaluation value of the candidate MV described above.

[0459] Figure 47 illustrates an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in Figure 47, in first pattern matching, two MVs (MV0, MV1) are derived by searching for the most matching pair of two blocks in two different reference pictures (Ref0, Ref1) that are along the motion trajectory of the current block (Cur block).

[0460] Specifically, for the current block, the difference is derived between the reconstructed image at a specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval. An evaluation value is then calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs should be selected as the best candidate MV.

[0461] Under the assumption of a continuous motion trajectory, the MV (MV0, MV1) pointing to two reference blocks is proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, then the first pattern matching derives a mirror-symmetric bidirectional MV.

[0462] [MV Derivation > FRUC > Template Matching] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (blocks adjacent to the current block in the current picture (e.g., blocks above and / or to the left)) and the block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the candidate MV described above.

[0463] Figure 48 illustrates an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in Figure 48, in the second pattern matching, the MV of the current block is derived by searching in the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic).

[0464] Specifically, for the current block, the difference between the reconstructed image of either or both of the left-adjacent and upper-adjacent encoded regions and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and an evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs should be selected as the best candidate MV.

[0465] Information indicating whether or not to apply such a FRUC mode (e.g., called the FRUC flag) may be signaled at the CU level. Furthermore, if the FRUC mode is applied (e.g., the FRUC flag is true), information indicating the applicable pattern matching method (first pattern matching or second pattern matching) may be signaled at the CU level.

[0466] Furthermore, the signaling of this information is not limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or subblock level).

[0467] [MV Derivation > Affine Mode] The affine mode is a mode in which the MV is generated using an affine transformation. For example, the MV may be derived on a subblock basis based on the MVs of multiple adjacent blocks. This mode is sometimes called the affine motion compensation prediction mode.

[0468] Figure 49A is a diagram illustrating an example of deriving the MV of a subblock based on the MV of multiple adjacent blocks.

[0469] In Figure 49A, the current block includes, for example, 16 subblocks consisting of 4x4 pixels. Here, the motion vector v of the upper left corner control point of the current block is determined based on the MV of the adjacent blocks. 0 The following is derived, and similarly, the motion vector v of the upper right corner control point of the current block based on the MV of the adjacent subblock. 1 The following is derived. Then, by the following equation (A), the two motion vectors v 0 and v 1 Project the motion vector (v) of each subblock within the current block. x ,v y ) is derived.

[0470]

[0471] Here, x and y represent the horizontal and vertical positions of the subblock, respectively, and w represents a predetermined weighting coefficient.

[0472] Information indicating such an affine mode (e.g., called an affine flag) may be signaled at the CU level. However, the signaling of this information indicating an affine mode is not limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or subblock level).

[0473] Furthermore, such affine modes may include several modes with different methods for deriving the MV of the upper-left and upper-right corner control points. For example, there are two affine modes: the affine inter (also called the affine normal inter) mode and the affine merge mode.

[0474] Figure 49B illustrates an example of deriving the subblock unit MV in affine mode using three control points.

[0475] In Figure 49B, the current block includes, for example, 16 subblocks consisting of 4x4 pixels. Here, the motion vector v of the upper left corner control point of the current block is determined based on the MV of the adjacent blocks.0 is derived. Similarly, based on the MVs of adjacent blocks, the motion vector v of the upper right control point of the current block 1 is derived, and based on the MVs of adjacent blocks, the motion vector v of the lower left control point of the current block 2 is derived.

[0476] Then, by the following formula (B), the three motion vectors v 0 , v 1 and v 2 are projected to derive the motion vectors (v x , v y ) of each sub-block within the current block.

[0477]

[0478] Here, x and y respectively indicate the horizontal position and vertical position of the center of the sub-block, and w and h indicate predetermined weight coefficients. w may indicate the width of the current block, and h may indicate the height of the current block.

[0479] The affine mode using different numbers of control points (for example, two and three) may be switched and signaled at the CU level. Note that the information indicating the number of control points of the affine mode used at the CU level may be signaled at other levels (for example, sequence level, picture level, slice level, block level, CTU level or sub-block level).

[0480] Also, the affine mode having such three control points may include several modes with different derivation methods of the MVs of the upper left, upper right and lower left control points. For example, the affine mode having three control points includes two modes, the affine inter mode and the affine merge mode, similar to the affine mode having two control points described above.

[0481] Note that in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.

[0482] [MV Derivation > Affine Mode > Control Point] Figures 50A, 50B, and 50C are conceptual diagrams illustrating an example of MV derivation for a control point in affine mode.

[0483] In affine mode, as shown in Figure 50A, for example, the predicted MV for each control point of the current block is calculated based on multiple MVs corresponding to blocks encoded in affine mode from among the encoded blocks A (left), B (top), C (upper right), D (lower left), and E (upper left) adjacent to the current block.

[0484] Specifically, these blocks are examined in the order of encoded block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left), and the first valid block encoded in affine mode is identified. Based on the multiple MVs corresponding to this identified block, the MV of the control point of the current block is calculated.

[0485] For example, as shown in Figure 50B, if block A adjacent to the left of the current block is encoded in affine mode with two control points, the motion vector v projected onto the upper left and upper right corners of the encoded block containing block A is... 3 and v 4 The following is derived. And the derived motion vector v 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The result is calculated.

[0486] For example, as shown in Figure 50C, if block A adjacent to the left of the current block is encoded in affine mode with three control points, then the motion vector v projected onto the upper left, upper right, and lower left corners of the encoded block containing block A is... 3 , v 4 and v 5 The following is derived. And the derived motion vector v3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 The result is calculated.

[0487] The MV derivation method shown in Figures 50A to 50C may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.

[0488] Figures 51A and 51B are conceptual diagrams illustrating another example of the derivation of the control point MV in affine mode.

[0489] Figure 51A is a diagram illustrating an affine mode having two control points.

[0490] In this affine mode, as shown in Figure 51A, the MV selected from the respective MVs of the encoded blocks A, B, and C adjacent to the current block is the motion vector v of the upper left corner control point of the current block. 0 It is used as follows. Similarly, the MV selected from the respective MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v of the upper right corner control point of the current block. 1 It is used as such.

[0491] Figure 51B is a diagram illustrating an affine mode having three control points.

[0492] In this affine mode, as shown in Figure 51B, the MV selected from the respective MVs of the encoded blocks A, B, and C adjacent to the current block is the motion vector v of the upper left corner control point of the current block. 0 It is used as such.

[0493] Similarly, the MV selected from the respective MVs of the encoded blocks D and E adjacent to the current block is the motion vector v of the upper right corner control point of the current block. 1 It is used as follows. Furthermore, the MV selected from the respective MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v of the lower left corner control point of the current block. 2 It is used as such.

[0494] The MV derivation method shown in Figures 51A and 51B may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.

[0495] Here, for example, when switching between affine modes with different numbers of control points (e.g., two and three) at the CU level to create a signal, the number of control points may differ between the encoded block and the current block.

[0496] Figures 52A and 52B are conceptual diagrams illustrating an example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block.

[0497] For example, as shown in Figure 52A, the current block is encoded in affine mode, having three control points at the upper left corner, upper right corner, and lower left corner, while block A, adjacent to the left of the current block, has two control points.

[0498] In this case, the motion vector v is projected onto the upper-left and upper-right corners of the encoded block containing block A. 3 and v 4 The following is derived. And the derived motion vector v 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The following is calculated. Furthermore, the derived motion vector v 0and v 1 From there, the motion vector v of the lower left corner control point. 2 This is calculated.

[0499] For example, as shown in Figure 52B, the current block is encoded in affine mode, having two control points at the upper left corner and the upper right corner, and block A adjacent to the left of the current block has three control points.

[0500] In this case, the motion vector v is projected onto the upper-left, upper-right, and lower-left corners of the encoded block containing block A. 3 , v 4 and v 5 The following is derived. And the derived motion vector v 3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The result is calculated.

[0501] The MV derivation method shown in Figures 52A and 52B may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.

[0502] [MV Derivation > Affine Mode > Affine Merge Mode] Figure 53 is a flowchart showing an example of the affine merge mode.

[0503] In affine merge mode, the interpretation unit 126 first derives the MV for each control point of the current block (step Sk_1). The control points are the upper left and upper right corners of the current block, as shown in Figure 49A, or the upper left, upper right, and lower left corners of the current block, as shown in Figure 49B. At this time, the interpretation unit 126 may encode MV selection information into a stream to identify two or three of the derived MVs.

[0504] For example, when using the MV derivation method shown in Figures 50A to 50C, the interpretation unit 126 examines the encoded blocks in the order of A (left), B (top), C (upper right), D (lower left), and E (upper left), as shown in Figure 50A, and identifies the first valid block encoded in affine mode.

[0505] The interpretation unit 126 derives the MV of the control point using the first valid block encoded in the identified affine mode. For example, if block A is identified and block A has two control points, as shown in Figure 50B, the interpretation unit 126 derives the motion vectors v of the upper left and upper right corners of the encoded block containing block A. 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 Calculate the result.

[0506] For example, the interpretation unit 126 generates motion vectors v of the upper left and upper right corners of the encoded block. 3 and v 4 By projecting this onto the current block, the motion vector v of the upper left corner control point of the current block is obtained. 0 And the motion vector v of the upper right corner control point 1 Calculate the result.

[0507] Alternatively, if block A is identified and block A has three control points, as shown in Figure 50C, the interpretation unit 126 predicts the motion vectors v of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. 3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 Calculate the result.

[0508] For example, the interpretation unit 126 generates motion vectors v of the upper left corner, upper right corner, and lower left corner of the encoded block. 3 , v 4 and v 5 By projecting this onto the current block, the motion vector v of the upper left corner control point of the current block is obtained. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 Calculate the result.

[0509] Furthermore, as shown in Figure 52A above, if block A is identified and block A has two control points, the MV of three control points may be calculated. Alternatively, as shown in Figure 52B above, if block A is identified and block A has three control points, the MV of two control points may be calculated.

[0510] Next, the interpretation unit 126 performs motion compensation for each of the multiple subblocks included in the current block. That is, the interpretation unit 126 calculates two motion vectors v for each of the multiple subblocks. 0 and v 1 Using the above equation (A), or the three motion vectors v 0 , v 1 and v 2 Using the above formula (B), the MV of the subblock is calculated as the affine MV (step Sk_2).

[0511] Then, the interpretation unit 126 performs motion compensation on the subblock using the affine MV and encoded reference picture (step Sk_3). Once steps Sk_2 and Sk_3 have been executed for each of the subblocks contained in the current block, the process of generating a predicted image using the affine merge mode for the current block is completed. In other words, motion compensation is performed on the current block, and a predicted image for the current block is generated.

[0512] In step Sk_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived for each control point using multiple MV derivation methods. The multiple MV derivation methods may be any combination of the MV derivation methods shown in Figures 50A to 50C, the MV derivation methods shown in Figures 51A and 51B, the MV derivation methods shown in Figures 52A and 52B, and other MV derivation methods.

[0513] The candidate MV list may also include candidate MVs for modes other than affine mode, where prediction is performed on a sub-block basis.

[0514] Furthermore, the candidate MV list may include, for example, a candidate MV for an affine merge mode having two control points and a candidate MV for an affine merge mode having three control points.

[0515] Alternatively, a candidate MV list may be generated containing candidate MVs for an affine merge mode having two control points, and a candidate MV list may be generated containing candidate MVs for an affine merge mode having three control points. Alternatively, a candidate MV list may be generated containing candidate MVs for one of the modes: an affine merge mode having two control points or an affine merge mode having three control points.

[0516] Candidate MVs may be, for example, the MVs of encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or they may be the MVs of any valid block among those blocks.

[0517] Alternatively, you may send an index indicating which candidate MV from the candidate MV list is being selected as MV selection information.

[0518] [MV Derivation > Affine Mode > Affine Intermode] Figure 54 is a flowchart showing an example of an affine intermode.

[0519] In affine intermode, first, the inter prediction unit 126 predicts the MV (v) of each of the two or three control points of the current block. 0 ,v 1 ) or (v 0 ,v 1 ,v 2 Derive the formula (step Sj_1). The control points are the upper left corner, upper right corner, or lower left corner of the current block, as shown in Figure 49A or Figure 49B.

[0520] For example, when using the MV derivation method shown in Figures 51A and 51B, the interpretation unit 126 selects the MV of any of the encoded blocks near each control point of the current block shown in Figure 51A or Figure 51B, thereby predicting the MV (v) of the control point of the current block. 0 ,v 1 ) or (v 0 ,v 1 ,v 2 The following is derived: At this time, the interpretation unit 126 encodes prediction MV selection information into a stream to identify two or three selected prediction MVs.

[0521] For example, the interpretation unit 126 may determine which block's MV from the encoded blocks adjacent to the current block to select as the predicted MV for the control point using cost evaluation or the like, and write a flag indicating which predicted MV was selected to the bitstream. In other words, the interpretation unit 126 outputs the predicted MV selection information, such as a flag, as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0522] Next, the interpretation unit 126 performs motion search (steps Sj_3 and Sj_4) while updating the predicted MV selected or derived in step Sj_1 (step Sj_2).

[0523] In other words, the interpretation unit 126 calculates the MV of each subblock corresponding to the updated predicted MV as an affine MV using the above-described equation (A) or equation (B) (step Sj_3). Then, the interpretation unit 126 performs motion compensation for each subblock using these affine MVs and encoded reference pictures (step Sj_4). The processes in steps Sj_3 and Sj_4 are performed for all blocks in the current block each time the predicted MV is updated in step Sj_2.

[0524] As a result, the interpretation unit 126 determines, for example, the predicted MV that yields the smallest cost in the motion search loop as the MV of the control point (step Sj_5). At this time, the interpretation unit 126 further encodes the difference between the determined MV and the predicted MV as the difference MV into the stream. In other words, the interpretation unit 126 outputs the difference MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0525] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).

[0526] In step Sj_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived for each control point using multiple MV derivation methods. The multiple MV derivation methods may be any combination of the MV derivation methods shown in Figures 50A to 50C, the MV derivation methods shown in Figures 51A and 51B, the MV derivation methods shown in Figures 52A and 52B, and other MV derivation methods.

[0527] The candidate MV list may also include candidate MVs for modes other than affine mode, where prediction is performed on a sub-block basis.

[0528] Furthermore, the candidate MV list may include candidate MVs for affine intermodes having two control points and candidate MVs for affine intermodes having three control points.

[0529] Alternatively, a candidate MV list may be generated containing candidate MVs for affine intermodes having two control points, and a candidate MV list may be generated containing candidate MVs for affine intermodes having three control points. Alternatively, a candidate MV list may be generated containing candidate MVs for one of the modes: an affine intermode with two control points or an affine intermode with three control points.

[0530] Candidate MVs may be, for example, the MVs of encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or they may be the MVs of any valid block among those blocks.

[0531] Additionally, an index indicating which candidate MV from the candidate MV list is being sent as predicted MV selection information.

[0532] [MV Derivation > GPM] In the above example, the interpretation unit 126 generates a prediction image for a rectangular current block. However, motion compensation may be performed for two regions defined by diagonal lines within the current block. These regions are also called partitions. This operating mode is called GPM (Geometric Partitioning Mode). GPM may also be expressed as geometric shape partitioning prediction mode or GPM mode, etc.

[0533] Figure 55A shows examples of two regions where a predicted image is generated by GPM. As shown in Figure 55A, GPM makes it possible to efficiently predict images with oblique object boundaries while maintaining a large CU size.

[0534] Figure 55B is a conceptual diagram showing the pattern of two regions defined in the GPM. The pattern of the two regions (first partition and second partition) defined in the current block in the GPM is indicated by an index that shows the angle of the lines and the distance from the center. There are 64 possible patterns, for example, as shown in Figure 55B. In addition, the shape of the partition may be a triangle, trapezoid, pentagon, or rectangle, etc.

[0535] The interpretation unit 126 specifies one single prediction motion vector each for the first partition and the second partition (the first MV and the second MV in Figure 55A).

[0536] Specifically, candidates are selected by index from a list of merge candidates for dual predictions, similar to the normal merge mode. If the index is even, a motion vector candidate from the L0 list is selected. If the index is odd, a motion vector candidate from the L1 list is selected. If no motion vector candidate to be selected exists in either the L0 or L1 list, a motion vector candidate from the other list is selected.

[0537] Furthermore, when deriving motion vectors for each of the two partitions, the operation is restricted to prevent the same motion vector candidate from being selected. Finally, motion compensation is performed on the first and second partitions using the first and second selected motion vectors, respectively.

[0538] The interpretation unit 126 then generates a predicted image for the current block by weighting near the boundary between the first and second partitions. Specifically, as shown in Figure 55C, it generates a predicted image for the current block by performing a weighted average so that the pixel values ​​of the two partitions are gradually mixed at the boundary.

[0539] Figure 55C is a conceptual diagram showing the weighted average of pixel values ​​at region boundaries. For example, when generating a predicted image of the current block, the pixel labeled "8" in Figure 55C may use only the pixel value based on the first partition, and the pixel labeled "0" may use only the pixel value from the second partition. In that case, for pixels "1" through "7" in Figure 55C, the pixel values ​​from the two partitions are mixed while varying the weight for each pixel position.

[0540] Furthermore, the two motion vectors used in GPM are stored for each subblock, which is determined by dividing the current block into 4x4 subblocks. Specifically, depending on whether the 4x4 subblock belongs to partition 1 only, partition 2 only, or both partition 1 and partition 2 (i.e., near the boundary), one or two motion vectors corresponding to each partition are stored.

[0541] Furthermore, in the case of a subblock located near a boundary, MV1 and MV2 are stored as dual predictive motion vectors. In this case, if both MV1 and MV2 are motion vectors selected from the same list, they are not stored as dual predictive motion vectors, and only one of them is stored as the motion vector for that subblock.

[0542] Figure 56 is a flowchart showing an example of the GPM mode.

[0543] In GPM mode, the interpretation unit 126 first divides the current block into a first partition and a second partition (step Sx_1). At this time, the interpretation unit 126 may encode partition information, which is information regarding the division into each partition, into the stream as prediction parameters. In other words, the interpretation unit 126 may output partition information as prediction parameters to the entropy coding unit 110 via the prediction parameter generation unit 130.

[0544] Next, the interpretation unit 126 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks surrounding the current block in time or space (step Sx_2). In other words, the interpretation unit 126 creates a candidate MV list.

[0545] Then, the interpretation unit 126 selects the candidate MV for the first partition and the candidate MV for the second partition from among the multiple candidate MVs obtained in step Sx_2 as the first MV and the second MV, respectively (step Sx_3).

[0546] In this case, the interpretation unit 126 may encode MV selection information for identifying the selected candidate MV as prediction parameters into the stream. In other words, the interpretation unit 126 may output MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.

[0547] Next, the interpretation unit 126 generates a first predicted image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Similarly, the interpretation unit 126 generates a second predicted image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).

[0548] Finally, the interpretation unit 126 generates a predicted image of the current block by weighting and adding the first predicted image and the second predicted image (step Sx_6).

[0549] [MV Derivation > ATMVP Mode] Figure 57 shows an example of ATMVP mode in which MV is derived on a subblock basis.

[0550] ATMVP mode is a mode classified as a merge mode. For example, in ATMVP mode, candidate MVs at the subblock level are registered in the candidate MV list used in normal merge mode.

[0551] Specifically, in ATMVP mode, first, as shown in Figure 57, the time MV reference block associated with the current block is identified in the encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block. Next, for each subblock within the current block, the MV used when encoding the area corresponding to that subblock within the time MV reference block is identified.

[0552] The identified motion video (MV) is then included in the candidate MV list as a candidate MV for the subblock of the current block. When a candidate MV for a subblock is selected from the candidate MV list, motion compensation using that candidate MV is performed for that subblock. This generates a predicted image for each subblock.

[0553] In the example shown in Figure 57, the block adjacent to the lower left of the current block was used as the peripheral MV reference block, but other blocks may be used. Also, the size of the subblock may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the subblock may be switched in units such as slices, bricks, or pictures.

[0554] [Motion Search > DMVR] Figure 58 shows the relationship between merge mode and DMVR.

[0555] The interpretation unit 126 derives the MV of the current block in merge mode (step Sl_1). Next, the interpretation unit 126 determines whether or not to perform an MV search, i.e., a motion search (step Sl_2). If the interpretation unit 126 determines not to perform a motion search (No. in step Sl_2), it determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). In other words, in this case, the MV of the current block is determined in merge mode.

[0556] On the other hand, if it is determined in step Sl_2 to perform a motion search (Yes in step Sl_2), the interpretation unit 126 derives the final MV for the current block by searching the surrounding region of the reference picture indicated by the MV derived in step Sl_1 (step Sl_3). In other words, in this case, the MV of the current block is determined by the DMVR.

[0557] Figure 59 is a conceptual diagram illustrating an example of a DMVR for determining MV.

[0558] First, for example in merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 list. Similarly, according to the candidate MV (L1), reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 list. A template is generated by taking the average of these reference pixels.

[0559] Next, using the template, the surrounding regions of candidate MVs for the first reference picture (L0) and the second reference picture (L1) are searched, and the MV with the minimum cost is determined as the final MV for the current block. The cost may be calculated, for example, using the difference between each pixel value in the template and each pixel value in the search region, as well as the candidate MV value.

[0560] Any process that can explore the vicinity of candidate MVs and derive the final MV can be used, even if it is not the exact process described here.

[0561] Figure 60 is a conceptual diagram illustrating another example of a DMVR for determining MV. Unlike the example of a DMVR shown in Figure 59, this example in Figure 60 calculates costs without generating a template.

[0562] First, the interpretation unit 126 searches around the reference blocks contained in the reference pictures in the L0 list and L1 list, respectively, based on the initial MV, which is a candidate MV obtained from the candidate MV list. For example, as shown in Figure 60, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1.

[0563] In motion search, the interpretation unit 126 first sets a search position relative to the reference picture in the L0 list. The difference vector indicating the set search position, specifically the difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to that search position, is MVd_L0.

[0564] The interpretation unit 126 then determines the search position in the reference picture of the L1 list. This search position is indicated by a difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. Specifically, the interpretation unit 126 determines the difference vector as MVd_L1 by mirroring MVd_L0. In other words, the interpretation unit 126 sets the search position to a position that is symmetrical to the position indicated by the initial MV in the reference pictures of the L0 list and the L1 list.

[0565] The interpretation unit 126 calculates a cost for each search location, such as the sum of the absolute differences in pixel values ​​within the block at that search location (SAD), and finds the search location that minimizes this cost.

[0566] Figure 61A is a diagram showing an example of motion search in a DMVR, and Figure 61B is a flowchart showing that same motion search example.

[0567] First, in Step 1, the interpretation unit 126 calculates the cost of the search position indicated by the initial MV (also called the starting point) and the eight surrounding search positions. Then, the interpretation unit 126 determines whether the cost of the search positions other than the starting point is the minimum. If the interpretation unit 126 determines that the cost of the search positions other than the starting point is the minimum, it moves to the search position with the minimum cost and proceeds to Step 2. On the other hand, if the cost of the starting point is the minimum, the interpretation unit 126 skips Step 2 and proceeds to Step 3.

[0568] In Step 2, the interpretation unit 126 uses the search position moved according to the processing result of Step 1 as a new starting point and performs a search similar to the processing in Step 1. The interpretation unit 126 then determines whether the cost of the search positions other than the starting point is the minimum. If the interpretation unit 126 determines that the cost of the search positions other than the starting point is the minimum, it proceeds to Step 4. On the other hand, if the interpretation unit 126 determines that the cost of the starting point is the minimum, it proceeds to Step 3.

[0569] In Step 4, the interpretation unit 126 treats the starting point's search position as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0570] In Step 3, the interpretation unit 126 determines the pixel position with the minimum cost based on the costs at the four points above, below, to the left and right of the starting point of Step 1 or Step 2, and sets that pixel position as the final search position. This decimal-precision pixel position is determined by weighting the vector of the four points above, below, to the left and right ((0,1), (0,-1), (-1,0), (1,0)) with the cost at each of the four search positions as the weight. The interpretation unit 126 then determines the difference between the position indicated by the initial MV and the final search position as the difference vector.

[0571] [Motion Compensation > BIO / OBMC / LIC] Motion compensation includes modes that generate a predictive image and then correct that predictive image. These modes include, for example, BIO, OBMC, and LIC, which are described below.

[0572] Figure 62 is a flowchart showing an example of predictive image generation.

[0573] The interpretation unit 126 generates a predicted image (step Sm_1) and corrects the predicted image according to one of the above modes (step Sm_2).

[0574] Figure 63 is a flowchart showing another example of predictive image generation.

[0575] The interpretation unit 126 derives the MV of the current block (step Sn_1). Next, the interpretation unit 126 generates a predicted image using the MV (step Sn_2) and determines whether or not to perform correction processing (step Sn_3). If the interpretation unit 126 determines that correction processing should be performed (Yes in step Sn_3), it generates the final predicted image by correcting the predicted image (step Sn_4).

[0576] In the LIC described later, brightness and color difference may be corrected in step Sn_4. On the other hand, if the interpretation unit 126 determines that no correction processing is necessary (No in step Sn_3), it outputs the predicted image as the final predicted image without correction (step Sn_5).

[0577] [Motion Compensation > OBMC] Interpretation images may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks. Specifically, interpretation images may be generated on a subblock basis within the current block by weighting and adding together a prediction image based on motion information obtained by motion search (within the reference picture) and a prediction image based on motion information of adjacent blocks (within the current picture).

[0578] Such interpretation (motion compensation) is sometimes called OBMC (overlapped block motion compensation) or OBMC mode.

[0579] In OBMC mode, information indicating the size of the subblock for OBMC (e.g., called the OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether or not to apply OBMC mode (e.g., called the OBMC flag) may be signaled at the CU level. Note that the signaling levels of this information are not limited to the sequence level and CU level, but may be other levels (e.g., picture level, slice level, brick level, CTU level, or subblock level).

[0580] The OBMC mode will be explained in more detail. Figures 64 and 65 are flowcharts and conceptual diagrams illustrating the overview of the predictive image correction process using OBMC.

[0581] First, as shown in Figure 65, a predicted image (Pred) is obtained using normal motion compensation with the MV assigned to the current block. In Figure 65, the arrow "MV" points to the reference picture, indicating what the current block of the current picture is referencing in order to obtain the predicted image.

[0582] Next, the previously derived MV (MV_L) for the encoded left-adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_L). The MV (MV_L) is indicated by an arrow “MV_L” pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by superimposing the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.

[0583] Similarly, the previously derived MV (MV_U) for the encoded upper adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_U). The MV (MV_U) is indicated by an arrow “MV_U” pointing from the current block to the reference picture. Then, the predicted image Pred_U is superimposed onto the predicted image that has undergone the first correction (for example, Pred and Pred_L) to perform a second correction of the predicted image.

[0584] This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block, with the boundaries with adjacent blocks blended (smoothed).

[0585] The above example is a two-pass correction method using left-adjacent and top-adjacent blocks, but the correction method may also be a three-pass or more-pass correction method using right-adjacent and / or bottom-adjacent blocks.

[0586] Furthermore, the area to be superimposed does not have to be the entire pixel area of ​​the block, but rather only a portion of the area near the block boundary.

[0587] This section describes OBMC's predictive image correction process for obtaining a single predictive image Pred by superimposing additional predictive images Pred_L and Pred_U onto a single reference picture.

[0588] However, if the predicted image is corrected based on multiple reference images, the same process may be applied to each of the multiple reference pictures. In such cases, by performing OBMC image correction based on multiple reference pictures, a corrected predicted image is obtained from each reference picture, and then the multiple corrected predicted images obtained are further superimposed to obtain the final predicted image.

[0589] In OBMC, the unit of a current block may be a PU unit, or it may be a subblock unit obtained by further dividing a PU.

[0590] One method for determining whether or not to apply OBMC is to use the obmc_flag signal, which indicates whether or not OBMC should be applied.

[0591] As a specific example, the encoding device 100 may determine whether the current block belongs to a region with complex motion. When the encoding device 100 determines that the current block belongs to a region with complex motion, it sets the value 1 as obmc_flag and performs encoding by applying OBMC. When the current block does not belong to a region with complex motion, it sets the value 0 as obmc_flag and performs block encoding without applying OBMC.

[0592] On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC according to the value.

[0593] [Motion Compensation > BIO] Next, a method for deriving the MV will be described. First, a mode for deriving the MV based on a model assuming a constant velocity linear motion will be described. This mode is sometimes called the BIO (bi-directional optical flow) mode. Also, this bi-directional optical flow may be denoted as BDOF instead of BIO.

[0594] FIG. 66 is a diagram for explaining a model assuming a constant velocity linear motion. In FIG. 66, (v x , v y ) indicates the velocity vector, and τ 0 , τ 1 respectively indicate the temporal distances between the current picture (Cur Pic) and the two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) indicates the MV corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) indicates the MV corresponding to the reference picture Ref 1 .

[0595] At this time, under the assumption of a constant velocity linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1, MVy 1 ) are, respectively, (v x τ 0 ,v y τ 0 ) and (-v x τ 1 , -v y τ 1 This can be expressed as follows, and the following optical flow equation holds:

[0596]

[0597] Here, I(k) represents the luminance value of the reference image k (k=0,1) after motion compensation. This optical flow equation shows that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on this optical flow equation and Hermitian interpolation, block-level motion vectors obtained from a candidate MV list, etc., may be corrected on a pixel-by-pixel basis.

[0598] Furthermore, the motion vector (MV) may be derived by the decoding device 200 using a method different from that used to derive the motion vector based on a model assuming uniform linear motion. For example, the motion vector may be derived on a sub-block basis based on the MVs of multiple adjacent blocks.

[0599] Figure 67 is a flowchart illustrating an example of inter prediction according to BIO. Figure 68 is a diagram illustrating an example of the configuration of the inter prediction unit 126 that performs the inter prediction according to BIO.

[0600] As shown in Figure 68, the interpretation unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a prediction image correction unit 126f. Note that the memory 126a may be a frame memory 122.

[0601] The interpretation unit 126 uses two different reference pictures (Ref) from the picture (Cur Pic) that contains the current block. 0 ,Ref1 Using this, two motion vectors (M0, M1) are derived. Then, the interpretation unit 126 uses these two motion vectors (M0, M1) to derive a predicted image of the current block (step Sy_1).

[0602] Note that the motion vector M0 is the reference picture Ref 0 The corresponding motion vector (MVx 0 , MVy 0 ) and the motion vector M1 is the reference picture Ref 1 The corresponding motion vector (MVx 1 , MVy 1 )

[0603] Next, the interpolated image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to create an interpolated image I of the current block. 0 The interpolated image deriving unit 126b also references memory 126a and uses motion vector M1 and reference picture L1 to derive the interpolated image I of the current block. 1 Derive the following (step Sy_2).

[0604] Here, interpolated image I 0 This is derived for the current block, the reference picture Ref 0 The image included is interpolated image I 1 This is derived for the current block, the reference picture Ref 1 This is an image included in [the collection / package].

[0605] Interpolated image I 0 and interpolated image I 1 Each of these may be the same size as the current block. Alternatively, interpolated image I 0 and interpolated image I 1 Each of these may be a larger image than the current block in order to properly derive the gradient image described later. Furthermore, interpolated image I 0 and I 1 This may include a motion vector (M0, M1) and a reference picture (L0, L1), and a predicted image derived by applying a motion compensation filter.

[0606] Furthermore, the gradient image derivation unit 126c generates an interpolated image I 0 and interpolated image I 1 From there, the gradient image of the current block (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) is derived (step Sy_3). Note that the horizontal gradient image is (Ix 0 , Ix 1 ) and the vertical gradient image is (Iy 0 , Iy 1 The gradient image derivation unit 126c may derive the gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image only needs to show the amount of spatial change in pixel values ​​along the horizontal or vertical direction.

[0607] Next, the optical flow derivation unit 126d interpolates the image (I) in units of multiple subblocks that constitute the current block. 0 , I 1 ) and gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 Using the above velocity vector, optical flow (v x ,v y Derive the following (step Sy_4).

[0608] Optical flow is a coefficient that corrects the spatial displacement of pixels, and may also be called a local motion estimate, corrected motion vector, or corrected weight vector. For example, a subblock may be a 4x4 pixel subCU. Note that the derivation of optical flow may be performed at other units, such as pixel units, rather than subblock units.

[0609] Next, the interpretation unit 126 calculates the optical flow (v x ,v y The predicted image of the current block is corrected using ). For example, the correction value derivation unit 126e uses optical flow (v x ,v yThe correction value for the pixel values ​​included in the current block is derived using (step Sy_5). The predicted image correction unit 126f may then correct the predicted image of the current block using the correction value (step Sy_6). The correction value may be derived for each pixel, or for multiple pixels or subblocks.

[0610] Note that the BIO processing flow is not limited to the processing disclosed in Figure 67. Only a portion of the processing disclosed in Figure 67 may be performed, different processing may be added or replaced, or the processing may be executed in a different order.

[0611] [Motion Compensation > LIC] Next, we will explain an example of a mode that generates a predicted image (prediction) using LIC (local illumination compensation).

[0612] Figure 69A is a diagram illustrating an example of a predictive image generation method using brightness correction processing by an LIC (Luminous Complex). Figure 69B is a flowchart illustrating an example of the predictive image generation method using the LIC.

[0613] First, the interpretation unit 126 derives MV from the encoded reference picture and obtains the reference image corresponding to the current block (step Sz_1).

[0614] Next, the interpretation unit 126 extracts information for the current block indicating how the luminance values ​​have changed between the reference picture and the current picture (step Sz_2). This extraction is performed based on the luminance pixel values ​​of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values ​​at the equivalent positions in the reference picture specified by the derived MV.

[0615] Then, the interpretation unit 126 calculates a brightness correction parameter using information indicating how the brightness value has changed (step Sz_3).

[0616] The interpretation unit 126 generates a predicted image for the current block by performing a brightness correction process that applies its brightness correction parameter to the reference image in the reference picture specified by MV (step Sz_4).

[0617] In other words, the predicted image, which is a reference image within the reference picture specified in MV, is corrected based on luminance correction parameters. This correction may involve correcting luminance or color difference. Specifically, color difference correction parameters may be calculated using information indicating how the color difference has changed, and the color difference correction process may be performed.

[0618] Note that the shape of the peripheral reference region in Figure 69A is just one example, and other shapes may be used.

[0619] Furthermore, although this explanation describes the process of generating a predicted image from a single reference picture, the process is similar when generating predicted images from multiple reference pictures. Alternatively, the brightness correction process may be applied to each reference picture obtained from the reference picture in the same manner as described above before generating the predicted image.

[0620] One method for determining whether to apply LIC is to use a signal called lic_flag, which indicates whether or not to apply LIC. As a specific example, the encoding device 100 determines whether the current block belongs to a region where a brightness change is occurring. If it belongs to a region where a brightness change is occurring, it sets lic_flag to a value of 1 and applies LIC to perform encoding. If it does not belong to a region where a brightness change is occurring, it sets lic_flag to a value of 0 and performs encoding without applying LIC.

[0621] On the other hand, the decoding device 200 may decode the lic_flag described in the stream and then switch whether or not to apply LIC depending on its value before performing the decoding.

[0622] Another way to determine whether to apply LIC is, for example, by checking whether LIC has been applied in surrounding blocks.

[0623] As a specific example, when the current block is being processed in merge mode, the interpretation unit 126 determines whether the surrounding encoded blocks selected during the MV derivation in merge mode were encoded with LIC applied. The interpretation unit 126 then switches whether to apply LIC and performs encoding based on the result. In this example as well, the same process is applied to the decoding device 200.

[0624] The LIC (Luminance Correction Processing) was explained using Figures 69A and 69B, but its details will be explained below.

[0625] First, the interpretation unit 126 derives an MV for obtaining the reference image corresponding to the current block from the reference picture, which is an encoded picture.

[0626] Next, the interpretation unit 126 extracts information indicating how the luminance values ​​have changed between the reference picture and the current picture, using the luminance pixel values ​​of the left-adjacent and upper-adjacent encoded peripheral reference regions and the luminance pixel values ​​at equivalent positions in the reference picture specified by MV, and calculates a luminance correction parameter.

[0627] For example, let p0 be the luminance pixel value of a pixel in the peripheral reference area of ​​the current picture, and let p1 be the luminance pixel value of a pixel in the peripheral reference area of ​​the reference picture at the same position as that pixel. The interpretation unit 126 calculates coefficients A and B as luminance correction parameters to optimize A × p1 + B = p0 for multiple pixels in the peripheral reference area.

[0628] Next, the interpretation unit 126 generates a predicted image for the current block by performing a brightness correction process on the reference image in the reference picture specified by MV using a brightness correction parameter. For example, let p2 be the brightness pixel value in the reference image, and p3 be the brightness pixel value of the predicted image after brightness correction processing. The interpretation unit 126 generates the predicted image after brightness correction processing by calculating A × p2 + B = p3 for each pixel in the reference image.

[0629] Furthermore, a portion of the peripheral reference region shown in Figure 69A may be used. For example, a region containing a predetermined number of pixels obtained by thinning out the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference region. Also, the peripheral reference region is not limited to the region adjacent to the current block, but may also be a region not adjacent to the current block.

[0630] Furthermore, in the example shown in Figure 69A, the peripheral reference area within the reference picture is the area specified by the MV of the current picture, relative to the peripheral reference area within the current picture, but it may also be an area specified by another MV. For example, this other MV may be the MV of the peripheral reference area within the current picture.

[0631] Although the operation of the encoding device 100 has been described here, the operation of the decoding device 200 is similar.

[0632] Furthermore, LIC may be applied not only to luminance but also to color difference. In this case, correction parameters may be derived individually for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.

[0633] Furthermore, LIC processing may be applied on a subblock basis. For example, correction parameters may be derived using the surrounding reference region of the current subblock and the surrounding reference region of the reference subblock within the reference picture specified by the MV of the current subblock.

[0634] [Generated Pictures] Generated pictures are images obtained through a generation process, unlike regular reference images (in other words, images decoded by intra / inter prediction). Here, as an example, we show how to generate a new reference picture by inputting one or more pictures into a neural network, but the method of generating a new reference picture is not limited to this.

[0635] The generation process may support the application of NN loop filter (NN LF) processing, which includes NN (Neural Network) processing, or it may support transformation processing (scaling, rotation, mirroring, shifting, etc.). Furthermore, the generation process may be a combination of multiple processes (for example, a process that generates a predicted image using an NN such as NN LF, a transformation process, and other processes).

[0636] The new reference picture generated by the generation process is also referred to as a generated reference picture, a generated picture, or a processed picture. Furthermore, such a new reference picture or its image may also be referred to as a generated reference image, a generated image, or a processed image. Adding the generated picture to the reference picture list and making it accessible may improve encoding efficiency.

[0637] First, we will explain the generated reference picture that is generated using the NN. The encoding device 100 and the decoding device 200 input one or more images to the NN and generate a generated picture.

[0638] Figure 70A is a conceptual diagram illustrating an example of a generated picture. As illustrated in Figure 70A, a new image (PicA') generated by inputting an image (PicA) into the NN may be used as a reference picture. For example, a reference picture generated by NN LF is an example of this.

[0639] Figure 70B is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 70B, a new image (PicAB') generated by inputting two images (PicA and PicB) into the NN may be used as the reference picture. Note that there are not limited to two images input into the NN; more than two images can be input.

[0640] Figure 70C is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 70C, if NN processing such as NN LF is not fast enough, a new image (PicA') smaller in size than the input image (PicA) may be used as a reference picture. In other words, the size of the input image and the size of the generated image may be different. Also, information about the size of the generated image may be included in the bitstream.

[0641] While this example shows generation at the picture level, the generation unit is not limited to this. The generated reference picture may be generated in predetermined units such as blocks, slices, or tiles contained within the picture.

[0642] Furthermore, the NN may be implemented using an NPU (Neural Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit).

[0643] Furthermore, as mentioned above, the newly generated image (PicA' or PicAB') may be used as a generated picture for inter prediction. Inter prediction using this generated picture may be called neural network intercoding (NN inter).

[0644] Furthermore, multiple neural networks (NNs) may be combined in the generation of the generated picture. In other words, for example, PicA and PicB may be input to NN1 to generate PicAB', and PicAB' may be input to another NN2 to generate PicAB''.

[0645] In typical image encoding on a CPU, a conventional pipeline is used, as shown in Figure 71A. Specifically, the entire encoding process is divided into several stages (0 to n), and processing is performed in block units within each stage. Furthermore, multiple stages are processed in parallel with each other.

[0646] In contrast, for generated pictures produced using NN processing, as can be seen from the configuration examples in Figures 71B and 71C, the image is first processed by the CPU, and then the data is passed to the NPU / GPU / TPU, etc. Then, NN processing is executed on the NPU / GPU / TPU side using a different pipeline than the CPU. In addition, processing may be executed on the NPU / GPU / TPU in units different from those of the CPU.

[0647] Due to overhead caused by data transfer and other processes, it may take some time for the generated picture produced by the NN processing to become available as a reference image for interpretation.

[0648] Patches may be defined to generate the generated picture by NN processing. A patch refers to a predetermined pixel region containing a predetermined number of pixels, as shown in Figure 72A. The shape of the patch may be a square, a non-square rectangle, or any other shape, as shown in Figure 72A. For example, a patch may have a non-rectangular shape such as a triangle or a rhombus.

[0649] Figure 72B is a conceptual diagram showing a patch containing an extended region. A patch containing an extended region includes an extended region in the vertical and / or horizontal directions relative to the overall size of the patch.

[0650] In the example shown in Figure 72B, the extension area has equal size in the vertical and horizontal directions, resulting in a square patch. The sizes of the vertical and horizontal extension areas are not limited to the example in Figure 72B and may be different. Also, the extension areas do not have to be equal, for example, at the top and bottom, or left and right. Furthermore, information regarding the shape or size of the patch may be added to the header area of ​​the bitstream, or it may be added to the bitstream as metadata (also called meta information).

[0651] Figure 72C is a conceptual diagram showing an image containing patches 1A and 1B that have overlapping regions. Patches 1A and 1B overlap each other across multiple pixels (=set of samples), as shown by the dotted lines in Figure 72C. Information indicating the shape (width and height) or size (number of pixels included in the overlapping region) of this overlapping region may be stored in the bitstream.

[0652] Furthermore, in the example in Figure 72C, each patch contains a 256x256 block. Therefore, each patch can also be described as a patch with an extended area relative to the 256x256 block in Figure 72B. In the example in Figure 72C, the 256x256 block in patch 1A and the 256x256 block in patch 1B are defined so as not to overlap.

[0653] Furthermore, while Figure 72C shows an example of two patches defined within a single picture, in an NN interface, for example, two patches may be defined in each of two pictures. In that case, the positions of the patches defined in the first picture and the second picture (for example, patch 1A in the first picture and patch 2A in the second picture, which is not shown) may be the same or different.

[0654] Figure 72D is a conceptual diagram showing the patches defined in each picture. Specifically, it shows multiple patches defined in an example such as an NN interface, including the first and third patches in the first picture, the second and fourth patches in the second picture, and the fifth and sixth patches in the third picture. The juxtaposed points of the third, fourth, and sixth patches are indicated by an "x" mark.

[0655] Figure 72E is a conceptual diagram showing an example of a method for generating the sixth patch. The sixth patch is obtained by cropping the seventh patch, which is generated by inputting the third and fourth patches into an NN generator. Similarly, the fifth patch may be obtained by cropping an image (the eighth patch, not shown) generated by inputting the first and second patches into an NN generator.

[0656] Furthermore, the first and second pictures may be ordinary reference pictures. In contrast, the third picture, which includes the fifth and sixth patches, may be a generated picture. In one embodiment, the encoding order and display order may be first picture → second picture → third picture. In another embodiment, the display order may be first picture → third picture → second picture.

[0657] Next, the generated picture produced using the transformation process will be described. The encoding device 100 and the decoding device 200 generate a generated picture by performing a transformation process on a single image. The single image input to the transformation process may be a normal reference image (in other words, an image decoded by intra / inter prediction) or an image generated by an NN. Furthermore, the image generated by the transformation process may be used as input to the NN.

[0658] Figure 73A is a conceptual diagram illustrating another example of a generated picture. As illustrated in Figure 73A, a new image (PicA') generated by scaling (enlarging or reducing) an image (PicA) may be used as the reference picture.

[0659] In the case of upscaling (enlarging), a new image PicA' is generated by enlarging the image and then cropping the resulting enlarged image. In the case of downscaling (reducing), a new image PicA' is generated by reducing the image and then padding the resulting reduced image. Note that enlargement can also be expressed as expansion.

[0660] Figure 73B is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73B, a new image (PicA') generated by rotating an image (PicA) may be used as a reference picture. For example, a new image PicA' is generated by rotating the image by an angle θ in a specified direction, and then cropping and padding the rotated image. The rotation direction may always be clockwise or always counterclockwise, and the rotation direction may also be specified as a parameter included in the bitstream along with the rotation angle.

[0661] Figure 73C is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73C, a new image (PicA') generated by mirroring an image (PicA) may be used as the reference picture. In Figure 73C, the left-hand sample of the input image (PicA) is mirrored horizontally.

[0662] Mirroring is not limited to the example above; the sample on the right may be mirrored to the left, or vertical mirroring may be performed (mirroring the upper sample to the lower side, or mirroring the lower sample to the upper side). In this example, when mirroring, half of the width x of the input image (PicA) (or half of the height y in the case of vertical mirroring) is always mirrored, and the width / height of the mirroring is uniformly defined, but the width / height of the mirroring may also be specified by parameters.

[0663] Figure 73D is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73D, a new image (PicA') generated by shifting an image (PicA) may be used as the reference picture. In Figure 73D, the input image (PicA) is shifted to the right. In other words, the left side of the input image is padded and the right side is cropped. The shift is not limited to this example; the right side of the sample may be shifted to the left, or a vertical shift (shifted upwards or downwards) may be performed.

[0664] In the above, shifting the image to the right corresponds to shifting the image sample, i.e., the image content, to the right, and to shifting the processing range of the image to the left. Similarly, shifting the image to the left, up, and down corresponds to shifting the image sample, i.e., the image content, to the left, up, and down, respectively, and to shifting the processing range of the image to the right, down, and up, respectively.

[0665] Furthermore, multiple transformations may be combined. For example, an image (PicA) may be scaled to produce an image (Scaled PicA), and then a shifted image (Shifted PicA) may be generated.

[0666] [NN Loop Filter] The NN Loop Filter (NN LF: Neural Network Loop Filter) corresponds to a filtering process performed on an input image (such as a PicA) using a neural network (NN). The image generated by the NN LF may be used as a reference image for interpretation, or it may be used as the display image as is.

[0667] [NN Inter Prediction] NN Inter Prediction (Neural Network Inter Prediction) corresponds to inter prediction performed using a neural network (NN). Specifically, in NN inter prediction, one or more images (PicA and / or PicB, etc.) are used as input to the NN, and an image is generated as a generated image (PicAB', etc.). The generated image is then used as a reference image for inter prediction. Since the image is generated using an NN, the aforementioned NN processing may also be used.

[0668] The generated reference image used in NN interpretation is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).

[0669] Note that while this example shows the use of newly generated images using a neural network as an example of NN interpretation, the NN interpretation process is not limited to this example. For example, images to which NN LF processing has been applied may also be used.

[0670] [Translation Inter Prediction] Translation Inter Prediction corresponds to an inter prediction performed using a transformation process. Specifically, in translation inter prediction, a transformation process is applied to one or more images (PicA and / or PicB, etc.) to generate an image (PicAB', etc.). This generated image is then used as the reference image for inter prediction. The transformation process described above may also be used as the transformation process.

[0671] The generated reference image referenced in the conversion interface prediction is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).

[0672] (First Embodiment) [Decoding Process] Figure 74 is a flowchart of the decoding process according to the first embodiment. First, the decoding device decodes one or more parameters and the first image from the bitstream (S101). For example, one or more parameters may be included in the bitstream header. One or more parameters may be included in, for example, VPS, SPS, PPS, SEI, SH (slice header), or VUI (Video usability Information). One or more parameters may include an ID (identifier) ​​for identifying the neural network used for NN processing (NN inter, or NN LF), and may also include information about the characteristics of the neural network used. When an NN inter is used, the first image may include one or more images that have been decoded before the second image.

[0673] Next, the decoding device generates a processed image by NN processing using the first image and the neural network (S102). Here, one or more parameters include information about the time during which the processed image can be accessed. One or more parameters may also include information about the input and output of the neural network. The neural network may include one or more convolutional layers or recurrent layers.

[0674] Furthermore, the decoding device decodes the second image from the bitstream (S103). Here, the processed image is not referenced during the decoding of the second image. In other words, the decoding device decodes the second image after decoding the first image and before the processed image becomes accessible. For example, steps S102 and S103 may be processed simultaneously or in parallel.

[0675] Next, the decoding device decodes the third image from the bitstream (S104). Here, the processed image is referenced during the decoding of the third image. That is, the decoding device decodes the third image after decoding the first image and after the processed image has become accessible. For example, the decoding device generates multiple prediction samples for interpretation using multiple samples (pixel values) of the processed image. Note that in the decoding of the third image, in addition to the processed image, the first and second images before the NN processing is applied may also be referenced.

[0676] For example, the decoding device may delete the processed image from the encoded picture buffer after decoding the third image. Alternatively, for example, the decoding device may generate a processed image using the first image and a neural network through NN processing, and then delete the first image from the encoded picture buffer. The images mentioned in steps S101, S102, S103, and S104 (the first image, the second image, the third image, and the processed image) may be slices, tiles, pictures, or subpictures. In another example, this image may refer to a part of an image represented by size information. In this case, the size information is signaled to the bitstream.

[0677] [Encoding Process] Figure 75 is a flowchart of the encoding process according to the first embodiment. First, the encoding device encodes one or more parameters and the first image into a bitstream (S201). One or more parameters may be included in the header of the bitstream. One or more parameters may be included, for example, in VPS, SPS, PPS, SEI, SH, or VUI. One or more parameters may include information about the characteristics of the neural network used, including an ID (identifier) ​​for identifying the neural network used for NN processing (NN inter, or NN LF). When an NN inter is used, the first image may include one or more images encoded before the second image.

[0678] Next, the encoding device generates a processed image by NN processing using the first image and the neural network (S202). Here, one or more parameters include information about the time during which the processed image is accessible. One or more parameters may also include information about the input and output of the neural network. The neural network may include one or more convolutional layers or recurrent layers.

[0679] Furthermore, the encoding device encodes the second image into a bitstream (S203). Here, the processed image is not referenced during the encoding process of the second image. In other words, the encoding device encodes the second image after encoding the first image and before the processed image becomes accessible. For example, steps S202 and S203 may be processed simultaneously or in parallel.

[0680] Next, the encoding device encodes the third image into a bitstream (S204). Here, the processed image is referenced during the encoding process of the third image. That is, the encoding device encodes the third image after the first image has been encoded and the processed image has become accessible. For example, the encoding device generates multiple prediction samples for interpretation using multiple samples (pixel values) of the processed image. Note that in the encoding process of the third image, in addition to the processed image, the first and second images before the NN processing is applied may also be referenced.

[0681] For example, the encoding device may delete the processed image from the encoded picture buffer after encoding the third image. Alternatively, for example, the encoding device may generate a processed image using the first image and a neural network through NN processing, and then delete the first image from the encoded picture buffer. The images referred to in steps S201, S202, S203, and S204 (first image, second image, third image, and processed image) may be slices, tiles, images, or sub-images. In another example, this image may refer to a part of an image represented by size information. In this case, the size information is signaled to the bitstream.

[0682] [NN LF] For example, NN LF refers to the process of filtering an image using a neural network for reference and display of interpretation. In the following example, one output image (processed image) is generated using NN LF with one image.

[0683] Figure 76 shows an example of the display order and encoding order of multiple images. Here, POC (Picture Order Count) is used as the display order. The encoding order refers to the order of encoding in the encoding device and the order of decoding in the decoding device. In other words, the encoding order is the same as the decoding order, and the encoding order can also be rephrased as the decoding order.

[0684] Furthermore, while the following explanation primarily uses the processing in the decoding device as an example, the processing in the encoding device is similar. In other words, the processing in the encoding device is the same as "encoding" in the following explanation, just with "decoding" replaced by "encoding".

[0685] In this example, the images POC0, POC2, and POC4 are reference images (reference pictures) that are referenced by other images, while the images POC1 and POC3 are non-reference images (non-reference pictures) that are not referenced by other images.

[0686] Figure 77 shows an example of an image stored in a buffer (encoded picture buffer). Figure 77 shows the target image to be decoded and the state of the buffer during periods T0 to T4. Figures 78 to 81 show examples of processing during periods T1 to T4, respectively. Note that each of periods T0 to T4 corresponds to, for example, one frame period (for example, 1 / 60 second).

[0687] At T0, the buffer is empty, and the decoding device begins decoding the image of POC0. This decoding process is performed, for example, by the CPU included in the decoding device.

[0688] At T1, the decoding of the POC0 image is complete, and the decoded POC0 image (decoded image) is stored in the buffer. The decoding device also starts NN processing (NN LF) on the POC0 image. This NN processing is performed, for example, by the GPU included in the decoding device. This NN processing generates the processed image of POC0. In other words, at T1, the processed image during NN processing (processed image before it can be referenced) is stored in the buffer before it is allowed to be referenced by other images. The decoding device also starts decoding the POC4 image. In the decoding process of the POC4 image, the POC0 image (decoded image) before NN processing (unprocessed) is referenced.

[0689] In T2, the decoding of the POC4 image is complete, and the decoded POC4 image is stored in the buffer. Also, the NN processing of the POC0 image is complete, and the processed image of POC0 (the referenced processed image), which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts NN processing on the POC4 image. The decoding device also starts decoding the POC2 image. In the decoding process of the POC2 image, the processed image of POC0 (which has already undergone NN processing) and the image of POC4 (which has not undergone NN processing) are referenced.

[0690] In T3, the decoding of the POC2 image is complete, and the decoded POC2 image is stored in the buffer. Also, the NN processing of the POC4 image is complete, and the processed POC4 image, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts NN processing on the POC2 image. The decoding device also starts decoding the POC1 image. In this decoding process of the POC1 image, the processed image of POC0 (which has already undergone NN processing), the processed image of POC4 (which has also undergone NN processing), and the image of POC2 before NN processing are referenced.

[0691] In T4, the decoding of the POC1 image is complete. Note that since the POC1 image is a non-referenced image, the decoded image of POC1 is not stored in the buffer. Also, the NN processing of the POC2 image is complete, and the processed image of POC2, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts decoding the POC3 image. In the decoding process of the POC3 image, the processed images of NN-processed POC0, NN-processed POC2, and NN-processed POC4 are referenced.

[0692] Thus, because processing time is required for NN processing by the GPU, the processed images cannot be immediately accessed during the subsequent image decoding process. Therefore, the encoding device signals information indicating the appropriate processing time to the bitstream. This allows the decoding device to use this information to know when each processed image will become accessible. As a result, it becomes possible to decode the processed images in real time on the GPU, improving the results of the filtering process.

[0693] The above example shows that the NN processing of one image is completed within one frame period (for example, 1 / 60th of a second). However, the same processing can be applied even when N frame periods (where N is an integer greater than or equal to 2) are required to complete the NN processing of one image.

[0694] If the NN processing of one image is completed within one frame period (e.g., 1 / 60 second), then, as in the example above, the processed image of a certain image (e.g., POC0) becomes accessible in the period following the period in which the decoding of that image was completed (e.g., T1) (e.g., T2).

[0695] On the other hand, if it takes N frames for the NN processing of a single image to be completed, the processed image of that image becomes accessible in a period N frames after the period in which the decoding of that image is completed. For example, if it takes 2 frames for the NN processing of a single image to be completed, the processed image of a certain image (e.g., POC0) becomes accessible in a period (e.g., T3) two frames after the period (e.g., T1) in which the decoding of that image (e.g., POC0) is completed.

[0696] Figures 82 and 83 show examples of the first, second, and third images. For example, the image of POC0 corresponds to one or more first images, the image of POC4 corresponds to the second image, the image of POC2 corresponds to the third image, and the processed image of POC0 corresponds to the first image after NN processing.

[0697] The following describes another example of operation when NN LF is used. Figure 84 shows an example of the display order and encoding order of multiple images. In this example, all images POC0 to POC4 are reference images that are referenced by other images.

[0698] Figure 85 shows an example of an image stored in a buffer (encoded picture buffer). Figure 85 shows the target image to be decoded and the state of the buffer during the period T0 to T4. Figures 86 to 89 show examples of processing during T1 to T4, respectively.

[0699] At T0, the buffer is empty, and the decoding unit begins decoding the image of POC0. At T1, the decoding of the image of POC0 is complete, and the decoded image of POC0 is stored in the buffer. The decoding unit also begins NN processing (NN LF) on the image of POC0. The decoding unit also begins decoding POC1. In the decoding process of the image of POC1, the POC0 image before NN processing is referenced.

[0700] In T2, the decoding of the POC1 image is complete, and the decoded POC1 image is stored in the buffer. Also, the NN processing of the POC0 image is complete, and the processed POC0 image, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts NN processing on POC1. The decoding device also starts decoding the POC2 image. In this decoding process for POC2, the processed image of POC0 (which has already undergone NN processing) and the image of POC1 (before NN processing) are referenced.

[0701] In T3, the decoding of the POC2 image is complete, and the decoded POC2 image is stored in the buffer. Also, the NN processing of the POC1 image is complete, and the processed POC1 image, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts NN processing on the POC2 image. The decoding device also starts decoding the POC3 image. In this POC3 decoding process, the NN-processed POC0 image, the NN-processed POC1 image, and the unprocessed POC2 image are referenced.

[0702] In T4, the decoding of the POC3 image is complete, and the decoded POC3 image is stored in the buffer. Also, the NN processing of the POC2 image is complete, and the processed POC2 image, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts NN processing on the POC3 image. The decoding device also starts decoding the POC4 image. In this POC4 decoding process, the processed images of NN-processed POC0, NN-processed POC1, NN-processed POC2, and the image of POC3 before NN processing are referenced.

[0703] Thus, because processing time is required for NN processing by the GPU, the processed images cannot be immediately accessed during the subsequent image decoding process. Therefore, the encoding device signals information indicating the appropriate processing time to the bitstream. This allows the decoding device to use this information to know when each processed image will become accessible. As a result, it becomes possible to decode the processed images in real time on the GPU, improving the results of the filtering process.

[0704] Figures 90 and 91 show examples of the first, second, and third images. For example, the image of POC0 corresponds to one or more first images, the image of POC1 corresponds to the second image, the image of POC2 corresponds to the third image, and the processed image of POC0 corresponds to the first image after NN processing.

[0705] [NN Inter] For example, NN Inter is an inter prediction process that uses a neural network to generate one or more reference images. In the following example, two images are used to generate one output image (processed image) using NN Inter.

[0706] Figure 92 shows an example of the display order and encoding order of multiple images. In this example, images POC0, POC1, POC2, and POC4 are reference images that are referenced by other images, while image POC3 is a non-reference image that is not referenced by other images.

[0707] Figure 93 shows an example of an image stored in a buffer (encoded picture buffer). Figure 93 shows the target image to be decoded and the state of the buffer during the period T0 to T4. Figures 94 to 97 show examples of processing during T1 to T4, respectively.

[0708] At T0, the buffer is empty, and the decoding unit begins decoding the image of POC0. At T1, the decoding of the image of POC0 is complete, and the decoded image of POC0 (decoded image) is stored in the buffer for reference. The decoding unit also begins decoding the image of POC4.

[0709] At T2, the decoding of the POC4 image is complete, and the decoded POC4 image is stored in the buffer. The decoding device also starts decoding the POC2 image. In the decoding process of the POC2 image, if interpretation is used, the POC0 image and the POC4 image are referenced. The POC0 image and the POC4 image are sent to the GPU. The GPU starts generating a new reference image, the processed image (POC1p), through NN interpretation. In other words, at T2, the processed image of POC1p during NN processing (the processed image before it can be referenced) is stored in the buffer before it is allowed to be referenced by other images. The POC1p image corresponds to the same time as the POC1 image.

[0710] At T3, the decoding of POC2 is complete, and the decoded image of POC2 is stored in the buffer. The NN processing (NN interprocessing) of the image of POC1p is complete, and the processed image of POC1p (a referenceable processed image) that is permitted to be referenced by other images is stored in the buffer. The decoding device also starts decoding the image of POC1. Interpretation is used in the decoding process of the image of POC1, and the images of POC0, POC2, POC4, and the processed image of POC1p are referenced. The decoding device also starts generating a new reference image, a processed image (POC3p), through NN interprocessing.

[0711] At T4, the decoding of POC1 is complete, and the decoded image of POC1 is stored in the buffer. The decoding device removes (deletes) the processed image of POC1p from the buffer. Also, the NN processing of the image of POC3p is complete, and the processed image of POC3p, which is permitted to be referenced by other images, is stored in the buffer. The decoding device also starts decoding the image of POC3. Interpretation is used in the decoding process of the image of POC3, and the images of POC0, POC1, POC2, POC4, and the processed image of POC3p are referenced.

[0712] Thus, the processed images for POC1p and POC3p are generated before decoding the images for POC1 and POC3, respectively. Therefore, the processed images for POC1p and POC3p can be used in interpretation for the images for POC1 and POC3, respectively. This increases the number of reference images available for each image. It also makes the reference image closest to the processed image available. These improvements can be made to the results of the interpretation process.

[0713] The above example shows that the NN processing of one image is completed within one frame period (for example, 1 / 60th of a second). However, the same processing can be applied even when N frame periods (where N is an integer greater than or equal to 2) are required to complete the NN processing of one image.

[0714] If the NN processing of one image is completed within one frame period (e.g., 1 / 60 second), then, as in the example above, the image referenced by the next decoded picture (e.g., POC1p) is generated in advance during the period (e.g., T2) when the decoding of a certain image (e.g., POC0 and POC4) is completed, so that the processed image of that image (e.g., POC1p) becomes accessible in the next period (e.g., T3).

[0715] On the other hand, if it takes N frames for the NN processing of one image to be completed, the processed image of a certain image becomes accessible in a period N frames after the period in which the decoding of that image is completed. For example, if it takes 2 frames for the NN processing of one image to be completed, the image that will be referenced in the next two decoded pictures (e.g., POC3p) is generated in advance during the period (e.g., T2) in which the decoding of a certain image (e.g., POC0 and POC4) is completed, so that the processed image of that image (e.g., POC3p) becomes accessible in a period two frames later (e.g., T4).

[0716] Figures 98 and 99 show examples of the first, second, and third images. For example, the images of POC0 and POC4 correspond to one or more first images, the image of POC2 corresponds to the second image, the image of POC1 corresponds to the third image, and the image of POC1p or POC3p corresponds to the processed image (first image after NN processing).

[0717] Note that the POC number used here can be replaced with another method indicating the encoding order of the images. For example, an image number, index, or ID may be used instead of the POC number.

[0718] The following describes another example of operation when an NN interface is used. Figure 100 shows an example of the display order and encoding order of multiple images. In this example, the multiple images include images POC0 to POCn. Also, the target image x is the image of POCx and is the image to be decoded.

[0719] Figure 101 shows a first example of the processing in this case. In the first example, when decoding the target image x, the decoding device starts generating a processed image of POC1p, which is used for interpretation of the image of POC1, which is the next image after the target image x in the encoding order.

[0720] Figure 102 shows a second example of the processing in this case. In the second example, when the decoding device decodes the target image x, it starts generating a POC2p image that is not the next image after the target image x in the encoding order. In this example, the POC2p image is the image two images after the target image x in the encoding order.

[0721] The encoding device may determine the POC (Proof of Concept) to generate the processed image based on the image content and signal information indicating the determined POC. The decoding device generates the processed image of the POC indicated by the information included in the bitstream. However, the encoding device does not need to signal information if the POC of the processed image is a predetermined default POC. In this case, if the decoding device does not include the information in the bitstream, it determines to generate the default POC for processing. For example, the default POC is the POC of the image following the target image in the encoding order.

[0722] Furthermore, the encoding device may signal information indicating two reference images to be used for NN interfacing into the bitstream. The decoding device generates a processed image using the two reference images indicated by this information included in the bitstream. Note that the encoding device does not need to signal this information into the bitstream if the two default reference images are used. In this case, if this information is not included in the bitstream, the decoding device performs NN interfacing using the two default reference images. For example, the two default reference images are the image specified at index 0 of reference picture list L0 and the image specified at index 0 of reference picture list L1. Reference picture lists L0 and L1 are lists of forward and backward reference pictures (reference images) used for bidirectional prediction. In this way, the amount of signaling data can be reduced by omitting information indicating the processed image or information indicating the images used to generate the processed image.

[0723] For example, the images of POC0 and POCn correspond to one or more first images, the target image x corresponds to the second image, the image of POC1 corresponds to the third image, and the images of POC1p or POC2p correspond to the processed images.

[0724] The following describes another example of operation when an NN interface is used. Figure 103 shows an example of the display order and encoding order of multiple images. In this example, images POC0, POC2, and POC4 are reference images that are referenced by other images, while images POC1 and POC3 are non-reference images that are not referenced by other images.

[0725] Figure 104 shows an example of an image stored in a buffer (encoded picture buffer). Figure 104 shows the target image to be decoded and the state of the buffer during the period T0 to T4. Figures 105 to 108 show examples of processing during T1 to T4, respectively.

[0726] At T0, the buffer is empty, and the decoding unit begins decoding the image of POC0. At T1, the decoding of the image of POC0 is complete, and the decoded image of POC0 is stored in the buffer for reference. The decoding unit also begins decoding the image of POC4.

[0727] In T2, the decoding of the POC4 image is complete, and the decoded POC4 image is stored in the buffer. The decoding device also starts generating a new reference image, the processed image (POC2p), using the POC0 and POC4 images through NN interprocessing. The decoding device also starts decoding the POC2 image. In this POC2 decoding process, if interpretation is used, the processed image of POC2p cannot be referenced, and the POC0 and POC4 images are referenced instead.

[0728] At T3, the decoding of the POC2 image is complete, and the decoded POC2 image is stored in the buffer. Also, the NN processing on the POC2p image is complete, and the processed POC2p image, which has been authorized to be referenced by the other images, is stored in the buffer. The decoding device also starts decoding the POC1 image. In the decoding process of the POC1 image, if interpretation is used, the POC0 image, the POC2 image, the POC4 image, and the processed POC2p image are referenced.

[0729] At T4, the decoding of POC1 is complete. The decoding device starts decoding the image of POC3. In this decoding process for POC3, if interpretation is used, the images of POC0, POC2, POC4, and the processed image of POC2p are referenced. Note that the processed image of POC2p may be listed in the reference picture list in any order.

[0730] Thus, the processed image of POC2p is generated before the images of POC1 and POC3 are decoded. Therefore, the processed image of POC2p can be used for interpretation of the images of POC1 and POC3. This increases the number of reference images available for each image. It also makes the reference image closest to the processed image available. These improvements can be made to the results of the interpretation process.

[0731] Figures 109 and 110 show examples of the first, second, and third images. For example, the images of POC0 and POC4 correspond to one or more first images, the image of POC2 corresponds to the second image, the image of POC1 corresponds to the third image, and the image of POC2p corresponds to the processed image (first image after NN processing).

[0732] Note that the POC number used here can be replaced with another method indicating the encoding order of the images. For example, an image number, index, or ID may be used instead of the POC number.

[0733] The following describes an example of the reference relationship when decoding the POC1 image in the decoding of multiple images shown in Figure 103. Figure 111 shows the first example of the reference relationship in this case. In this example, at T3, the decoding process of the POC1 image uses interpretation with the POC0 image, the POC2 image, and the processed image of POC2p.

[0734] Figure 112 shows a second example of the reference relationship in this case. In this example, at T3, the image of POC2 and the processed image of POC2p are used for interpretation of the image of POC1, but the image of POC0 is not used. As a result, the decoding device can improve the encoding quality by referencing the processed image while using fewer reference images.

[0735] Here, for example, the image of POC2 and the image of POC2p represent images displayed at the same time. For example, the images of POC0 and POC4 correspond to one or more first images, the image of POC2 corresponds to the second image, the image of POC1 corresponds to the third image, and the image of POC2p corresponds to the processed image (the first image after NN processing).

[0736] In this embodiment, an example of generating one processed image from two images as input was described as the processing of the NN interface, but it is not limited to this. For example, the same processing can be applied when generating one processed image from one image as input, or when generating one processed image from three or more images as input.

[0737] [First Example of Syntax] The first example of syntax (syntax) indicating the time or period during which a processed image can be referenced or used is described below. The POC number of the processed image may be the same as the current target image, or it may be the POC number of an image that has not yet been decoded. For example, the syntax is encoded in either the header or the data unit.

[0738] Figure 113 shows an example of this syntax. In Figure 113, nn_processed_delay indicates the delay period (delay time) until the NN-processed image becomes available. For example, nn_processed_delay = 0 indicates no delay. In this case, the decoder can immediately access or use the processed image.

[0739] nn_processed_delay=1 indicates a one-image delay. In this case, the decoding device can refer to or use the processed image after decoding the image following the target image.

[0740] nn_processed_delay=n indicates the n-image delay (where n is an integer greater than or equal to 1). In this case, the decoding device can refer to or use the processed image after decoding the nth image from the target image.

[0741] Figure 114 shows another example of syntax indicating the time or period during which a processed image can be referenced or used. The nn_processed_unit shown in Figure 114 represents the unit (granularity) represented by nn_processed_delay. Figure 115 shows an example of a unit represented by nn_processed_delay. For example, as shown in Figure 115, this unit includes picture, slice, tile, subpicture, and combinations thereof. It may also include only a portion of the units represented by nn_processed_delay. Here, slice, tile, and subpicture are units of an image in which a picture is divided.

[0742] In this case, nn_processed_delay indicates the delay period in units represented by nn_processed_delay. For example, if the unit is slices, nn_processed_delay = 1 indicates a delay of 1 slice. In this case, the decoding device can refer to or use the processed image (processed slice) after decoding the next slice of the target slice being processed.

[0743] Figure 116 shows another example of syntax indicating the time or period during which a processed image can be referenced or used. nn_processed_size indicates the data size of the processed image required by the decoder at one-frame intervals. For example, if encoding is performed assuming that one image is processed at two-frame intervals, the data size of the processed image is half the data size of a single image. Thus, the data size processed at one-frame intervals may be signaled instead of the delay period. For example, this data size may be determined based on the processing power or computational power of the neural network and represent the number of CTUs, coding units, or pixel samples processed by the neural network at one-frame intervals.

[0744] [Second example of syntax] Next, we will explain a second example of syntax. Figure 117 shows an example of this syntax. This example is for the case where two images are used in NN processing (for example, an NN interface). As shown in Figure 117, the syntax includes nn_inputA, nn_inputB and bias_offset.

[0745] nn_inputA and nn_inputB represent two images, for example, showing the POC numbers of the two images. bias_offset is information that represents the processed image and shows the POC number. For example, bias_offset is information for calculating the POC number to be assigned to the processed image generated by the NN interface from at least one of nn_inputA and nn_inputB. For example, bias_offset shows the difference (offset) between the POC number shown in nn_inputA and the POC number of the processed image.

[0746] Here, if the relationship between bias_offset and nn_inputA and nn_inputB is a predetermined default relationship, bias_offset does not need to be signaled to the bitstream. Alternatively, bias_offset may be set to a special value such as 0. In this case, the decoder performs the default processing.

[0747] For example, the default process is to set the POC number of the processed image to the average of the POC numbers of the two images. For example, this case applies when the POC number of the processed image is 8, nn_inputA = 0, and nn_inputB = 16.

[0748] Even in such cases, bias_offset may be signaled. For example, in this case, bias_offset is set to 8. The decoding device calculates 8 as the POC number of the processed image by adding 8, indicated by bias_offset, to 0, indicated by inputA.

[0749] Furthermore, if the POC number of the processed image is 4, and nn_inputA = 0 and nn_inputB = 16, then bias_offset = 4 is set. The decoding device calculates 4 as the POC number of the processed image by adding 4, which is indicated by bias_offset, to 0, which is indicated by inputA.

[0750] Furthermore, if the POC number of the processed image is 12, and nn_inputA = 8 and nn_inputB = 16, then bias_offset = 4 is set. The decoding device calculates 12 as the POC number of the processed image by adding 4, which is shown in bias_offset, to 8, which is shown in inputA.

[0751] Furthermore, if the POC number of the processed image is 12, and nn_inputA = 0 and nn_inputB = 16, then bias_offset = 12 is set. The decoding device calculates 12 as the POC number of the processed image by adding 12, indicated by bias_offset, to 0, indicated by inputA. This example is used, for example, when the image of POC0 is of higher quality than the image of POC8.

[0752] Furthermore, if the POC number of the processed image is 2, and nn_input1 = 0 and nn_input2 = 4, then bias_offset = 2 is set. The decoding device calculates 2 as the POC number of the processed image by adding 2, which is shown in bias_offset, to 0, which is shown in inputA.

[0753] Note that the POC number used here can be replaced with another method indicating the encoding order of the images. For example, an image number, index, or ID may be used instead of the POC number.

[0754] Furthermore, while this example shows the use of bias_offset, which indicates the difference, as information indicating the POC number of the processed image, the POC number of the processed image can also be used directly for signaling.

[0755] Furthermore, while this example shows the calculation of the POC number of the processed image by adding bias_offset to the value indicated by nn_inputA, it is not limited to this. For example, the POC number of the processed image may be calculated by adding bias_offset to the value indicated by nn_inputB. Alternatively, the POC number of the processed image may be calculated by adding bias_offset to the value of nn_inputA or nn_inputB, either earlier or later in the decoding order.

[0756] Next, we will describe a variation where two images are used in NN processing. Figure 118 shows an example of the syntax in this case. In the example shown in Figure 118, the syntax includes present_flag in addition to the signals shown in Figure 118. present_flag indicates whether nn_inputA and nn_inputB are signaled (whether they are included in the bitstream).

[0757] When present_flag=1, nn_inputA and nn_inputB are encoded, and the two images represented by nn_inputA and nn_inputB are used in the NN processing.

[0758] If present_flag=0, nn_inputA and nn_inputB are not encoded. In this case, the decoder generates the processed image using the default reference image set. For example, the default reference image set includes the image specified at index 0 of reference picture list L0 and the image specified at index 0 of reference picture list L1. Also, if reference picture list L1 is unavailable, the default reference image set for the reference image includes the two images specified at index 0 and index 1 of reference picture list L0.

[0759] In another example, two or more images may be used in the NN processing. In this case, information indicating all two or more images is signaled during the NN processing.

[0760] In another example, only one image may be used in the NN processing. In this case, information indicating that only one image is being used in the NN processing is signaled.

[0761] Next, we will explain an example of syntax for generating a processed image by directly applying NN processing to the target image (e.g., NN LF). Figure 119 shows an example of the syntax in this case. As shown in Figure 119, the syntax includes nn_inputA.

[0762] nn_inputA indicates the image used for NN processing, for example, the POC number of the image.

[0763] In this case, nn_inputA is set for each POC number on which NN processing is performed. That is, for POCXp (X = 0, 1, 2, 3, 4, 5...), X nn_inputA are set.

[0764] Figure 120 shows a modified example of the syntax used when generating a processed image by directly applying NN processing to the target image. The syntax shown in Figure 120 includes present_flag in addition to the signals shown in Figure 119. present_flag indicates whether or not nn_inputA is signaled (whether or not it is included in the bitstream).

[0765] If present_flag=1, nn_inputA is encoded, and the image represented by nn_inputA is used for NN processing.

[0766] If present_flag = 0, nn_inputA is not encoded. In this case, the decoder generates the processed image using the default image. For example, the default image is the image that was most recently decoded.

[0767] In another example, two or more images may be used in the NN processing. In this case, information indicating all two or more images is signaled during the NN processing.

[0768] [Third example of syntax] Next, we will explain the third example of syntax. Figure 121 shows an example of this syntax. This example is for the case where one image is used in NN processing (for example, NN LF). As shown in Figure 121, the syntax includes is_nn_processed_flag[i][j].

[0769] is_nn_processed_flag[i][j] is a flag that indicates whether each reference image in the reference picture list is pre-NN processed or post-NN processed. Here, i represents reference picture list 0 or 1, and j represents one of the reference images in the reference picture list.

[0770] For example, this syntax can be found in either a header or a data unit.

[0771] This is_nn_processed_flag may be associated with each of the multiple images in the encoded picture buffer, regardless of whether the image is subjected to NN processing or not. Figure 122 shows an example of is_nn_processed_flag associated with multiple images stored in a buffer.

[0772] For example, if is_nn_processed_flag[i][j] = 0, the image associated with this flag is before NN processing. If is_nn_processed_flag[i][j] = 1, the image associated with this flag is after NN processing.

[0773] [Variations] Although the first, second, and third examples of syntax were explained individually above, two or more of these may be applied together. Alternatively, two or more of these may be applied in combination.

[0774] NN Inter and NN LF may be combined. The processed image may include the processed image of NN Inter and the processed image of NN LF. Information indicating a common period for the period during which the processed image of NN Inter can be referenced and the period during which the processed image of NN LF can be referenced may be signaled. Alternatively, information indicating the period during which the processed image of NN Inter can be referenced and information indicating the period during which the processed image of NN LF can be referenced may be signaled separately.

[0775] NN processing may also involve processing using neural networks other than NN inter and NN LF. For example, other NN processing may include intra prediction, inter prediction, transformation, quantization, entropy coding, or reconstruction.

[0776] The information indicating the period during which the processed image can be accessed may indicate the delay between the decoding time and the display time, rather than the delay between the decoding time and the time the processed image becomes accessible.

[0777] The time required to generate the processed image using the neural network may be signaled to the bitstream. This time may be determined based on the processing speed of the neural network.

[0778] One or more parameters may include the POC number of the processed image. For example, the decoding process for the second and third images is an interpretation process. In another example, the decoding process for the second and third images is a loop filtering process. The decoding process may include any combination of multiple decoding processes or decoding processes including neural network processing.

[0779] Furthermore, while the above explanation mainly described the processing in the decoding device, the processing in the encoding device is similar. In other words, the processing in the encoding device is the same as in the above explanation, but with "decoding" replaced by "encoding".

[0780] [Summary of the First Embodiment] As described above, the decoding device according to the embodiment performs the processing shown in Figure 123. Figure 123 is a flowchart of the decoding process by the decoding device. The decoding device applies NN (neural network) processing using a neural network to the decoded first image (S301), decodes the second image which is decoded later than the first image by referring to the first image before NN processing without referring to the first image after NN processing (S302), and decodes the third image which is decoded later than the second image by referring to the first image after NN processing (S303).

[0781] According to this, the decoding device does not refer to the first image after NN processing when decoding the second image, but it does refer to the first image after NN processing when decoding the third image. This allows the decoding of the second image to be performed appropriately even if the generation of the first image after NN processing is not completed at the time of decoding the second image. Furthermore, encoding efficiency can be improved by using the first image after NN processing for the third image. In this way, the decoding device can use the time-consuming NN processing for decoding multiple images, thereby improving encoding efficiency.

[0782] For example, the decoding device performs NN processing on the first image in parallel with the decoding of the second image. This allows the decoding device to generate the first image after NN processing early. Furthermore, when performing parallel processing, the first image after NN processing may not be immediately available for other processing (e.g., decoding). Even in such cases, the NN processing can be used in the image decoding process without any problems.

[0783] For example, each of the first image, second image, and third image is a picture, or a subpicture, tile, or slice that is a sub-region of a picture. For example, the decoder obtains information from the stream regarding the delay amount from the time of the first image to the time when the first image after NN processing becomes accessible (e.g., nn_processed_delay, nn_processed_size, bias_offset, or is_nn_processed_flag), and the delay amount information is obtained from the stream's sequence header, picture header, slice header, or extended information area (e.g., SEI). Based on this, the decoder can determine the time when the first image after NN processing becomes accessible based on the delay amount information.

[0784] For example, NN processing is a loop filtering process (e.g., NN LF) that generates an image by applying a filtering process using a neural network to the first image. This allows for improvement of the image quality of the reconstructed image through filtering using a neural network.

[0785] For example, the decoding device obtains information from the stream regarding the delay amount from the time of the first image to the time when the first image after NN processing becomes accessible. This delay information corresponds to the respective images referenced in the second and third images, and is a flag (e.g., is_nn_processed_flag) that indicates whether the corresponding image is the image before or after NN processing.

[0786] According to this, the decoding device can determine the time at which the first image after NN processing becomes accessible, based on information regarding the delay amount. Furthermore, the decoding device can easily determine whether each image is an image before or after NN processing using this flag.

[0787] For example, NN processing is an interpredictive reference image generation process (e.g., NN inter) that uses a neural network to generate an image corresponding to a display time different from any of the first images, from one or more first images including the first image. This allows for improved coding efficiency through interpredictive processing using a neural network.

[0788] For example, the interpretation and reference image generation process generates an image corresponding to the display time of the second or third image. This allows for improved coding efficiency through interpretation and prediction processing using a neural network.

[0789] For example, the decoding device obtains information from the stream regarding the delay amount from the time of the first image to the time when the first image after NN processing becomes accessible. This delay information includes a value that identifies each of the one or more first images (e.g., nn_inputA and / or nn_inputB) and a value that identifies the image corresponding to the display time of the image generated by the NN processing (e.g., bias_offset). Based on this, the decoding device can determine, based on the delay information, which of the one or more first images are used for NN processing and which are generated by the NN processing.

[0790] For example, the decoding device obtains information from the stream regarding the delay amount from the time of the first image to the time when the first image after NN processing becomes accessible. The information regarding the delay amount (e.g., nn_inputA and / or nn_inputB, and bias_offset) indicates the POC (Picture Order Count) value assigned to each of the one or more first images, and the POC value corresponding to the display time of the image generated by NN processing. Based on this, the decoding device can determine, based on the information regarding the delay amount, one or more first images to be used for NN processing and the image generated by NN processing.

[0791] For example, in NN processing, one or more first images, including the first image, are used, and information about the delay amount (e.g., nn_processed_delay) indicates the number of images from the last first image to the third image in the decoding order among the one or more first images. This allows for a reduction in the amount of data related to the delay amount.

[0792] For example, in NN processing, one or more first images, including the first image, are used, and the information regarding the delay amount (e.g., nn_processed_delay) represents the difference between the POC (Picture Order Count) value assigned to the last first image in the decoding order among the one or more first images, and the POC value assigned to the third image. This reduces the amount of data related to the delay amount.

[0793] For example, information regarding the delay amount (e.g., nn_processed_size) is information regarding the area of ​​a picture that can be processed by a neural network (NN) in the time interval of one frame. Based on this, the decoder can calculate the delay amount using the information regarding the delay amount.

[0794] For example, information about the delay amount (e.g., nn_processed_delay) is information about the processing time required for NN processing of one picture. This allows for a reduction in the amount of data related to the delay amount.

[0795] The decoding device has a configuration similar to that of the decoding device 200 shown in Figure 5, for example. The decoding device 200 comprises a processor b1 (or circuit) and a memory b2, and the processor b1 (or circuit) uses the memory b2 to perform the above processing during operation.

[0796] Furthermore, the encoding device according to the embodiment performs the processing shown in Figure 124. Figure 124 is a flowchart of the encoding process by the encoding device. The encoding device applies NN (neural network) processing using a neural network to the encoded first image (S401), encodes the second image which is encoded after the first image by referring to the first image before NN processing without referring to the first image after NN processing (S402), and encodes the third image which is encoded after the second image by referring to the first image after NN processing (S403).

[0797] According to this, the encoding device does not refer to the first image after NN processing when encoding the second image, but it does refer to the first image after NN processing when encoding the third image. This allows the encoding of the second image to be performed appropriately even if the generation of the first image after NN processing is not completed at the time of encoding the second image. Furthermore, encoding efficiency can be improved by using the first image after NN processing for the third image. In this way, the encoding device can use the time-consuming NN processing for encoding multiple images, thereby improving encoding efficiency.

[0798] For example, the NN processing of the first image is performed in parallel with the encoding of the second image. This allows the encoding device to generate the first image after NN processing early. Furthermore, when performing parallel processing, the first image after NN processing may not be immediately available for other processing (e.g., encoding). Even in such cases, the NN processing can be used for image encoding without any problems.

[0799] For example, the first image, second image, and third image are each a picture, or a subpicture, tile, or slice which are sub-regions of a picture. For example, the encoding device writes information about the delay amount from the time of the first image to the time when the NN-processed image becomes accessible (e.g., nn_processed_delay, nn_processed_size, bias_offset, or is_nn_processed_flag) to a stream, and the delay amount information is written in the sequence header, picture header, slice header, or extended information area (e.g., SEI) of the stream. In this way, the encoding device...

Claims

1. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, decodes computing power information associated with a neural network from a bitstream, and decodes an image from the bitstream using the neural network based on the computing power information.

2. The decoding device according to claim 1, wherein the circuit further determines in the operation whether the computing power of the decoding device satisfies the computing power indicated by the computing power information, and if it is determined that the computing power of the decoding device satisfies the computing power indicated by the computing power information, generates a first image from the bitstream using the neural network, and decodes a second image from the bitstream by referring to the first image.

3. The decoding device according to claim 2, wherein the bitstream further includes timing information, and the circuit generates the first image from the bitstream using the neural network at the timing indicated by the timing information when it is determined that the computing power of the decoding device satisfies the computing power indicated by the computing power information.

4. The bitstream further includes calculation count information indicating the total number of calculations performed on an input image input to the neural network, and the circuit, in the determination, derives the GPU processing time of the decoder from the total number of calculations and the GPU (Graphics Processing Unit) inference speed of the decoder, and determines whether the GPU processing time of the decoder satisfies the computing capability indicated by the computing capability information, thereby determining whether the computing capability of the decoder satisfies the computing capability indicated by the computing capability information.

5. The decoding device according to claim 1, wherein the computing power information indicates a GPU profile, and the GPU profile defines a set of decoding tools comprising a neural network.

6. The decoding device according to claim 1, wherein the computing power information indicates one of a plurality of levels of the GPU profile, and each of the plurality of levels defines an upper limit of the computing power of the GPU.

7. The decoding device according to claim 1, wherein the computing power information indicates one of a plurality of tiers of GPU profiles, and each of the plurality of tiers defines an upper limit on the computing power of the GPU.

8. The decoding device according to claim 1, wherein the computing power information includes at least one of the following: the maximum resolution of the input image input to the neural network, the maximum inference speed of the GPU, the maximum memory of the GPU, the accuracy of the neural network, and the maximum GPU processing time.

9. The decoding device according to claim 1, wherein the computing power information is included in VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), picture header, slice header, APS (Adaptation Parameter Set), SEI (Supplemental Enhancement Information), VUI (Video Usability Information), tile header, or metadata.

10. The decoding device according to claim 3, wherein the timing information is included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

11. The decoding device according to claim 4, wherein the calculation count information is included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

12. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit generates encoded data by encoding an image using a neural network during operation, and generates a bitstream including the encoded data and computing power information associated with the neural network.

13. The encoding device according to claim 12, wherein the circuit further determines in the operation whether the computing power of the encoding device satisfies the computing power indicated by the computing power information, and if it is determined that the computing power of the encoding device satisfies the computing power indicated by the computing power information, generates a first image using the neural network, and encodes a second image by referring to the first image.

14. The encoding apparatus according to claim 12, wherein the bitstream further includes timing information indicating the timing at which the decoding apparatus generates a first image from the bitstream using the neural network.

15. The encoding device according to claim 12, wherein the bitstream further includes calculation count information indicating the total number of calculations performed on the input image input to the neural network.

16. The encoding apparatus according to claim 12, wherein the computing power information indicates a GPU profile, and the GPU profile defines a set of encoding tools comprising a neural network.

17. The encoding device according to claim 12, wherein the computing power information indicates one of a plurality of levels of the GPU profile, and each of the plurality of levels defines an upper limit of the computing power of the GPU.

18. The encoding device according to claim 12, wherein the computing power information indicates one of a plurality of tiers of GPU profiles, and each of the plurality of tiers defines an upper limit on the computing power of the GPU.

19. The encoding device according to claim 12, wherein the computing power information includes at least one of the following: the maximum resolution of the input image input to the neural network, the maximum inference speed of the GPU, the maximum memory of the GPU, the accuracy of the neural network, and the maximum GPU processing time.

20. The encoding device according to claim 12, wherein the computing power information is included in VPS (Video Parameter Set), SPS (Sequence Parameter Set), PPS (Picture Parameter Set), picture header, slice header, APS (Adaptation Parameter Set), SEI (Supplemental Enhancement Information), VUI (Video Usability Information), tile header, or metadata.

21. The encoding device according to claim 14, wherein the timing information is included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

22. The encoding apparatus according to claim 15, wherein the calculation count information is included in VPS, SPS, PPS, picture header, slice header, APS, SEI, VUI, tile header, or metadata.

23. A decoding method for decoding computing power information associated with a neural network from a bitstream, and decoding an image from the bitstream using the neural network based on the computing power information.

24. An encoding method that generates encoded data by encoding an image using a neural network, and generates a bitstream including the encoded data and computing power information associated with the neural network.