Encoding device, decoding device, encoding method, and decoding method
By generating a second reference image through neural networks or transformation processes, the proposed solution addresses inefficiencies in existing video coding technologies, enhancing encoding efficiency and image quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2026-01-13
- Publication Date
- 2026-07-23
AI Technical Summary
Existing video coding technologies face challenges in improving encoding efficiency, image quality, processing load, and circuit size, particularly in selecting optimal elements and operations such as filters, blocks, motion vectors, and reference pictures.
The proposed solution involves generating a second reference image from a first reference image using a neural network or predetermined transformation processes, such as scaling, shifting, rotation, or mirroring, to enhance prediction accuracy and reduce prediction residuals, thereby improving encoding efficiency.
This approach allows for improved encoding efficiency by generating a reference image closer to the current picture, reducing prediction residuals, and enhancing overall coding efficiency.
Smart Images

Figure JP2026000682_23072026_PF_FP_ABST
Abstract
Description
Encoding device, decoding device, encoding method, and decoding method
[0001] This disclosure relates to encoding devices, etc.
[0002] Video coding technology has advanced from H.261 and MPEG-1 to H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding). With this advancement, there is a constant need to provide improvements and optimizations to video coding technology to handle the ever-increasing volume of digital video data in various applications. This disclosure relates to further advancements, improvements, and optimizations in video coding.
[0003] Non-patent document 1 relates to an example of a conventional standard concerning the video coding technology described above.
[0004] H. 266 (ISO / IEC 23090-3) / VVC (Versatile Video Coding)
[0005] Regarding the encoding methods described above, there is a need for proposals for new methods to improve encoding efficiency, image quality, processing load, circuit size, or to appropriately select elements or actions such as filters, blocks, size, motion vectors, reference pictures, or reference blocks.
[0006] This disclosure provides a configuration or method that can contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, improved processing applied to the decoded image, or provision of new processing. This disclosure may also include configurations or methods that can contribute to other benefits not mentioned above.
[0007] For example, an encoding device according to one aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit uses the memory to encode a parameter set into a bitstream, encode a first image into the bitstream, derives a first reference image by reconstructing the first image, and generates a second reference image from the first reference image using the parameter set.
[0008] Each embodiment, or any part thereof, of the present disclosure enables at least one of the following: improved encoding efficiency, improved image quality, reduced encoding / decoding processing load, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment, or any part thereof, of the present disclosure enables appropriate selection of components / operations such as filters, blocks, sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. The present disclosure also includes disclosures of configurations or methods that may provide benefits other than those mentioned above, such as configurations or methods that improve encoding efficiency while suppressing an increase in processing load.
[0009] Further advantages and effects of one aspect of this disclosure will be made apparent from the specification and drawings. Such advantages and / or effects may be obtained by several embodiments and features described in the specification and drawings, but not all of them are necessarily provided to obtain one or more advantages and / or effects.
[0010] These general or specific embodiments may be implemented as a system, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of system, method, integrated circuit, computer program, and recording medium.
[0011] A configuration or method relating to one aspect of this disclosure may contribute to one or more of the following: improved encoding efficiency, improved image quality, reduced processing load, reduced circuit size, improved processing speed, and appropriate selection of elements or operations. A configuration or method relating to one aspect of this disclosure may also contribute to other benefits not mentioned above.
[0012] This is a schematic diagram showing an example of the configuration of a transmission system according to an embodiment. This is a block diagram showing an example of the implementation of an encoding device according to an embodiment. This is a block diagram showing an example of the configuration of an encoding device according to an embodiment. This is a flowchart showing an example of the overall encoding process by the encoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the encoding device according to an embodiment. This is a block diagram showing an example of the implementation of a decoding device according to an embodiment. This is a block diagram showing an example of the configuration of a decoding device according to an embodiment. This is a flowchart showing an example of the overall decoding process by the decoding device according to an embodiment. This is a flowchart showing an example of the processing for each block by the decoding device according to an embodiment. This is a schematic diagram showing another example of the configuration of a transmission system according to an embodiment. This is a schematic diagram showing an example of the configuration of a transmitting device according to an embodiment. This is a block diagram showing an example of the implementation of a transmitting device according to an embodiment. This is a schematic diagram showing an example of the configuration of a receiving device according to an embodiment. This is a block diagram showing an example of the implementation of a receiving device according to an embodiment. This is a schematic diagram showing an example of the configuration of a bitstream generation device according to an embodiment. This is a block diagram showing an example of the implementation of a bitstream generation device according to an embodiment. This is a schematic diagram showing an example of the configuration of a storage medium and a computer according to an embodiment. This is a schematic diagram showing an example of the configuration of a computer according to an embodiment. This is a diagram showing an example of the hierarchical structure of data in a stream. This is a diagram showing an example of the configuration of a slice. This is a diagram showing an example of the configuration of a tile. This is a diagram showing an example of the configuration of a time-scalable stream. This is a diagram showing an example of the configuration of a stream using multilayer encoding. This is a diagram showing an example of block partitioning. This is a diagram showing an example of the configuration of the split section. This is a diagram showing an example of a split pattern. This is a diagram showing an example of a syntax tree for a split pattern. This is a diagram showing another example of a syntax tree for a split pattern. This is a block diagram showing an example of the configuration of the loop filter section. This is a diagram showing an example of the shape of the filter used in ALF (adaptive loop filter). This is a diagram showing another example of the shape of the filter used in ALF. This is a diagram showing another example of the shape of the filter used in ALF. This is a diagram showing an example of CCALF. This is a diagram showing the filter shape of CCALF.This is a block diagram showing an example of a detailed configuration of the loop filter section that functions as a DBF. This is a flowchart showing an example of the processing of the DBF processing section. This is a flowchart showing another example of the processing of the DBF processing section. This is a diagram showing an example of a DBF having filter characteristics symmetric with respect to the block boundary. This is a diagram to explain an example of a block boundary on which DBF processing is performed. This is a diagram showing an example of the block boundary strength Bs value. This is a block diagram showing an example of the configuration of the loop filter section. This is a flowchart showing an example of the processing performed in the prediction section of the encoding device. This is a flowchart showing another example of the processing performed in the prediction section of the encoding device. This is a flowchart showing another example of the processing performed in the prediction section of the encoding device. This is a diagram showing an example of each reference picture. This is a conceptual diagram showing an example of a reference picture list. This is a conceptual diagram showing another example of a reference picture. This is a conceptual diagram showing another example of generating a generated reference picture. This is a conceptual diagram showing another example of generating a generated reference picture. This is a flowchart showing the basic processing flow of inter prediction. This is a flowchart showing an example of MV derivation. This is a flowchart showing another example of MV derivation. This is a diagram showing an example of the classification of each mode of MV derivation. This is a diagram showing an example of the classification of each mode of MV derivation. This is a flowchart showing an example of inter prediction using normal inter mode. This is a flowchart showing an example of inter prediction using normal merge mode. This is a diagram illustrating an example of MV derivation using normal merge mode. This is a diagram illustrating an example of MV derivation using HMVP (History-based Motion Vector Prediction / Predictor) mode. This is a flowchart illustrating an example of FRUC (frame rate up conversion). This is a diagram illustrating an example of pattern matching (bilateral matching) between two blocks along a motion trajectory. This is a diagram illustrating an example of pattern matching (template matching) between a template in the current picture and a block in a reference picture. This is a diagram illustrating an example of MV derivation at the subblock level in affine mode using two control points.This is a diagram illustrating an example of deriving the MV per subblock in affine mode using three control points. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a conceptual diagram illustrating an example of deriving the MV of control points in affine mode. This is a diagram illustrating an affine mode with two control points. This is a diagram illustrating an affine mode with three control points. This is a conceptual diagram illustrating an example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block. This is a conceptual diagram illustrating another example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block. This is a flowchart illustrating an example of processing in affine merge mode. This is a flowchart illustrating an example of processing in affine inter mode. This is a diagram illustrating an example of two regions where a predicted image is generated by GPM. This is a conceptual diagram showing the pattern of two regions defined by GPM. This is a conceptual diagram showing the weighted average of pixel values at region boundaries. This is a flowchart illustrating an example of GPM mode. This figure shows an example of the ATMVP (Advanced Temporal Motion Vector Prediction / Predictor) mode in which MV is derived for each subblock. This figure shows the relationship between merge mode and DMVR (decoder motion vector refinement). This is a conceptual diagram to explain an example of DMVR. This is a conceptual diagram to explain another example of DMVR for determining MV. This figure shows an example of motion search in DMVR. This is a flowchart showing an example of motion search in DMVR. This is a flowchart showing an example of predictive image generation. This is a flowchart showing another example of predictive image generation. This is a flowchart to explain an example of predictive image correction processing by OBMC (overlapped block motion compensation). This is a conceptual diagram to explain an example of predictive image correction processing by OBMC. This figure shows a model that assumes uniform linear motion.This is a flowchart showing an example of inter prediction according to BIO. This is a diagram showing an example of the configuration of an inter prediction unit that performs inter prediction according to BIO. This is a diagram to explain an example of a prediction image generation method using brightness correction processing by LIC (local illumination compensation). This is a flowchart showing an example of a prediction image generation method using brightness correction processing by LIC. This is a conceptual diagram showing an example of a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing pipeline processing. This is a block diagram showing an example of a configuration used in NN processing. This is a block diagram showing another example of a configuration used in NN processing. This is a conceptual diagram showing a patch used in NN processing. This is a conceptual diagram showing an extended patch that includes an extended region. This is a conceptual diagram showing an image containing two patches with overlapping regions. This is a conceptual diagram showing a patch defined in each picture. This is a conceptual diagram showing an example of a patch generation method. This is a conceptual diagram showing another example of a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a flowchart showing an example of an encoding process including the generation of a generated reference image by an encoding device according to an embodiment. This is a flowchart showing an example of a decoding process including the generation of a generated reference image by a decoding device according to an embodiment. This is a diagram showing an example of a syntax structure used when generating a generated reference image. This is a conceptual diagram showing an example of a method for generating a generated reference image when there is one input image. This is a conceptual diagram showing an example of a method for generating a generated reference image when there are multiple input images. This is a diagram showing an example of a syntax structure used when generating a generated reference image using a conversion method. This is a diagram showing another example of a syntax structure used when generating a generated reference image using a conversion method. This is a table showing the relationship between a conversion method and a value. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture. This is a conceptual diagram showing another example of generating a generated picture.This is a flowchart showing an example of encoding processing including a method for managing generated reference images by an encoding device according to an embodiment. This is a flowchart showing an example of decoding processing including a method for managing generated reference images by a decoding device according to an embodiment. This is a diagram showing an example of a syntax structure used when reassigning a reference index. This is a diagram showing an example of a syntax structure used when reassigning a reference index. This is a diagram showing an example of a syntax structure used when reassigning a reference index. This is a conceptual diagram showing the POC assignment result when no generated reference image is generated for a reconstructed reference image. This is a conceptual diagram showing the POC number reassignment result when there is one generated reference image during RPL construction. This is a diagram showing the POC number reassignment result when two generated reference images are generated for a reconstructed reference image. This is a conceptual diagram showing the POC assignment result when no generated reference image is generated for a reconstructed reference image. This is a conceptual diagram showing the POC assignment result when there is one generated reference image during RPL construction. This is a conceptual diagram showing the result of recalculating the POC assignment when there is one generated reference image. This is a diagram showing an example where POC'4 is set as a long-term reference image when two generated reference images are generated for a reconstructed reference image. This figure shows the order of short-term and long-term reference images for POC'6. This figure shows another example of assigning a reference index. This figure shows the syntax structure used when assigning a reference index. This figure shows the syntax structure used when assigning a reference index. This figure shows the result of reallocating the reference index when there are no generated reference images. This figure shows the result of reallocating the reference index when one generated reference image is generated for a reconstructed reference image. This figure shows the result of reallocating the reference index when two generated reference images are generated for a reconstructed reference image. This figure shows an example where POC1 is set as the long-term reference image when two generated reference images are generated for a non-generated reference image. This figure shows the order of short-term and long-term reference images for POC2 and the current picture with idx=0. This figure shows the result of reallocating the reference index when there are no generated reference images.This figure shows the result of reallocating the reference index when one generated reference image is created from a reconstructed reference image. This figure shows the result of reallocating the reference index when two generated reference images are created from a reconstructed reference image. This figure shows the L0 list. This figure shows the syntax structure when reference indices assigned according to different rules are given to the reconstructed reference image and the generated reference image. This flowchart shows an example of encoding processing including the process of selecting a collated image by the encoding device according to the embodiment. This flowchart shows an example of decoding processing including the process of selecting a collated image by the decoding device according to the embodiment. This figure shows an example of the overall configuration of a content supply system that realizes a content distribution service. This figure shows an example of the configuration of a content distribution system. This figure shows an example of a web page display screen. This figure shows an example of a web page display screen. This figure shows an example of a smartphone. This is a block diagram showing an example of a smartphone configuration.
[0013] [Introduction] This disclosure relates to video encoding, and more particularly to systems, components, and methods for encoding and decoding video. For example, this disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method for interpretation using a generated reference image.
[0014] For example, an encoding device generates a predicted image of the current block by referencing a reference image stored in frame memory and making a prediction for the current block. By improving the prediction accuracy of the generated predicted image, the prediction residual can be reduced, and as a result, the amount of code can be reduced. In other words, encoding efficiency can be improved. However, there are cases where a suitable predicted image is not generated from the existing reference image, and encoding efficiency deteriorates.
[0015] Therefore, the encoding device of Example 1 comprises a circuit and a memory connected to the circuit, the circuit uses the memory to encode a parameter set into a bitstream, encode a first image into the bitstream, derive a first reference image by reconstructing the first image, and generate a second reference image from the first reference image using the parameter set.
[0016] This allows for the generation of a predicted image using a second reference image generated from the first reference image. If the second reference image is closer to the current picture than the first reference image, the prediction residual can be reduced. Therefore, it may be possible to improve coding efficiency.
[0017] Furthermore, the encoding device in Example 2 is the encoding device in Example 1, and the second reference image is generated by inputting the first reference image into a neural network that transforms and outputs an input image, and the parameter set may include information indicating the neural network.
[0018] This can sometimes make it possible to generate a second reference image that is closer to the current picture than the first reference image using a neural network. For example, it may be possible to select a reference image with content more similar to the current picture from the list of reference pictures, thus improving encoding efficiency.
[0019] Furthermore, the encoding device in Example 3 is the encoding device in Example 1 or 2, the second reference image is generated by performing a predetermined transformation process on the first reference image, and the parameter set may include information indicating the predetermined transformation process.
[0020] This may make it possible to generate a second reference image that is closer to the current picture than the first reference image using a predetermined conversion process. For example, it may be possible to select a reference image with content more similar to the current picture from the reference picture list, thus improving encoding efficiency.
[0021] Furthermore, the encoding device of Example 4 is the encoding device of Example 3, and the predetermined conversion process may include at least one of the following: a scaling process to enlarge or reduce the first reference image; a shifting process to shift the first reference image in one direction; a rotation process to rotate the first reference image; a padding process to replace the pixel values in the first reference image; and a mirroring process to invert the first reference image with respect to predetermined axis coordinates.
[0022] This means that, with the appropriate conversion process, it may be possible to generate a second reference image that is closer to the current picture than the first reference image.
[0023] Furthermore, the encoding device in Example 5 may be any of the encoding devices in Examples 1 to 4, and the process of generating the second reference image from the first reference image may be performed on a picture-by-picture basis.
[0024] This can sometimes reduce the number of times the first reference image is read when generating the second reference image.
[0025] Furthermore, the encoding device in Example 6 may be any of the encoding devices in Examples 1 to 5, and may generate the second reference image based on the first reference image and a third reference image derived by reconstructing a second image different from the first image.
[0026] This approach, by using two reference images, can sometimes effectively improve the prediction accuracy of the second reference image.
[0027] Furthermore, the encoding device in Example 7 is any of the encoding devices in Examples 1 to 6, and the parameter set may include information indicating a generation method for generating the second reference image from the first reference image.
[0028] This can sometimes enable the decoding device to properly decode the bitstream.
[0029] Furthermore, the decoding device of Example 8 comprises a circuit and a memory connected to the circuit, the circuit uses the memory to decode a parameter set from a bitstream, derives a first reference image by decoding a first image from the bitstream, and generates a second reference image from the first reference image using the parameter set.
[0030] This allows for the generation of a predicted image using a second reference image generated from the first reference image. If the second reference image is closer to the current picture than the first reference image, the prediction residual can be reduced. Therefore, it may be possible to improve coding efficiency.
[0031] Furthermore, the decoding device in Example 9 is the decoding device in Example 8, wherein the second reference image is generated by inputting the first reference image into a neural network that transforms and outputs an input image, and the parameter set may include information indicating the neural network.
[0032] This can sometimes make it possible to generate a second reference image that is closer to the current picture than the first reference image using a neural network. For example, it may be possible to select a reference image with content more similar to the current picture from the list of reference pictures, thus improving encoding efficiency.
[0033] Furthermore, the decoding device in Example 10 is the decoding device in Example 8 or 9, wherein the second reference image is generated by performing a predetermined conversion process on the first reference image, and the parameter set may include information indicating the predetermined conversion process.
[0034] This may make it possible to generate a second reference image that is closer to the current picture than the first reference image using a predetermined conversion process. For example, it may be possible to select a reference image with content more similar to the current picture from the reference picture list, thus improving encoding efficiency.
[0035] Furthermore, the decoding device of Example 11 is the decoding device of Example 10, and the predetermined conversion process may include at least one of the following: a scaling process to enlarge or reduce the first reference image; a shifting process to shift the first reference image in one direction; a rotation process to rotate the first reference image; a padding process to replace the pixel values in the first reference image; and a mirroring process to invert the first reference image with respect to predetermined axis coordinates.
[0036] This can sometimes make it possible to generate a second reference image that is closer to the current picture than the first reference image, provided that appropriate conversion processing is used. For example, it may be possible to select a reference image with content more similar to the current picture from the reference picture list, thus improving encoding efficiency.
[0037] Furthermore, the decoding device in Example 12 may be any decoding device from Examples 8 to 11, and the process of generating the second reference image from the first reference image may be performed on a picture-by-picture basis.
[0038] This can sometimes reduce the number of times the first reference image is read when generating the second reference image.
[0039] Furthermore, the decoding device in Example 13 may be any decoding device from Examples 8 to 12, which generates the second reference image based on the first reference image and a third reference image derived by reconstructing a second image different from the first image.
[0040] This approach, by using two reference images, can sometimes effectively improve the prediction accuracy of the second reference image.
[0041] Furthermore, the decoding device in Example 14 is any decoding device from Examples 1 to 13, and the parameter set may include information indicating a generation method for generating the second reference image from the first reference image.
[0042] This can sometimes enable the decoding device to properly decode the bitstream.
[0043] Furthermore, the encoding method of Example 15 involves encoding a parameter set into a bitstream, encoding a first image into the bitstream, reconstructing the first image to derive a first reference image, and generating a second reference image from the first reference image using the parameter set.
[0044] This allows for the generation of a predicted image using a second reference image generated from the first reference image. If the second reference image is closer to the current picture than the first reference image, the prediction residual can be reduced. Therefore, it may be possible to improve coding efficiency.
[0045] Furthermore, the decoding method of Example 16 involves decoding a parameter set from a bitstream, deriving a first reference image by decoding a first image from the bitstream, and generating a second reference image from the first reference image using the parameter set.
[0046] This allows for the generation of a predicted image using a second reference image generated from the first reference image. If the second reference image is closer to the current picture than the first reference image, the prediction residual can be reduced. Therefore, it may be possible to improve coding efficiency.
[0047] Furthermore, the encoding device of Example 17 comprises an input unit, a division unit, an intra-prediction unit, an inter-prediction unit, a loop filter unit, a conversion unit, a quantization unit, an entropy encoding unit, and an output unit.
[0048] The current picture is input to the input unit. The division unit divides the current picture into multiple blocks.
[0049] The intra prediction unit generates a prediction signal for the current block contained in the current picture using the reference pixels contained in the current picture. The inter prediction unit generates a prediction signal for the current block contained in the current picture using the reference blocks contained in a reference picture different from the current picture. The loop filter unit applies a filter to the reconstructed blocks of the current block contained in the current picture.
[0050] The conversion unit converts the prediction error between the original signal of the current block included in the current picture and the predicted signal generated by the intra-prediction unit or the inter-prediction unit to generate conversion coefficients. The quantization unit quantizes the conversion coefficients to generate quantization coefficients. The entropy coding unit applies variable-length coding to the quantization coefficients to generate an encoded bitstream. The encoded bitstream, which includes the quantization coefficients to which variable-length coding has been applied and control information, is output from the output unit.
[0051] Then, the entropy coding unit encodes the parameter set into a bitstream, encodes the first image into the bitstream, the inter-prediction unit or the intra-prediction unit derives a first reference image by reconstructing the first image, and generates a second reference image from the first reference image using the parameter set.
[0052] Furthermore, the decoding device of Example 18 comprises an input unit, an entropy decoding unit, an inverse quantization unit, an inverse transform unit, an intra prediction unit, an inter prediction unit, a loop filter unit, and an output unit.
[0053] The input unit receives an encoded bitstream. The entropy decoding unit applies variable-length decoding to the encoded bitstream to derive quantization coefficients. The inverse quantization unit dequantizes the quantization coefficients to derive conversion coefficients. The inverse transformation unit inversely transforms the conversion coefficients to derive prediction errors.
[0054] The intra prediction unit generates a prediction signal for the current block contained in the current picture using the reference pixels contained in the current picture. The inter prediction unit generates a prediction signal for the current block contained in the current picture using the reference blocks contained in a reference picture different from the current picture.
[0055] The loop filter unit applies a filter to the reconstructed block of the current block included in the current picture. Then, the current picture is output from the output unit.
[0056] Then, the entropy decoding unit decodes the parameter set from the bitstream, the inter-prediction unit or the intra-prediction unit derives a first reference image by decoding the first image from the bitstream, and generates a second reference image from the first reference image using the parameter set.
[0057] Furthermore, these comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or non-temporary recording media such as computer-readable CD-ROMs, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0058] [Definition of Terms] Each term may be defined as follows, for example:
[0059] (1) Image: A unit of data composed of a collection of pixels, consisting of pictures and smaller blocks, and includes both still images and videos.
[0060] (2) Chroma Chroma represents a sample sequence or single sample that represents one of two color difference signals. The two color difference signals may be represented by the symbols Cb and Cr. Chroma is an adjective, for example, represented by the symbols Cb and Cr, or U and V. The term chrominance can also be used instead of chroma.
[0061] (3) Luminance (luma) Luminance refers to a sample sequence or a single sample representing a monochrome signal associated with the primary colors. Luminance is an adjective represented by the symbols Y or L. The term luminance can also be used instead of lumina.
[0062] (4) Component A component represents a single array or sample. A component may be, for example, one of the three arrays (i.e., luminance and two color differences) that make up a color format picture, or a single sample, or one array or a single sample that makes up a monochrome format picture.
[0063] (5) A picture is an image processing unit composed of a set of picture samples, and is sometimes called a frame or field. A picture can be a set of luminance samples, a set of two color difference samples, or a set of both luminance samples and a set of two color difference samples. A set of samples can also be described as a matrix of samples.
[0064] (5-1) I-Picture An I-picture is a picture encoded using only intra-prediction. I-pictures can be decoded independently. I-pictures are also called intra-pictures, intra-frames, keypictures, or keyframes.
[0065] (5-2) P-Picture A P-Picture is a picture encoded using single prediction. In other words, only one picture is referenced to encode the P-Picture. Single prediction is also called unidirectional prediction.
[0066] (5-3) B-Picture A B-picture is a picture encoded using bidirectional prediction. Bidirectional prediction is also expressed as dual prediction and means a prediction that references multiple pictures. Bidirectional prediction may refer to multiple pictures in different directions from the picture being processed (for example, a forward picture and a backward picture in different temporal directions), or it may refer to pictures in the same direction from the picture being processed.
[0067] Note that P-pictures and B-pictures are also called inter-pictures or inter-frames.
[0068] (6) A block is a processing unit of a set containing a specific number of pixels, and its name is not restricted, as shown in the following examples. It is also not restricted in shape, and includes not only rectangles made up of M x N pixels and squares made up of M x M pixels, but also other shapes.
[0069] (Examples of blocks) Slice / Tile / Brick CTB (Coding Tree Block) / CTU (Coding Tree Unit) / Superblock / Segment / Basic division unit / CU / Processing block unit / Prediction block unit (PU) / Translation block unit (TU) / Unit / Subblock / VPDU / Hardware processing division unit
[0070] A CTB is a processing unit into which components are divided. A CTB contains samples. A CTB is also called a CTU, superblock, or basic division unit. A CTB may be, for example, an N x N square block.
[0071] A Coding Block (CB) is a processing unit obtained by dividing a Coding Block (CTB). For example, it may be an M x N rectangular block. A CB is also called a Coding Unit (CU).
[0072] (7) A pixel / sample is the smallest unit of a point that constitutes an image, and includes not only pixels at integer positions but also pixels at decimal positions that are generated based on pixels at integer positions. A pixel / sample is a fundamental element that constitutes a picture.
[0073] (8) Pixel value / sample value: An intrinsic value of a pixel, which includes not only luminance value, chrominance value, and RGB gradation, but also depth value, or binary values of 0 or 1.
[0074] (9) Flags A flag is a variable or a single-bit syntax element. For example, a flag can only take one of two values: 0 or 1.
[0075] (10) A symbol or code used to transmit signal information, which includes not only discretized digital signals but also analog signals that take continuous values.
[0076] (11) Stream / Bitstream: Refers to a sequence of digital data or a flow of digital data. A stream / bitstream may consist of a single stream, or it may be divided into multiple layers and composed of multiple streams. It also includes cases where data is transmitted via serial communication over a single transmission path, as well as cases where data is transmitted via packet communication over multiple transmission paths.
[0077] (12) In the case of difference / difference scalar quantities, it is sufficient that the difference operation is included in addition to the simple difference (x - y), and this includes the absolute value of the difference (|x - y|), the squared difference (x^2 - y^2), the square root of the difference (√(x - y)), the weighted difference (ax - by: a, b is a constant), and the offset difference (x - y + a: a is the offset). Note that the difference operation also includes cases where bit shifts are performed (x >> 1 - y >> 1).
[0078] (13) In the case of summation scalar quantities, it is sufficient that the sum operation is included in addition to the simple sum (x + y), and this includes the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), weighted sum (ax + by: a, b is a constant), and offset sum (x + y + a: a is the offset). Note that the sum operation also includes cases where bit shifts are performed (x >> 1 + y >> 1).
[0079] (14) Based on: This includes cases where factors other than the subject being based on are taken into consideration. It also includes cases where the result is obtained not only by obtaining a direct result, but also by obtaining an intermediate result.
[0080] (15) Using: This includes cases where elements other than the target of use are taken into consideration. It also includes cases where the result is obtained not only by obtaining the result directly, but also by obtaining the result via an intermediate result.
[0081] (16) Prohibit, forbid. This can be rephrased as not being allowed. Also, not prohibiting or allowing something does not necessarily mean it is an obligation.
[0082] (17) To restrict (limit, restriction / restrict / restricted) This can be rephrased as not being allowed. Also, not being prohibited or being permitted does not necessarily mean being obligated. Furthermore, it is sufficient if it is prohibited in part quantitatively or qualitatively, and it also includes cases where it is prohibited entirely.
[0083] (18) MV (motion vector) A two-dimensional vector used in interpretation, which is a value that indicates the offset from the coordinates of the decoded image to the coordinates of the reference image.
[0084] (19) Data assigned to a specific column or row in a database, such as an index table or list. For example, a reference index is an index for a reference picture list, where one reference picture is assigned to the reference picture index.
[0085] [Explanation of Descriptions] In the drawings, the same reference number indicates the same or similar component. Also, the size and relative position of components in the drawings are not necessarily depicted to a constant scale.
[0086] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all general or specific examples. The numerical values, shapes, materials, components, arrangement and connection forms of components, steps, relationships and sequences of steps shown in the following embodiments are examples only and are not intended to limit the scope of the claims.
[0087] The following describes embodiments of encoding and decoding devices. These embodiments are examples of encoding and decoding devices to which the processes and / or configurations described in each aspect of this disclosure can be applied. The processes and / or configurations can also be implemented in encoding and decoding devices different from those in the embodiments. For example, with respect to the processes and / or configurations applicable to the embodiments, one of the following may be implemented:
[0088] (1) Any of the multiple components of the encoding or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.
[0089] (2) In the encoding or decoding device of the embodiment, any changes such as addition, replacement, or deletion of functions or processes performed by some of the multiple components of the encoding or decoding device may be made. For example, any of the functions or processes may be replaced or combined with other functions or processes described in any of the embodiments of this disclosure.
[0090] (3) In the methods performed by the encoding or decoding apparatus of the embodiment, any modifications such as additions, replacements, and deletions may be made to some of the processes included in the method. For example, any of the processes in the method may be replaced with or combined with other processes described in any of the embodiments of this disclosure.
[0091] (4) Some of the multiple components constituting the encoding or decoding device of the embodiment may be combined with components described in any of the embodiments of this disclosure, or with components that have some of the functions described in any of the embodiments of this disclosure, or with components that perform some of the processing performed by the components described in any of the embodiments of this disclosure.
[0092] (5) Components that provide some of the functions of the encoding or decoding device of the embodiment, or components that perform some of the processing of the encoding or decoding device of the embodiment, may be combined with or replaced with components described in any of the embodiments of this disclosure, components that provide some of the functions described in any of the embodiments of this disclosure, or components that perform some of the processing described in any of the embodiments of this disclosure.
[0093] (6) In a method performed by an encoding or decoding device of an embodiment, any of the processes included in the method may be replaced or combined with a process described in any of the embodiments of the present disclosure, or any of the similar processes.
[0094] (7) Some of the processes included in the methods performed by the encoding or decoding device of the embodiment may be combined with the processes described in any of the embodiments of this disclosure.
[0095] (8) The methods of carrying out the processes and / or configurations described in each aspect of the present disclosure are not limited to the encoding or decoding devices of the embodiments. For example, the processes and / or configurations may be carried out in devices used for purposes other than the video encoding or video decoding disclosed in the embodiments.
[0096] [System Configuration] Figure 1 is a schematic diagram showing an example of the configuration of the transmission system according to this embodiment.
[0097] The transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, an encoding device 100, a network Nw, and a decoding device 200, as shown in Figure 1.
[0098] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.
[0099] The original image input to the encoding device 100 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. Furthermore, the image is a higher-level concept than sequences, pictures, and blocks, and is not limited by spatial and temporal domains unless otherwise specified. The image consists of a sequence of pixels or pixel values, and the signal or pixel values representing the image are also called samples.
[0100] Furthermore, the stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal. In addition, the encoding device 100 may be called an image encoding device or a video encoding device, and the encoding method by the encoding device 100 may be called an encoding method, an image encoding method, or a video encoding method.
[0101] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting.
[0102] Furthermore, the network Nw may be replaced by a storage medium that records streams such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).
[0103] The decoding device 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted by the network Nw. For example, the decoding device decodes the stream according to a decoding method that corresponds to the encoding method by the encoding device 100.
[0104] The decoding device 200 may also be called an image decoding device or a video decoding device, and the decoding method performed by the decoding device 200 may be called a decoding method, an image decoding method, or a video decoding method.
[0105] [Encoding device] Next, the encoding device 100 according to the embodiment will be described.
[0106] [Implementation Example of Encoding Device] Figure 2 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, the multiple components of the encoding device 100 shown in Figure 3, which will be described later, are implemented by the processor a1 and memory a2 shown in Figure 2.
[0107] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit for encoding images. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits. Furthermore, for example, processor a1 may play the role of multiple components of the encoding device 100 shown in Figure 3, which will be described later, excluding the component for storing information.
[0108] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to encode an image is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.
[0109] For example, memory a2 may store the image to be encoded, or it may store a stream corresponding to the encoded image. Alternatively, memory a2 may store a program for processor a1 to encode the image.
[0110] Furthermore, for example, memory a2 may play the role of an information storage component among the multiple components of the encoding device 100 shown in Figure 3, which will be described later. Specifically, memory a2 may play the role of the block memory 118 and frame memory 122 shown in Figure 3, which will be described later. More specifically, memory a2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).
[0111] The generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be encoded. Whether or not to encode them is determined appropriately based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.
[0112] Furthermore, in the encoding device 100, not all of the multiple components shown in Figure 3 described later are to be implemented, nor are all of the multiple processes described above to be performed. Some of the multiple components shown in Figure 3 described later may be included in other devices, and some of the multiple processes described later may be performed by other devices.
[0113] The following is a general description of the encoding device 100, followed by a description of the components included in the encoding device 100.
[0114] [Example of Encoding Device Configuration] Figure 3 is a block diagram showing an example of the configuration of an encoding device 100 according to an embodiment. The encoding device 100 encodes images in block units.
[0115] As shown in Figure 3, the encoding device 100 is a device that encodes an image in block units and comprises a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, and a prediction unit. The prediction unit comprises an intra-prediction unit 124, an inter-prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130.
[0116] These components are implemented, for example, by the processor a1 and memory a2 of the encoding device 100 shown in Figure 2 above.
[0117] [Overall Encoding Process Flow] Figure 4A is a flowchart showing an example of the overall encoding process by the encoding device 100. Figure 4B is a flowchart showing an example of the block-by-block processing by the encoding device 100.
[0118] First, the division unit 102 of the encoding device 100 divides each picture contained in the original image into multiple blocks. For example, the division unit 102 first divides the picture into blocks of a fixed size (for example, 128 x 128 pixels) (step Sa_1a). These fixed-size blocks are sometimes called coding tree units (CTUs).
[0119] Then, the division unit 102 selects a division pattern for the fixed-size block and further divides the fixed-size block into multiple blocks to constitute the selected division pattern (step Sa_2a).
[0120] The further divided blocks are of a variable size, for example, 64 x 64 pixels or less. The vertical and horizontal pixel counts of the variable size can be any combination of 4, 8, 16, 32, or 64. These variable-sized blocks are sometimes called coding units (CUs), prediction units (PUs), or transformation units (TUs).
[0121] In various implementation examples, CU, PU, and TU do not need to be distinguished, and some or all of the blocks within a picture may be processing units for CU, PU, or TU.
[0122] Then, the encoding device 100 performs the processing shown in steps Sa_3 to Sa_11 in Figure 4B for each of the multiple blocks (step Sa_3a). Here, an example of a CTU size of 128 x 128 pixels is given, but other sizes are also possible. For example, the size of the CTU may be 256 x 256 pixels, or a larger size such as 512 x 512 pixels. In this case, the vertical and horizontal size of the divided blocks, i.e., the number of pixels, may be greater than 64 pixels, such as 128 pixels or 256 pixels.
[0123] Furthermore, the division unit 102 outputs parameters indicating the division pattern to the conversion unit 106, the inverse conversion unit 114, the intra-prediction unit 124, the inter-prediction unit 126, and the entropy coding unit 110. The conversion unit 106 may convert the prediction residuals based on these parameters, and the intra-prediction unit 124 and the inter-prediction unit 126 may generate a prediction image based on these parameters. The entropy coding unit 110 may also perform entropy coding on these parameters.
[0124] Then, the encoding device 100 processes each of the multiple blocks (step Sa_3a). Specifically, the encoding device 100 processes steps Sa_3 to Sa_11 as shown in Figure 4B. Then, the encoding device 100 determines whether or not the encoding of the entire picture is complete (step Sa_4a), and if it determines that it is not complete (No. in step Sa_4a), it repeats the processing from step Sa_2a.
[0125] Next, the processing for each block will be explained (Sa_3 to Sa_11 in Figure 4B). First, the prediction unit generates a predicted image of the current block (step Sa_3). Specifically, the prediction unit generates a predicted image of the current block by referring to a reconstructed image generated by encoding and then decoding other blocks. The predicted image is also called a predicted signal, predicted block, or predicted sample.
[0126] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.
[0127] The predicted image is, for example, an intra-prediction image (intra-prediction signal) generated based on intra-prediction, or an inter-prediction image (inter-prediction signal) generated based on inter-prediction.
[0128] Intra prediction is a method of predicting the current block by referring to blocks within the current picture, and is also called in-screen prediction. Specifically, the intra prediction unit 124 performs intra prediction by referring to the pixel values (e.g., luminance values or chrominance values, etc.) of blocks adjacent to the current block that are included in the current picture stored in the block memory 118. As a result, the intra prediction unit 124 generates an intra prediction image and outputs the intra prediction image to the prediction control unit 128.
[0129] Inter-prediction is a method of predicting the current block by referring to a reference picture different from the current picture, and is also called inter-screen prediction. Specifically, the inter-prediction unit 126 generates an inter-predicted image by referring to a reference picture stored in the frame memory 122 and performing inter-prediction of the current block, and outputs the inter-predicted image to the prediction control unit 128.
[0130] Note that prediction processing using the reconstructed image may not be performed on some blocks, such as the first block of the first image to be encoded. In that case, the original image will be output in the next subtraction process without performing subtraction on the original image. The interpretation unit 126 may also generate a predicted image by performing prediction processing without using a reference image.
[0131] Next, the subtraction unit 104 subtracts the predicted image (the predicted image input from the prediction control unit 128) from the original image in block units that are input from the division unit 102 and divided by the division unit 102. In other words, the subtraction unit 104 generates the difference between the current block and the predicted image as the predicted residual (step Sa_4).
[0132] The prediction residual is also called the prediction error. The original image is the input signal to the encoding device 100, and is, for example, a signal representing the image of each picture that makes up the video (e.g., a luminance (luma) signal and two chroma (chroma) signals). If the prediction process is skipped, the prediction residual becomes the value of the original image. For example, the prediction process is skipped for the first block in the processing order.
[0133] Next, the conversion unit 106 applies a conversion process to the predicted residual to generate conversion coefficients and outputs the conversion coefficients to the quantization unit 108 (step Sa_5).
[0134] The transformation process performed by the transformation unit 106 is, for example, an orthogonal transformation such as a discrete cosine transform (DCT) or discrete sine transform (DST) that transforms the predicted residual in the spatial domain into transformation coefficients in the frequency domain. However, other transformation processes such as weblet transforms, non-orthogonal transforms, or transformation processes expressed by matrix operations such as NSST (non-separable secondary transform) as defined in VVC may also be used.
[0135] The conversion unit 106 may perform a conversion process selected from among a plurality of candidate conversion processes. The conversion unit 106 may output conversion coefficients generated by sequentially applying a plurality of conversion processes to the predicted residual, or it may output the predicted residual as is without performing any conversion processes.
[0136] Depending on the processing performed by the conversion unit 106, information indicating whether or not a conversion process is applied to the predicted residuals, and / or information indicating the conversion type, may be encoded in the bitstream. This information is, for example, signaled at the CU level, but is not limited to the CU level and may be encoded at other levels (e.g., sequence level, picture level, slice level, brick level, or CTU level).
[0137] The quantization unit 108 quantizes the conversion coefficients output from the conversion unit 106 (step Sa_6). Specifically, the quantization unit 108 quantizes the conversion coefficients based on the quantization parameter (QP) corresponding to the conversion coefficients. The quantization unit 108 then outputs a plurality of quantized conversion coefficients (hereinafter referred to as quantization coefficients) of the current block to the entropy coding unit 110 and the inverse quantization unit 112. In other words, the quantization unit 108 generates quantization coefficients and outputs them to the entropy coding unit 110 and the inverse quantization unit 112.
[0138] The quantization parameter (QP) is a parameter that defines the quantization step (quantization width). For example, if the value of the quantization parameter increases, the quantization step also increases. In other words, if the value of the quantization parameter increases, the error in the quantization coefficient (quantization error) increases.
[0139] Furthermore, the quantization unit 108 may quantize the conversion coefficients based on a quantization matrix. In other words, quantization parameters (QP) and / or a quantization matrix may be used for quantization. The quantization parameters (QP) and the quantization matrix may be encoded, for example, at the sequence level, picture level, slice level, brick level, or CTU level.
[0140] The quantization unit 108 performs quantization processing in a predetermined scanning order. This predetermined scanning order is the order for quantization / inverse quantization of the conversion coefficients. For example, the predetermined scanning order is defined as ascending order of frequency (from low frequency to high frequency) or descending order (from high frequency to low frequency).
[0141] Next, the entropy coding unit 110 generates a stream (step Sa_7) by coding (specifically, entropy coding) the plurality of quantization coefficients and the prediction parameters related to the generation of the predicted image. For this entropy coding, for example, CABAC (Context-based Adaptive Binary Arithmetic Coding), which is used in VVC, may be used. Alternatively, a coding method that is a modified version of CABAC used in VVC may be used.
[0142] Furthermore, the processing of the entropy encoding unit 110 does not necessarily have to be performed after the quantization process; for example, it may be performed outside the loop of processing for each block.
[0143] Next, the inverse quantization unit 112 and the inverse transformation unit 114 restore the predicted residuals by performing inverse quantization and inverse transformation on a plurality of quantization coefficients (steps Sa_8 and Sa_9).
[0144] Next, the summing unit 116 reconstructs the current block by adding the predicted image to the recovered predicted residual (step Sa_10). This generates a reconstructed image. The reconstructed image is also called a reconstructed block, and the reconstructed image generated by the encoding device 100 is also called a local decoded block or local decoded image.
[0145] Next, the loop filter unit 120 performs filtering on the reconstructed image as needed (step Sa_11). Specifically, the loop filter unit 120 applies loop filtering to the reconstructed image output from the adder unit 116 and outputs the filtered reconstructed image to the frame memory 122.
[0146] Loop filters are filters used to reduce block noise that occurs at block boundaries, and include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), and sample adaptive offset (SAO). Here, loop filters are described as in-loop filters used within the coding loop, but loop filters may also be out-loop filters used outside the coding loop.
[0147] In the example described above, the encoding device 100 selects one division pattern for a fixed-size block and encodes each block according to that division pattern. However, it may also encode each block according to multiple division patterns. In this case, the encoding device 100 may evaluate the cost of each of the multiple division patterns and, for example, select the stream obtained by encoding according to the division pattern with the smallest cost as the final output stream.
[0148] Furthermore, the processes in steps Sa_1a to Sa_4a and Sa_3 to Sa_11 may be performed sequentially by the encoding device 100, some of these processes may be performed in parallel, and the order may be changed.
[0149] The encoding process performed by such an encoding device 100 is a hybrid encoding using predictive encoding and transformative encoding. Furthermore, predictive encoding is performed by an encoding loop consisting of a subtraction unit 104, a transformer unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transformer unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra-prediction unit 124, an inter-prediction unit 126, and a prediction control unit 128. In other words, the prediction processing unit consisting of the intra-prediction unit 124 and the inter-prediction unit 126 constitutes a part of the encoding loop.
[0150] [Decoding Device] Next, a decoding device 200 capable of decoding the stream output from the encoding device 100 will be described.
[0151] [Implementation Example of Decryption Device] Figure 5 is a block diagram showing an implementation example of the decoding device 200. The decoding device 200 includes a processor b1 and memory b2. For example, the multiple components of the decoding device 200 shown in Figure 6, which will be described later, are implemented by the processor b1 and memory b2 shown in Figure 5.
[0152] Processor b1 is a circuit that performs information processing and is a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit for decoding streams. Processor b1 may be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits. Furthermore, for example, processor b1 may play the role of multiple components of the decoding device 200 shown in Figure 6, etc., described later, excluding the component for storing information.
[0153] Memory b2 is a dedicated or general-purpose memory in which information for the processor b1 to decode the stream is stored. Memory b2 may be an electronic circuit and may be connected to the processor b1. Memory b2 may also be included in the processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.
[0154] For example, memory b2 may store an image or a stream. Alternatively, memory b2 may store a program for processor b1 to decode the stream.
[0155] Furthermore, for example, memory b2 may play the role of an information storage component among the multiple components of the decoding device 200 shown in Figure 6 below. Specifically, memory b2 may play the role of a block memory 210 and a frame memory 214 shown in Figure 6 below. More specifically, memory b2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).
[0156] The generated parameters disclosed in the embodiments may be stored in memory b2 and referenced during processing by processor b1. The parameters may be generated by a decoding process or not.
[0157] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 6, etc., described later, nor does it need to perform all of the processes described above. Some of the components shown in Figure 6, etc., described later, may be included in other devices, and some of the processes described above may be performed by other devices.
[0158] The following is a general description of the decoding device 200, followed by a description of the components included in the decoding device 200. Note that detailed descriptions of some of the components included in the decoding device 200 that perform the same processing as the components included in the encoding device 100 may be omitted.
[0159] [Example of Decoding Device Configuration] Figure 6 is a block diagram showing an example of the configuration of a decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes a stream, which is an encoded image, in block units.
[0160] As shown in Figure 6, the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transformation unit 206, an addition unit 208, a block memory 210, a loop filter unit 212, a frame memory 214, a prediction unit, a prediction control unit 220, a prediction parameter generation unit 222, and a division determination unit 224. The prediction unit includes an intra-prediction unit 216 and an inter-prediction unit 218.
[0161] These components are implemented, for example, by the processor b1 and memory b2 of the decoding device 200 shown in Figure 5 above.
[0162] Furthermore, components included in the decoding device 200 may perform the same processing as components included in the encoding device 100, and their explanation may be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, adder unit 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, adder unit 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120, respectively.
[0163] [Overall Decryption Process Flow] Figure 7A is a flowchart showing an example of the overall decryption process by the decryption device 200. Figure 7B is a flowchart showing an example of the block-by-block processing by the decryption device 200.
[0164] First, the entropy decoding unit 202 of the decoding device 200 acquires the stream output from the encoding device 100. Then, the entropy decoding unit 202 decodes (specifically, entropy decodes) the encoded quantization coefficients and prediction parameters of the current block contained in the bitstream (step Sp_1a).
[0165] The division determination unit 224 determines the division pattern for each of the multiple fixed-size blocks (128 x 128 pixels) contained in the picture based on the parameters input from the entropy decoding unit 202 (step Sp_2a). This division pattern is the division pattern selected by the encoding device 100. The decoding device 200 then performs block-specific processing for each of the multiple blocks. Specifically, it performs the processing in steps Sp_1 to Sp_5 for each of the multiple blocks (step Sp_3a).
[0166] The decoding device 200 then determines whether or not the decoding of the entire picture is complete (step Sp_4a). If it determines that it is not complete (No. in step Sp_4a), it repeats the process from step Sp_2a.
[0167] Next, we will explain the processing for each block (Sp_1 to Sp_5 in Figure 7B).
[0168] First, the inverse quantization unit 204 inversely quantizes the quantization coefficients of the current block, which are input from the entropy decoding unit 202. Specifically, for each of the quantization coefficients of the current block, the inverse quantization unit 204 inversely quantizes the quantization coefficient based on the quantization parameter corresponding to that quantization coefficient (step Sp_1).
[0169] For example, the inverse quantization unit 204 may acquire a parameter indicating whether or not to perform inverse quantization, and quantization parameters (such as difference quantization parameters and QP index). The inverse quantization unit 204 may then decide whether or not to perform inverse quantization based on the acquired parameters, and may perform the inverse quantization process if it is decided to perform inverse quantization.
[0170] The inverse quantization unit 204 then outputs the inversely quantized quantization coefficients (i.e., transformation coefficients) of the current block to the inverse transformation unit 206.
[0171] Next, the inverse transform unit 206 restores the predicted residual by inversely transforming the transformation coefficients, which are input from the inverse quantization unit 204 (step Sp_2). Here, inverse transform refers to the reverse transformation process of the transformation process described in the encoding device 100. In other words, the inverse transform unit 206 can also be said to be performing a transformation process.
[0172] Next, the prediction unit, consisting of the intra-prediction unit 216, the inter-prediction unit 218, and the prediction control unit 220, generates a predicted image of the current block (step Sp_3). Here, the intra-prediction unit 216 and the inter-prediction unit 218 may perform the same processing as described above on the encoding device 100. Note that the predicted image needs to be generated before the addition process because it is used in the subsequent addition process, but it does not necessarily have to be done after the prediction residual restoration process. For example, the prediction image generation process may be performed in parallel with the prediction residual restoration process.
[0173] Next, the summing unit 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted image to the predicted residual (step Sp_4). In other words, the summing unit 208 generates a reconstructed image of the current block by adding the predicted image to the predicted residual. The summing unit 208 then outputs the reconstructed image of the current block to the block memory 210 and the loop filter unit 212.
[0174] The block memory 210 is a block referenced in intra prediction and is a storage unit for storing blocks within the current picture. Specifically, the block memory 210 stores the reconstructed image output from the adder 208.
[0175] The loop filter unit 212 performs filtering on the reconstructed image (step Sp_5). The filtered reconstructed image is output to the frame memory 214 and the display device, etc.
[0176] In Figures 6 and 7B, the loop filter unit 212 processes the reconstructed image within the loop; that is, it outputs the filtered reconstructed image to the frame memory 214. Some or all of the processing in the loop filter unit 212 may be performed outside the loop. That is, filtering may be performed before outputting to a display device or the like.
[0177] Furthermore, the processes in steps Sp_1a to Sp_4a and Sp_1 to Sp_5 may be performed sequentially by the decoding device 200, some of these processes may be performed in parallel, or the order of these processes may be changed.
[0178] [Another Example of System Configuration] Figure 8 is a schematic diagram showing another example of the configuration of the transmission system according to this embodiment. The transmission system Trs is a system that transmits generated streams and receives transmitted streams. Such a transmission system Trs includes, for example, a transmitting device 2000, a network Nw, and a receiving device 3000, as shown in Figure 8. Note that the transmission system Trs does not need to have all of these components and may consist of only some of the devices.
[0179] The network Nw may be the Internet, a wide-area network (WAN), a local area network (LAN), or a combination thereof. The network Nw is not necessarily limited to a bidirectional communication network; it may also be a unidirectional communication network transmitting broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Furthermore, the network Nw may be replaced by a storage medium that records streams, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc®).
[0180] The transmitting device 2000 transmits the stream over the network Nw. The transmitting device 2000 may also take the encoded stream as input and output the acquired stream. Alternatively, the transmitting device 2000 may have the encoding processing configuration disclosed as the encoding device 100, take the original image before encoding as input, encode the acquired image to generate a stream, and transmit it.
[0181] When the transmitting device 2000 generates a stream, the original input image before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image. The stream includes, for example, an encoded image and control information for decoding that encoded image. This encoding compresses the image.
[0182] The transmitting device 2000 can reduce the load on stream transmission by reducing the amount of data in the stream, which includes the encoded image.
[0183] The receiving device 3000 receives an encoded stream from the network Nw. The receiving device 3000 may store the received stream in memory, transfer the received stream to another device, or perform any processing on the received stream. For example, the receiving device 3000 may have a decoding process configuration disclosed as the decoding device 200 and decode the received stream.
[0184] The receiving device 3000 can reduce the delay of the stream by reducing the amount of data in the received stream.
[0185] [Transmitting Device] Figure 9 is a schematic diagram showing an example configuration of the transmitting device 2000 according to this embodiment. The transmitting device 2000 is a device that transmits a stream.
[0186] [Implementation Example of Transmitting Device] Next, a transmitting device 2000 according to the embodiment will be described. Figure 10 is a block diagram showing an implementation example of the transmitting device 2000. The transmitting device 2000 includes a processor a1 and a transmitting unit a3 that outputs a stream. The transmitting device 2000 may also further include a memory a2 as shown in Figure 2, or a receiving unit that acquires a video or stream.
[0187] Processor a1 is a circuit that performs information processing. Processor a1 may also be a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that transmits a stream. Processor a1 may also be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.
[0188] Memory a2 is a dedicated or general-purpose memory in which information for the processor a1 to transmit a stream is stored. Memory a2 may be an electronic circuit and may be connected to the processor a1. Memory a2 may also be included in the processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.
[0189] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to send the stream.
[0190] The stream transmitted by the transmitting device 2000 may be the stream (also called a bitstream) described in this disclosure.
[0191] Furthermore, in the transmitting device 2000, the processor a1 may perform some or all of the roles of the transmitting unit a3. Alternatively, the transmitting device 2000 may not have a transmitting unit a3, and the processor a1 may output the stream.
[0192] [Receiving device] Figure 11 is a schematic diagram showing an example of the configuration of a receiving device 3000 according to this embodiment. The receiving device 3000 is a device that receives a stream.
[0193] [Implementation Example of Receiving Device] Next, a receiving device 3000 according to the embodiment will be described. Figure 12 is a block diagram showing an implementation example of the receiving device 3000. The receiving device 3000 includes a processor b1 and a receiving unit b4 that receives a stream. The receiving device 3000 may also further include a memory b2 as shown in Figure 5, or a transmitting unit that transmits a video or stream.
[0194] Processor b1 is a circuit that performs information processing. Processor b1 may also be a circuit that can access memory b2. For example, processor b1 is a dedicated or general-purpose electronic circuit that receives a stream. Processor b1 may also be a processor such as a CPU. Alternatively, processor b1 may be a collection of multiple electronic circuits.
[0195] Memory b2 is a dedicated or general-purpose memory that stores information for processor b1 to receive streams. Memory b2 may be an electronic circuit and may be connected to processor b1. Memory b2 may also be included in processor b1. Memory b2 may also be a collection of multiple electronic circuits. Memory b2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory b2 may also be a non-volatile memory or a volatile memory.
[0196] For example, memory b2 may store the input stream or the decoded image. Alternatively, memory b2 may store a program for processor b1 to receive the stream.
[0197] The stream received by the receiving device 3000 may be the stream (also called a bitstream) described in this disclosure.
[0198] Furthermore, in the receiving device 3000, the processor b1 may perform some or all of the functions of the receiving unit b4. Alternatively, the receiving device 3000 may not have a receiving unit b4, and the processor b1 may receive the stream.
[0199] [Bitstream Generation Device] Figure 13 is a schematic diagram showing an example of the configuration of a bitstream generation device according to this embodiment. The bitstream generation device 1000 is a device that generates a bitstream and transmits the generated stream.
[0200] An image is input to the bitstream generator 1000. The bitstream generator 1000 generates a stream by encoding the input image and outputs the stream. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding process.
[0201] The original image input to the bitstream generator 1000 before encoding is also called the original image, original signal, or original sample. The image may be a moving image or a still image.
[0202] Furthermore, an image is a broader concept than sequences, pictures, and blocks, and is not limited to spatial or temporal domains unless otherwise specified. An image consists of a sequence of pixels or pixel values, and the signal or pixel values representing that image are also called samples. A stream may also be called a bitstream, encoded bitstream, compressed bitstream, or encoded signal.
[0203] Furthermore, the bitstream generation device may also be called the encoding device, image encoding device, or video encoding device, and the method for generating a bitstream by the bitstream generation device may be called the encoding method, image encoding method, or video encoding method.
[0204] [Implementation Example of Bitstream Generation Device] Next, a bitstream generation device 1000 according to an embodiment will be described. Figure 14 is a block diagram showing an implementation example of the bitstream generation device 1000. The bitstream generation device 1000 includes a processor a1 and a memory a2. The bitstream generation device 1000 may also include an input unit for inputting moving images, and an output unit for outputting the generated bitstream.
[0205] Processor a1 is a circuit that performs information processing and is a circuit that can access memory a2. For example, processor a1 is a dedicated or general-purpose electronic circuit that generates a bitstream. Processor a1 may be a processor such as a CPU. Alternatively, processor a1 may be a collection of multiple electronic circuits.
[0206] Memory a2 is a dedicated or general-purpose memory in which information for processor a1 to generate a bitstream is stored. Memory a2 may be an electronic circuit and may be connected to processor a1. Memory a2 may also be included in processor a1. Memory a2 may also be a collection of multiple electronic circuits. Memory a2 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory a2 may also be a non-volatile memory or a volatile memory.
[0207] For example, memory a2 may store the input image or the generated stream. Memory a2 may also store a program for processor a1 to generate the bitstream.
[0208] The bitstream generator 1000 may have the same configuration as the encoding device 100 in this disclosure. Specifically, the bitstream generator 1000 may implement the multiple components shown in Figure 3. However, not all of the multiple components shown in Figure 3 are to be implemented, nor are all of the multiple processes to be performed. Some of the multiple components shown in Figure 3 may be included in other devices, and some of the multiple processes described above may be performed by other devices.
[0209] Furthermore, for example, processor a1 may perform the role of multiple components of the encoding device 100 shown in Figure 3, excluding the component for storing information. Also, for example, memory a2 may perform the role of the component for storing information among the multiple components of the encoding device 100 shown in Figure 3.
[0210] Specifically, memory a2 may function as the block memory 118 and frame memory 122 shown in Figure 3. More specifically, memory a2 may store reconstructed images (specifically, reconstructed blocks or reconstructed pictures, etc.).
[0211] Furthermore, during operation, processor a1 uses memory a2 to generate a bitstream that includes parameters for causing the decoding device 200 to execute processing, and / or parameters for switching the processing of the decoding device 200 or other specific processing. For example, processor a1 generates these parameters, includes the generated parameters in the bitstream, and generates the bitstream.
[0212] Here, the "parameters" to be executed by the decoding device 200 may be parameters for executing any decoding process exemplified in this disclosure as a process to be executed. For example, the "parameters" to be executed by the decoding device 200 may be any parameter in the syntax described in this disclosure. Furthermore, the "process of the decoding device 200" or "other specific process" that can be switched may be any decoding process exemplified in this disclosure as a process performed by obtaining the syntax.
[0213] Furthermore, the generated parameters disclosed in the embodiments may be stored in memory a2 and referenced in the processing of processor a1. The generated parameters may be stored in memory a2 and may or may not be included in the bitstream. Whether or not to include them in the bitstream is appropriately determined based on the relationship between the increase in the amount of code due to encoding the parameters and the reduction in the processing load of the decoding device 200 due to receiving the parameters.
[0214] [Storage Medium] Figure 15A is a schematic diagram showing an example of the configuration of a storage medium and a computer according to this embodiment. The storage medium 4000 is a medium for storing bitstreams. The storage medium 4000 is connected to a processor b1 included in the computer 5000 and outputs a bitstream, which is an executable instruction for the computer 5000, to the computer 5000. For example, when the computer 5000 reads the bitstream from the storage medium 4000, the storage medium 4000 outputs the bitstream to the computer 5000.
[0215] The encoded stream stored in the storage medium 4000 is an instruction that the computer 5000 can execute. The processor b1 in the computer 5000 generates a reconstructed image by executing a bitstream decoding process based on the acquired bitstream. In this way, the computer 5000, having acquired a stream from the storage medium 4000, can decode the bitstream based on the parameters and other information contained in the encoded stream.
[0216] Furthermore, the storage medium 4000 may be the memory b2 described herein. For example, the memory b2 may store a program for the processor b1 to generate a bitstream. The storage medium 4000 may be a non-transitory computer-readable medium. The storage medium 4000 is also referred to as a recording medium.
[0217] The storage medium 4000 may include an input / output unit for inputting and outputting bitstreams. The storage medium 4000 may also include an input unit for receiving bitstreams and an output unit for outputting bitstreams.
[0218] Furthermore, the computer 5000 may also be the decoding device 200. In other words, the computer 5000 may be read as the decoding device 200.
[0219] Figure 15B is a schematic diagram showing an example of the configuration of a computer according to this embodiment. The schematic diagram shown in Figure 15B differs from the example in Figure 15A in that the storage medium 4000 is included in the computer 5000. The storage medium 4000 and the processor b1 may be the same as those described in Figure 15A.
[0220] Here, the bitstream includes parameters that indicate instructions for the computer 5000 to execute. The "parameters" may include parameters that cause the computer 5000 to perform processing, and / or parameters that switch the processing or other specific processing of the computer 5000.
[0221] The “parameter” may be any parameter that performs any decryption process exemplified in this disclosure as an example of a process to be performed. For example, it may be any parameter in the syntax described in this disclosure. The process to be performed and the process to be switched may be any decryption process exemplified in this disclosure as a process performed by obtaining the syntax.
[0222] [Examples of Parameters] In the above, for the bitstream generator 1000, "parameters that cause the decoding device 200 to execute processing, and / or parameters that switch the processing of the decoding device 200 or other specific processing" are described. Also, for the storage medium 4000, "parameters that cause the computer 5000 to execute processing, and / or parameters that switch the processing of the computer 5000 or other specific processing" are described. These parameters may be, for example, the following parameters.
[0223] The "parameters" may include, for example, parameters indicating the method of dividing the block. The computer 5000 may divide the block based on the parameters indicating the method of dividing the block and perform any of the processes illustrated in this disclosure on the divided blocks.
[0224] The "parameters" may include, for example, parameters indicating the prediction mode to be applied. The computer 5000 may determine the prediction mode to be applied based on the parameters indicating the prediction mode to be applied, and may perform any prediction process exemplified in this disclosure corresponding to the determined prediction mode.
[0225] The "parameters" may include, for example, a parameter indicating whether or not process y can be performed within range x. The computer 5000 may determine whether or not process y can be performed on a target block based on the parameter indicating whether or not process y can be performed within range x.
[0226] A flag (parameter) indicating whether or not process y can be performed within range x is shown, for example, as xxx_yyy_enabled_flag / xxx_yyy_disabled_flag. Here, yyy may represent process y, which is indicated by the flag as either or not. Process y may also be any process exemplified in this disclosure.
[0227] xxx may represent the level to which parameters are assigned. For example, xxx may correspond to sps, pps, ph, sh, cu (Coding Unit), or block, and this level may indicate the range x. More specifically, if it is sps, the range x is a sequence; if it is pps, the range x is a picture; if it is ph, the range x is a picture; and if it is sh, the range x is a slice. Note that the levels to which parameters are assigned are not limited to these.
[0228] If the parameter indicates that process y is not feasible, then process y will not be performed in the blocks included in range x. If the parameter indicates that process y is feasible, then process y is feasible in the blocks included in range x. In other words, if the parameter indicates that process y is feasible, then process y may or may not be performed in the blocks included in range x.
[0229] The flag indicating whether or not process y can be performed may be defined for multiple headers. For example, process y may be permitted for a sequence, but not for a particular picture within that sequence.
[0230] In the above case, the sps_yyy_enabled_flag (or sps_yyy_disabled_flag) for the sequence may have a value indicating that process y can be performed. Furthermore, the ph_yyy_enabled_flag (or ph_yyy_disabled_flag) for a certain picture included in the sequence may have a value indicating that process y cannot be performed.
[0231] As described above, the levels correspond to sps, pps, ph, sh, cu (Coding Unit), or blocks, etc. An example definition for each level is shown below. Note that the definitions for each level can be changed.
[0232] The description of `sps` (Sequence Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more CLVS (Coded Layer Video Sequences). Here, the number of zero or more CLVSs is determined by the content of the syntax elements contained in the PPS referenced by the syntax elements contained in each picture header. In other words, `sps` corresponds to the sequence level.
[0233] The description of pps (Picture Parameter Set) corresponds to a syntax structure containing syntax elements that apply to zero or more encoded pictures. Here, zero or more encoded pictures are determined by the syntax elements contained in each picture header. In other words, pps corresponds to the picture level.
[0234] ph (PH: Picture Header) corresponds to a syntax structure containing syntax elements that apply to all slices of an encoded picture. In other words, ph corresponds to the picture level.
[0235] The entry for sh (SH: Slice Header) corresponds to the portion of the encoded slice that contains all tiles or data elements related to CTU rows within a slice. In other words, sh corresponds to the slice level.
[0236] The notation `cu` (Coding Unit) corresponds to the coding block of the sample and the syntax structure used to encode that sample. In other words, `cu` corresponds to the coding unit level or the block level.
[0237] Here, the above coding blocks correspond to (i) the coding block for the luminance sample and the two corresponding coding blocks for the chrominance sample of a picture having three sample sequences in single-tree mode, (ii) the coding block for the luminance sample of a picture having three sample sequences in dual-tree mode, (iii) the two coding blocks for the chrominance sample of a picture having three sample sequences in dual-tree mode, or (iv) the coding block for the sample of a monochrome picture.
[0238] Although a flag indicating whether or not process y can be performed has been described above, a flag indicating whether or not to perform process y (xxx_yyy_flag) may be used in a similar manner. Furthermore, the flag indicating whether or not process y can be performed and the flag indicating whether or not to perform process y may be used in combination.
[0239] [Data Structure] Figure 16 shows an example of the hierarchical structure of data in a stream. A stream includes, for example, a video sequence. This video sequence includes, for example, a VPS (Video Parameter Set), an SPS (Sequence Parameter Set), a PPS (Picture Parameter Set), an SEI (Supplemental Enhancement Information), and multiple pictures, as shown in Figure 16(a).
[0240] VPS includes encoding parameters common to multiple layers in a video composed of multiple layers, and encoding parameters related to the multiple layers included in the video, or to individual layers.
[0241] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of the picture. Multiple SPSs may exist.
[0242] The PPS includes parameters used for the picture, i.e., encoding parameters that the decoding device 200 references to decode each picture in the sequence. For example, the encoding parameters may include a reference value for the quantization width used to decode the picture and a flag indicating the application of weighted prediction. There may be multiple PPSs. Also, the SPS and PPS are sometimes simply referred to as parameter sets.
[0243] The picture may include a picture header and one or more slices, as shown in Figure 16(b). The picture header includes encoding parameters that the decoding device 200 refers to in order to decode the one or more slices.
[0244] A slice includes a slice header and one or more Coding Tree Units (CTUs), as shown in Figure 16(c). The slice header includes coding parameters that the decoding device 200 references to decode the one or more CTUs.
[0245] A picture may not contain slices, but instead contain tile groups. In this case, a tile group may contain one or more tiles, and each tile may contain one or more CTUs. Furthermore, a tile may contain one or more slices, and a slice may contain one or more tiles.
[0246] A CTU is also called a superblock or basic partitioning unit. Such a CTU includes a CTU header and one or more CUs (Coding Units), as shown in Figure 16(d). The CTU header contains coding parameters that the decoding device 200 refers to in order to decode one or more CUs.
[0247] A CU may be divided into multiple smaller CUs. Furthermore, as shown in Figure 16(e), a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating the predicted residual, which will be described later.
[0248] Note that a CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but in SBT, for example, as described later, it may include multiple TUs smaller than the CU. Also, a CU may be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes that CU. A VPDU is a fixed unit that can be processed in one stage when performing pipeline processing in hardware, for example.
[0249] Note that a stream does not necessarily have to have some of the hierarchical levels shown in Figure 16. Also, the order of these hierarchical levels may be changed, and any hierarchical level may be replaced by another hierarchical level.
[0250] Furthermore, the picture that is currently being processed by a device such as the encoding device 100 or the decoding device 200 is called the current picture. If the processing is encoding, the current picture is synonymous with the picture to be encoded; if the processing is decoding, the current picture is synonymous with the picture to be decoded.
[0251] Furthermore, a block, such as CU or CU, that is currently being processed by a device such as an encoding device 100 or a decoding device 200 is called the current block. If the processing is encoding, the current block is synonymous with the block to be encoded; if the processing is decoding, the current block is synonymous with the block to be decoded.
[0252] [Picture Composition: Slices / Tiles] To decode pictures in parallel, pictures may be composed of slices or tiles.
[0253] A slice is the basic encoding unit that makes up a picture. A picture is composed of, for example, one or more slices. A slice consists of one or more consecutive CTUs.
[0254] Figure 17 shows an example of a slice configuration. For example, a picture contains 11 x 8 CTUs and is divided into four slices (slices 1-4). Slice 1 consists of, for example, 16 CTUs, slice 2 consists of, for example, 21 CTUs, slice 3 consists of, for example, 29 CTUs, and slice 4 consists of, for example, 22 CTUs. Here, each CTU in the picture belongs to one of the slices.
[0255] The shape of a slice is a horizontal division of the picture. The boundaries of a slice do not have to be at the edges of the screen, but can be anywhere among the CTU boundaries within the screen. The processing order (encoding order or decoding order) of the CTUs within a slice is, for example, the raster scan order. A slice also includes a slice header and encoded data. The slice header may describe the characteristics of the slice, such as the CTU address of the beginning of the slice and the slice type.
[0256] A tile is a rectangular area that makes up a picture. Each tile may be assigned a number called a TileId in the order of the raster scan.
[0257] Figure 18 shows an example of a tile configuration. For example, a picture contains 11 x 8 CTUs and is divided into four rectangular tile regions (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used.
[0258] If tiles are not used, multiple CTUs within a picture are processed, for example, in raster scan order. If tiles are used, at least one CTU in each of the multiple tiles is processed, for example, in raster scan order. For example, as shown in Figure 18, the processing order of the multiple CTUs contained in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.
[0259] Note that one tile may contain one or more slices, and one slice may contain one or more tiles.
[0260] A picture may be composed of tilesets. A tileset may contain one or more tile groups, or one or more tiles. A picture may consist of only one of a tileset, a tile group, or a tile. For example, the order in which multiple tiles are scanned in raster order for each tileset is defined as the basic coding order of the tiles. Within each tileset, a collection of one or more tiles whose basic coding order is consecutive is defined as a tile group. Such a picture may be composed of the division unit 102 (see Figure 3) described later.
[0261] [Temporal Scalable Encoding] Figure 19 shows an example of a temporally scalable stream configuration.
[0262] The encoding device 100 may generate a temporally scalable stream by encoding multiple pictures in multiple layers, as shown in Figure 19. For example, the encoding device 100 achieves scalability by encoding each picture in its own layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called temporally scalable encoding.
[0263] This allows the decoding device 200 to switch the frame rate of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.
[0264] As a result, the decoding device 200 can freely switch between decoding the same content in low-frame-rate and high-frame-rate formats. For example, a user of the stream can watch part of the video using their smartphone while on the go, and then watch the rest of the video using an internet TV or other device after returning home.
[0265] Each of the aforementioned smartphones and devices incorporates a decoding device 200, which may or may not have the same performance. In this case, if the device decodes up to the upper layers of the stream, the user can view high-frame-rate video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different frame rates, thereby reducing the processing load.
[0266] In temporally scalable coding, layers are sometimes referred to as temporal layers or temporal sublayers.
[0267] [Multilayer Encoding] Figure 20 shows an example of a stream configuration using multilayer encoding. Each of AU0 to AU3 corresponds to an AU (Access Unit). Each of Layer0 to Layer3 corresponds to a layer in multilayer encoding. Each PU (Picture Unit) corresponds to a picture. Each of POC_1 to POC_3 corresponds to a POC (Picture Of Count) and indicates the display order.
[0268] As shown in Figure 20, the encoding device 100 may generate a stream in which spatial resolution, image quality, or multiplexed content can be scalably changed by encoding multiple pictures into multiple layers. For example, the encoding device 100 achieves scalability by encoding pictures layer by layer, with enhancement layers existing above the base layer. This type of encoding of each picture is called multi-layer encoding.
[0269] This allows the decoding device 200 to switch the spatial resolution, image quality, or multiplexed content of the image displayed by decoding the stream. In other words, the decoding device 200 decides which layers to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth.
[0270] As a result, the decoding device 200 can freely switch between decoding low-resolution and high-resolution content, low-quality and high-quality content, and basic and customized content for the same content. For example, a user of the stream can watch part of the video on their smartphone while on the go, and then watch the rest of the video on an internet TV or other device after returning home.
[0271] Furthermore, each of the aforementioned smartphones and devices incorporates a decoding device 200 with identical or different performance characteristics. In this case, if the device decodes up to the upper layers of the stream, the user can view high-definition video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different image quality, thereby reducing the processing load.
[0272] [Metadata] Each layer of the stream may include metadata based on statistical information of the image. The decoding device 200 may generate high-resolution video by super-resolution the pictures of each layer based on the metadata. Super-resolution may be either an improvement in the signal-to-noise ratio (SN) at the same resolution, or an increase in resolution. The metadata may include information for identifying linear or nonlinear filter coefficients used in the super-resolution process, or information for identifying parameter values in the filtering process, machine learning, or least-squares operation used in the super-resolution process.
[0273] Alternatively, the picture may be divided into tiles or the like, depending on the meaning of each object within the picture. In this case, the decoding device 200 may decode only a portion of the picture by selecting the tiles to be decoded.
[0274] Furthermore, object attributes (such as person, car, or ball) and their position within the picture (such as their coordinate position within the same picture) may be stored as metadata. In this case, the decoding device 200 can identify the position of a desired object based on the metadata and determine the tile containing that object. For example, the metadata is stored using a data storage structure different from pixel data, such as SEI in HEVC. This metadata may indicate, for example, the position, size, or color of the main object.
[0275] Furthermore, metadata may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units. This allows the decoding device 200 to obtain information such as the time when a specific person appears in the video, and by using that time and the information in the picture units, it can identify the picture in which the object exists and the position of the object within that picture.
[0276] [Dividing section] The dividing section 102 first divides the picture into blocks of a fixed size (e.g., 128 x 128 pixels). Then, the dividing section 102 divides each of the fixed-size blocks into blocks of a variable size (e.g., 64 x 64 pixels or less) based on, for example, a recursive quadtree and / or binary tree block partitioning.
[0277] Figure 21 shows an example of block partitioning in the embodiment. In Figure 21, solid lines represent block boundaries due to quadtree block partitioning, and dashed lines represent block boundaries due to binary tree block partitioning.
[0278] Here, block 10 is a 128x128 pixel square block. This block 10 is first divided into four 64x64 pixel square blocks (quadtree block partitioning).
[0279] The top-left 64x64 pixel square block is further divided vertically into two rectangular blocks, each consisting of 32x64 pixels, and the left 32x64 pixel rectangular block is further divided vertically into two rectangular blocks, each consisting of 16x64 pixels (binary tree block partitioning). As a result, the top-left 64x64 pixel square block is divided into two 16x64 pixel rectangular blocks 11 and 12 and a 32x64 pixel rectangular block 13.
[0280] The 64x64 pixel square block in the upper right is horizontally divided into two rectangular blocks 14 and 15, each consisting of 64x32 pixels (binary tree block partitioning).
[0281] The 64x64 pixel square block in the lower left is divided into four square blocks, each consisting of 32x32 pixels (quadtree block partitioning). Of these four 32x32 pixel square blocks, the upper left and lower right blocks are further divided.
[0282] The 32x32 pixel square block in the upper left is vertically divided into two rectangular blocks, each consisting of 16x32 pixels. The rectangular block on the right, each consisting of 16x32 pixels, is further horizontally divided into two 16x16 pixel square blocks (binary tree block partitioning). The 32x32 pixel square block in the lower right is horizontally divided into two rectangular blocks, each consisting of 32x16 pixels (binary tree block partitioning).
[0283] As a result, the 64x64 pixel square block in the lower left is divided into a 16x32 pixel rectangular block 16, two 16x16 pixel square blocks 17 and 18, two 32x32 pixel square blocks 19 and 20, and two 32x16 pixel rectangular blocks 21 and 22.
[0284] The 64x64 pixel block 23 in the lower right corner will not be divided.
[0285] As described above, in Figure 21, block 10 is divided into 13 variable-sized blocks 11 to 23 based on recursive quadtree and binary tree block partitioning. Such partitioning is sometimes called QTBT (quad-tree plus binary tree) partitioning.
[0286] In Figure 21, one block was divided into four or two blocks (quadrutree or binary tree block partitioning), but partitioning is not limited to these. For example, one block may be divided into three blocks (ternary tree block partitioning). Partitioning that includes such ternary tree block partitioning is sometimes called MTT (Multi-Type Tree) partitioning.
[0287] Furthermore, while this example demonstrates a binary tree block partitioning method where a block is divided into two rectangular blocks of the same size in a 1:1 ratio, binary tree block partitioning is not limited to this. For example, a block may be divided into two rectangular blocks of different sizes in a 1:2 or 1:3 ratio. In such cases, parameters indicating these ratios may be used as block partitioning information.
[0288] Figure 22 shows an example of the configuration of the division unit 102. As shown in Figure 22, the division unit 102 may include a block division determination unit 102a. The block division determination unit 102a may perform the following processing as an example.
[0289] The block division determination unit 102a collects block information from, for example, the block memory 118 or the frame memory 122, and determines the division pattern based on that block information. The division unit 102 divides the original image according to the division pattern and outputs one or more blocks obtained by the division to the subtraction unit 104.
[0290] Furthermore, the block division determination unit 102a outputs parameters indicating the division pattern described above to the conversion unit 106, the inverse conversion unit 114, the intra prediction unit 124, the inter prediction unit 126, and the entropy coding unit 110. The conversion unit 106 may convert the prediction residuals based on these parameters, and the intra prediction unit 124 and the inter prediction unit 126 may generate a prediction image based on these parameters. The entropy coding unit 110 may also perform entropy coding on these parameters.
[0291] The parameters related to the splitting pattern may be written to the stream as follows, for example.
[0292] Figure 23 shows examples of division patterns. Division patterns include, for example, quadripartition (QT), which divides the block into two horizontally and two vertically; tripartition (HT or VT), which divides the block in the same direction in a 1:2:1 ratio; duplicate division (HB or VB), which divides the block in the same direction in a 1:1 ratio; and no division (NS).
[0293] Note that in the case of four divisions and no division, the division pattern does not have a block division direction, while in the case of two divisions and three divisions, the division pattern has division direction information.
[0294] Figures 24A and 24B show examples of syntax trees for split patterns. In the example in Figure 24A, first there is information indicating whether or not to perform a split (S: Split flag), then there is information indicating whether or not to perform a four-way split (QT: QT flag). Next there is information indicating whether to perform a three-way split or a two-way split (TT: TT flag or BT: BT flag), and finally there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag).
[0295] Furthermore, the same division process may be repeatedly applied to each of the one or more blocks obtained by the division using such a division pattern. That is, as an example, the determination of whether or not to divide, whether or not to divide into four, whether or not to divide horizontally or vertically, and whether or not to divide into three or two may be performed recursively, and the results of the determinations performed may be encoded into a stream according to the encoding order disclosed in the syntax tree shown in Figure 24A.
[0296] Furthermore, in the syntax tree shown in Figure 24A, the information is arranged in the order of S, QT, TT, and Ver, but it may also be arranged in the order of S, QT, Ver, and BT. In other words, in the example in Figure 24B, first there is information indicating whether or not to perform a split (S: Split flag), then there is information indicating whether or not to perform a four-way split (QT: QT flag). Next there is information indicating the direction of the split (Ver: Vertical flag or Hor: Horizontal flag), and finally there is information indicating whether to perform a two-way split or a three-way split (BT: BT flag or TT: TT flag).
[0297] Note that the division patterns described here are just examples; you may use other division patterns, or only a part of the division patterns described.
[0298] [Loop Filtering Section in Encoding Process] The loop filter section 120 in the encoding device 100 applies a filter to the reconstructed image output from the adder 116 and outputs the filtered reconstructed image to the frame memory 122. Here, the filter performed by the loop filter section 120 is a filter used within the encoding loop (also called a loop filter or in-loop filter).
[0299] Filtering is a technique that reduces encoding noise generated by quantization processing within the encoding loop. Filtering can directly reduce visually noticeable block distortion or ringing distortion, improving subjective and / or objective performance. Furthermore, the loop filter unit 120 prevents image quality degradation from propagating between frames by performing filtering within the loop.
[0300] Filters include, for example, adaptive loop filters (ALF), deblocking filters (DF or DBF), sample adaptive offset (SAO), luminance mapping with chroma scaling (LMCS), or any combination thereof.
[0301] Figure 25 is a block diagram showing an example of the configuration of the loop filter unit 120. The loop filter unit 120 comprises, for example, an LMCS processing unit 120d, a DBF processing unit 120a, a SAO processing unit 120b, and an ALF processing unit 120c, as shown in Figure 25.
[0302] The LMCS processing unit 120d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 120a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 120b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 120c applies ALF processing to the reconstructed image after SAO processing. Details of ALF and DBF will be described later.
[0303] SAO processing is a process that improves image quality by reducing ringing (a phenomenon in which pixel values are distorted in a wave-like manner around edges) and correcting pixel value misalignment. Examples of SAO processing include edge offset processing and band offset processing.
[0304] LMCS processing is a process that reassigns codewords assigned to rarely occurring pixel values to codewords assigned to frequently occurring pixel values when there is a large bias in the pixel value distribution of the luminance signal of the original image. Examples of LMCS processing include mapping the luminance signal using a piecewise linear model based on the pixel value distribution of the original image, and scaling the residual color difference signal according to the luminance pixel values.
[0305] Note that the loop filter unit 120 does not necessarily have to include all of the processing units disclosed in Figure 25, and may include only some of them. Also, the loop filter unit 120 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 25.
[0306] [Loop Filter Section > Adaptive Loop Filter] In ALF, a least-squares error filter is applied to remove encoding distortion. For example, for each 2x2 pixel subblock within the current block, one filter selected from several filters is applied based on the direction and activity of the local gradient.
[0307] Specifically, first, subblocks (for example, 2x2 pixel subblocks) are classified into multiple classes (for example, 15 or 25 classes). The classification of subblocks is done, for example, based on the direction of the gradient and the activity level. In a specific example, a classification value C (for example, C = 5D + A) is calculated using the gradient direction value D (for example, 0 to 2 or 0 to 4) and the gradient activity value A (for example, 0 to 4). Then, based on the classification value C, the subblocks are classified into multiple classes.
[0308] The gradient direction value D is derived, for example, by comparing gradients in multiple directions (e.g., horizontal, vertical, and two diagonal directions). The gradient activation value A is derived, for example, by adding the gradients in multiple directions and quantizing the sum.
[0309] Based on the results of this classification, a filter for the subblock is determined from among multiple filters.
[0310] For example, circularly symmetrical shapes are used for filters in ALF. Figures 26A to 26C show several examples of filter shapes used in ALF. Figure 26A shows a 5x5 diamond-shaped filter, Figure 26B shows a 7x7 diamond-shaped filter, and Figure 26C shows a 9x9 diamond-shaped filter.
[0311] Information indicating the shape of the filter is typically signaled at the picture level. However, the signaling of information indicating the shape of the filter is not limited to the picture level; it may be at other levels (e.g., sequence level, slice level, brick level, CTU level, or CU level).
[0312] The on / off status of the ALF may be determined, for example, at the picture level or the CU level. For example, the decision to apply the ALF may be made at the CU level for luminance, and at the picture level for color difference. Information indicating whether the ALF is on or off is usually signaled at the picture level or the CU level.
[0313] Note that the signaling of ALF on / off information is not limited to picture level or CU level, but may be at other levels (e.g., sequence level, slice level, brick level, or CTU level). If the ALF on / off information read from the stream indicates that ALF is on, one filter is selected from several filters based on the direction and activity of the local gradient, and the selected filter is applied to the reconstructed image.
[0314] Furthermore, as described above, one filter is selected from among several filters and ALF processing is applied to the subblock. For each of these filters (for example, up to 15 or 25 filters), the set of coefficients used in that filter is usually signaled at the picture level. However, the signaling of the coefficient set is not limited to the picture level and may be at other levels (for example, sequence level, slice level, brick level, CTU level, CU level, or subblock level).
[0315] [Loop Filter Section > CCALF] CC-ALF (Cross Component Adaptive Loop Filter) is a type of restoration filter that uses correlations between color components. The filtering results of eight luminance signals corresponding to the positions of the color difference signal being processed are added to the color difference signal. To enable parallel processing with ALF, the luminance signal before the application of ALF is used for filtering.
[0316] Figure 26D shows a configuration diagram of an example of a CCALF in this disclosure. Figure 26E shows an example of a filter shape of a CCALF in this disclosure.
[0317] One example of CC-ALF operates by applying a linear diamond-shaped filter (Figures 26D and 26E) to the luminance channel of each chromatic difference component. For example, the filter coefficients are transmitted via APS, scaled by a factor of 2^10, and rounded for fixed-point representation. The application of the filter is controlled by a variable block size and indicated by a context-encoded flag received for each block of samples. The block size and CC-ALF enable flag are received at the slice level for each chromatic difference component.
[0318] The syntax and semantics of CC-ALF are provided in the encoding standard. Block sizes of 16x16, 32x32, 64x64, and 128x128 may also be supported for color difference samples.
[0319] [Loop Filter Section > DBF] In DBF processing, the loop filter section 120 reduces distortion occurring at block boundaries by applying a filter to the block boundaries of the reconstructed image.
[0320] Figure 27A is a block diagram showing an example of the detailed configuration of the DBF processing unit 120a. The DBF processing unit 120a includes, for example, a boundary determination unit 1201, a filter determination unit 1203, a filter processing unit 1205, a processing determination unit 1208, a filter characteristic determination unit 1207, and switches 1202, 1204, and 1206.
[0321] Figure 27B is a flowchart showing an example of the processing performed by the DBF processing unit 120a.
[0322] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0b).
[0323] Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1b). If it determines to perform DBF, it proceeds to the next determination. If the DBF processing unit 120a determines not to perform DBF, it terminates the process without performing DBF (step S_2b).
[0324] Specifically, for example, the DBF processing unit 120a may determine in the boundary determination unit 1201 whether or not the target pixel for DBF processing is located near a block boundary, and if the target pixel is not located near a block boundary, it may decide not to perform DBF. Alternatively, the DBF processing unit 120a may calculate a Bs value and decide not to perform DBF according to the Bs value. For example, if the Bs value is equal to 0, it indicates that DBF will not be performed, and if it is 1 or greater, it indicates that there is a possibility of performing DBF.
[0325] Next, the DBF processing unit 120a determines whether or not to perform DBF on the target pixel in the filter determination unit 1203 (step S_3b). In other words, the DBF processing unit 120a determines the type of filter to be applied. Specifically, for example, the filter determination unit 1203 determines whether or not to perform DBF processing on the target pixel based on the pixel values of at least one surrounding pixel located around the target pixel. The image before filtering is an image consisting of the target pixel and at least one surrounding pixel located around that target pixel.
[0326] The filter determination unit 1203 may compare a value based on the pixel value of at least one peripheral pixel, and / or the difference between multiple pixels, etc., with a threshold value, and decide whether to execute the DBF to be determined if the value is greater than the threshold value, and whether to execute the DBF to be determined if the value is less than or equal to the threshold value. The threshold value is, for example, β or tC determined based on the quantization parameter QP. In other words, in the DBF process, for example, the DBF to be performed is selected based on quantization.
[0327] In one example of the process for determining whether or not to perform the DBF to be judged, it is first determined whether or not to perform DBF(a). If it is determined that DBF(a) should be performed, DBF(a) is performed (step S_4b). If it is determined that DBF(a) should not be performed, the next determination is made. In the next determination, it is determined whether or not to perform DBF(b). If it is determined that DBF(b) should be performed, DBF(b) is performed (step S_5b). If it is determined that DBF(b) should not be performed, DBF is not performed (step S_2b).
[0328] Furthermore, if it is decided not to implement DBF(b), there may be multiple DBF candidates, such as deciding whether or not to implement DBF(c). Here, DBF(a), which has been decided to implement, may be further subdivided, and it may be decided whether or not to implement DBF(a_1) or DBF(a_2).
[0329] Here, DBF(a), DBF(b), DBF(a_1), and DBF(a_2) may be short filter, long filter, weak filter, and strong filter, respectively.
[0330] Figure 27C is a flowchart showing another example of the processing performed by the DBF processing unit 120a.
[0331] First, the DBF processing unit 120a determines the boundary to be processed in the boundary determination unit 1201 (step S_0c). Next, the DBF processing unit 120a determines whether or not to perform DBF (step S_1c), and if it determines to perform DBF, it selects the DBF to perform (step S_2c). Then, the DBF processing unit 120a performs the selected DBF (step S_3c). If the DBF processing unit 120a determines in S_1c not to perform DBF, it terminates the process without performing DBF (step S_4c).
[0332] Note that the processing of S_1c may be the same as or different from the processing of S_1b. Also, the processing of S_2c may be the same as or different from the processing of S_3b.
[0333] Figure 28 shows an example of a DBF (Digital Block Filter) with symmetrical filter characteristics with respect to block boundaries. DBF is a process that reduces distortion at block boundaries by correcting the sample values on both sides of the block boundary. The number of pixels to be filtered and the number of pixels used for filtering vary depending on the type of filter. The higher the filter strength, the greater the number of pixels to be filtered and the greater the number of pixels used for filtering.
[0334] When DBF processing is performed on the block boundary between block P and block Q adjacent to block P, the pixels to be processed are pn (n=0, ..., n) in block P and qm (m=0, ..., m) in block Q. Pixels with a value of n or m of 0 are adjacent to the boundary, while pixels with larger values are further from the boundary.
[0335] For example, in a strong filter, as shown in Figure 28, processing is performed on pixels p0 to p2 in block P and pixels q0 to q2 in block Q. The respective pixel values of pixels q0 to q2 are changed to pixel values q'0 to q'2 by performing the calculation shown in the following equation.
[0336] q'0=(p1+2×p0+2×q0+2×q1+q2+4) / 8 q'1=(p0+q0+q1+q2+2) / 4 q'2=(p0+q0+q1+3×q2+2×q3+4) / 8
[0337] In the above equations, p0 to p2 and q0 to q2 are the pixel values of pixels p0 to p2 and pixels q0 to q2, respectively. Also, q3 is the pixel value of pixel q3, which is adjacent to pixel q2 on the opposite side of the block boundary. Furthermore, the coefficient multiplied by the pixel value of each pixel used in the deblocking filter process on the right-hand side of each of the above equations is the filter coefficient.
[0338] Furthermore, in DBF processing, clipping may be performed to ensure that the pixel value after calculation does not change beyond a threshold. In this clipping process, the pixel value after calculation using the above formula is clipped to "pre-calculation pixel value ± a × threshold (where a is an integer)" using a threshold determined from the quantization parameters. In other words, the amount of change in the pixel value before and after calculation is clipped so that it is within the range of ± a × threshold (where a is an integer). This prevents excessive smoothing. Specifically, for example, the pixel values q'0 to q'2 are expressed by the following formula.
[0339] q'0=Clip3 (q0-3×tC, q0+3×tC, (p1+2×p0+2×q0+2×q1+q2+4)>>3) q'1=Clip3 (q1-2×tC, q1+2×tC, (p0+q0+q1+q2+2)>>2) q'2=Clip3 (q2-1×tC, q2+1×tC, (p0+q0+q1+3×q2+2×q3+4)>>3)
[0340] Furthermore, Clip3 may be defined as follows:
[0341]
[0342] Figure 29 is a diagram illustrating an example of a block boundary where DBF processing is performed. Figure 30 is a diagram illustrating an example of a block boundary strength Bs value.
[0343] The block boundaries on which DBF processing is performed are, for example, the CU, PU, or TU boundaries of an 8x8 pixel block as shown in Figure 29. DBF processing is performed in units of, for example, 4 rows or 4 columns. First, for blocks P and Q shown in Figure 29, the Boundary Strength (Bs) value is determined as shown in Figure 30. The Bs value is determined independently for each color component.
[0344] In Figure 30, conditions listed higher up have higher priority, and evaluation is performed in order from the highest priority condition. If a condition is met, evaluation of lower priority conditions is not performed. The magnitude and range of block distortion differ depending on the boundary, and excessive smoothing can cause problems such as blurring. Therefore, in this example, the Bs value is determined according to the prediction modes of blocks P and Q, the presence or absence of motion vectors and transformation coefficients.
[0345] Furthermore, based on the determined Bs value, it may be decided whether or not to perform DBF processing of different strengths even for block boundaries belonging to the same image. DBF processing for chrominance signals is performed when the Bs value is 2. DBF processing for luminance signals is performed when the Bs value is 1 or greater and predetermined conditions are met. Note that the criteria for determining the Bs value are not limited to the example shown in Figure 30, and may be determined based on other parameters.
[0346] The above method for determining the Bs value and the method for performing DBF processing based on the Bs value are examples only, and the method for determining the Bs value and the method for performing DBF processing based on the Bs value are not limited to the above example.
[0347] In DBF processing, it is determined whether or not to apply a filter to each of the luminance sample and chrominance sample, and the type of filter to be applied. Furthermore, in DBF processing, it is determined whether or not to apply a filter to each of the vertical and horizontal boundaries contained in each of the luminance sample and chrominance sample, and the type of filter to be applied.
[0348] Some or all of the processing may be performed on either the luminance sample or the chrominance sample, or on both. The luminance sample and the chrominance sample may undergo the same processing, or they may undergo different processing.
[0349] [Loop Filtering Unit in Decoding Process] The loop filtering unit 212 in the decoding device 200 applies a loop filter to the reconstructed image generated by the summing unit 208, and outputs the filtered reconstructed image to the frame memory 214 and the display device, etc.
[0350] Figure 31 is a block diagram showing an example of the configuration of the loop filter unit 212. The loop filter unit 212 has a configuration similar to that of the loop filter unit 120 of the encoding device 100. The loop filter unit 212 includes, for example, an LMCS processing unit 212d, a DBF processing unit 212a, a SAO processing unit 212b, and an ALF processing unit 212c, as shown in Figure 31.
[0351] The LMCS processing unit 212d performs luminance mapping and color difference scaling on the reconstructed image. The DBF processing unit 212a performs DBF processing on the reconstructed image after LMCS processing. The SAO processing unit 212b performs SAO processing on the reconstructed image after DBF processing. In addition, the ALF processing unit 212c applies ALF processing to the reconstructed image after SAO processing.
[0352] Furthermore, the loop filter unit 212 does not necessarily have to include all of the processing units disclosed in Figure 31, and may include only some of them. Also, the loop filter unit 212 may perform the above-mentioned processing in an order different from the processing order disclosed in Figure 31.
[0353] [Prediction Unit (Intra Prediction Unit, Inter Prediction Unit, Prediction Control Unit)] Figure 32 is a flowchart showing an example of processing performed in the prediction unit of the encoding device 100. For example, the prediction unit consists of all or some of the components of the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128. The prediction processing unit includes, for example, the intra prediction unit 124 and the inter prediction unit 126.
[0354] The prediction unit generates a predicted image of the current block (step Sb_1). The predicted image may be, for example, an intra-prediction image (intra-prediction signal) or an inter-prediction image (inter-prediction signal). Specifically, the prediction unit generates a predicted image of the current block using a reconstructed image already obtained by generating predicted images for other blocks, generating prediction residuals, generating quantization coefficients, restoring prediction residuals, and adding the predicted images.
[0355] The reconstructed image may be, for example, the image of the reference picture, or it may be the image of an encoded block (i.e., the other block mentioned above) within the current picture, which is the picture containing the current block. An encoded block within the current picture is, for example, an adjacent block to the current block.
[0356] Figure 33 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.
[0357] The prediction unit generates a predicted image using a first method (step Sc_1a), a second method (step Sc_1b), and a third method (step Sc_1c). The first, second, and third methods are different methods for generating predicted images, and may be, for example, an interpretation method, an intraprediction method, and other prediction methods. These prediction methods may use the reconstructed images described above.
[0358] Next, the prediction unit evaluates the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_2). For example, the prediction unit calculates a cost C for each of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c, and evaluates the predicted images by comparing the costs C of those predicted images.
[0359] The cost C is calculated using the R-D optimization model formula, for example, C = D + λ × R. In this formula, D is the coding distortion of the predicted image, which can be expressed as, for example, the sum of the absolute differences between the pixel values of the current block and the pixel values of the predicted image. R is the bitrate of the stream, and λ is, for example, the Lagrange multiplier.
[0360] Next, the prediction unit selects one of the predicted images generated in steps Sc_1a, Sc_1b, and Sc_1c (step Sc_3). In other words, the prediction unit selects a method or mode for obtaining the final predicted image. For example, the prediction unit selects the predicted image with the smallest cost C based on the cost C calculated for those predicted images. Alternatively, the evaluation in step Sc_2 and the selection of the predicted image in step Sc_3 may be based on parameters used in the coding process.
[0361] The encoding device 100 may signal information to identify the selected prediction image, scheme, or mode into a stream. This information may be, for example, a flag. Based on this information, the decoding device 200 can generate a prediction image according to the scheme or mode selected by the encoding device 100.
[0362] In the example shown in Figure 33, the prediction unit generates predicted images using each method and then selects one of the predicted images. However, the prediction unit may also select a method or mode based on the parameters used in the encoding process described above before generating those predicted images, and then generate the predicted images according to that method or mode.
[0363] For example, the first method and the second method are intra-prediction and inter-prediction, respectively, and the prediction unit may select the final predicted image for the current block from the predicted images generated according to these prediction methods.
[0364] Figure 34 is a flowchart showing another example of the processing performed in the prediction unit of the encoding device 100.
[0365] First, the prediction unit generates a predicted image by intra-prediction (step Sd_1a) and then generates a predicted image by inter-prediction (step Sd_1b). The predicted image generated by intra-prediction is also called an intra-prediction image, and the predicted image generated by inter-prediction is also called an inter-prediction image.
[0366] Next, the prediction unit evaluates both the intra-predicted image and the inter-predicted image (step Sd_2). The cost C mentioned above may be used for this evaluation. The prediction unit may then select the prediction image with the smallest cost C from the intra-predicted image and the inter-predicted image as the final prediction image for the current block (step Sd_3). In other words, a prediction method or mode for generating the prediction image for the current block is selected.
[0367] [Prediction Control Unit] The prediction control unit 128 selects either an intra-prediction image (an image or signal output from the intra-prediction unit 124) or an inter-prediction image (an image or signal output from the inter-prediction unit 126), and outputs the selected prediction image to the subtraction unit 104 and the addition unit 116.
[0368] [Prediction Parameter Generation Unit] The prediction parameter generation unit 130 may output information regarding intra-prediction, inter-prediction, and the selection of a predicted image in the prediction control unit 128 as prediction parameters to the entropy coding unit 110. The entropy coding unit 110 may generate a stream based on the prediction parameters input from the prediction parameter generation unit 130 and the quantization coefficients input from the quantization unit 108. The prediction parameters may be used by the decoding device 200.
[0369] The decoding device 200 may receive and decode the stream and perform the same processing as the prediction processing performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0370] The prediction parameter may include a selection prediction signal (e.g., MV, prediction type, or prediction mode used in the intra prediction unit 124 or the inter prediction unit 126), or any index, flag, or value based on or indicating the prediction process performed in the intra prediction unit 124, the inter prediction unit 126, and the prediction control unit 128.
[0371] [Inter Prediction Unit] The inter prediction unit 126 generates a predicted image (inter prediction image) by performing inter prediction (also called inter-frame prediction) of a current block by referring to a reference picture stored in the frame memory 122 that is different from the current picture.
[0372] Inter prediction is performed in units of a current block or a current sub-block within a current block. A sub-block is included in a block and is a unit smaller than the block. The size of the sub-block may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the sub-block may be switched in units such as slices, bricks, or pictures.
[0373] For example, the inter prediction unit 126 performs motion search in the reference picture for a current block or a current sub-block, and finds the reference block or sub-block that most matches the current block or current sub-block. Then, the inter prediction unit 126 obtains motion information (e.g., a motion vector) for compensating the motion or change from the reference block or sub-block to the current block or sub-block.
[0374] The inter prediction unit 126 performs motion compensation (or motion prediction) based on the motion information, and generates an inter prediction image of the current block or sub-block. The inter prediction unit 126 outputs the generated inter prediction image to the prediction control unit 128.
[0375] Motion compensation is a process of generating a predicted image using one or more reference pictures and one or more motion vectors. The predicted image is generated by an interpolation process that uses sample values of a region specified by a motion vector and its peripheral region as a region corresponding to a current block in the reference picture.
[0376] When each element in the horizontal and vertical directions of the motion vector is represented by an integer value, the sample value of the corresponding region may be used as it is without interpolation processing. Also, in motion compensation, a predicted image may be generated using a reconstructed image included in the current picture without using one or more reference pictures.
[0377] In the present disclosure, the term motion compensation is used, but this term may be rephrased as the term interpolation. This rephrasing may be applied throughout the specification.
[0378] The motion information used for motion compensation may be signaled as an inter-predicted image in various forms. For example, a motion vector may be signaled. As another example, a difference between a motion vector and a predicted motion vector may be signaled.
[0379] [Reference Picture List] FIG. 35 is a diagram showing an example of each reference picture, and FIG. 36 is a conceptual diagram showing an example of a reference picture list. The reference picture list is a list showing one or more reference pictures stored in the frame memory 122. In FIG. 35, a rectangle indicates a picture, an arrow indicates a reference relationship between pictures, the horizontal axis indicates time, I, P, and B in the rectangle indicate an intra-predicted picture, a single-predicted picture, and a bi-predicted picture, respectively, and the numbers in the rectangle indicate the decoding order.
[0380] As shown in Figure 35, the decoding order of each picture is I0, P1, B2, B3, B4, and the display order of each picture is I0, B3, B2, B4, P1. As shown in Figure 36, the reference picture list is a list representing candidates for reference pictures, and for example, one picture (or slice) may have one or more reference picture lists. For example, if the current picture is a single-prediction picture, one reference picture list is used, and if the current picture is a double-prediction picture, two reference picture lists are used.
[0381] In the examples in Figures 35 and 36, picture B3, which is the current picture currPic, has two reference picture lists, the L0 list and the L1 list. When the current picture currPic is picture B3, the candidate reference pictures for that current picture currPic are I0, P1, and B2, and each reference picture list (i.e., the L0 list and the L1 list) points to these pictures.
[0382] The interpretation unit 126 or the prediction control unit 128 specifies which picture in each reference picture list to actually reference using the reference picture index refIdxLx. In Figure 36, reference pictures P1 and B2 are specified by reference picture indices refIdxL0 and refIdxL1.
[0383] Such reference picture lists may be generated on a sequence, picture, slice, brick, CTU, or CU basis. Furthermore, the reference picture index indicating the reference picture referenced in interpretation among the reference pictures shown in the reference picture list may be encoded at the sequence, picture, slice, brick, CTU, or CU level. Additionally, a common reference picture list may be used across multiple interpretation modes.
[0384] [Generated Reference Picture] Figure 37A is a conceptual diagram showing another example of a reference picture. As illustrated in Figure 37A, a new image (PicA') generated based on an image (PicA) may be used as a reference picture. A reference picture generated in this way is defined as a generated reference picture. A generated reference picture may be used as a reference picture in the prediction process.
[0385] Additionally, the generated reference picture may be added to the same reference picture list as the reference picture list to which other reference pictures are added. Alternatively, a separate reference picture list may be generated, and the generated reference picture may be added to that separate reference picture list.
[0386] Figure 37B is a conceptual diagram illustrating another example of generating a generated reference picture. In the example in Figure 37B, a new image (PicAB') is generated by inputting two images (PicA and PicB) into an NPU (Neural Processing Unit) or a GPU (Graphics Processing Unit). The generated new image (PicAB') is then used as a generated reference picture for interpretation.
[0387] Interpretation using this generated reference picture may also be called neural network intercoding (NN inter). Note that the NPU or GPU is not limited to two images, but can receive two or more images as input.
[0388] Figure 37C is a conceptual diagram illustrating another example of generating a generated reference picture. As illustrated in Figure 37C, the size of the input image and the size of the generated image may differ. In this example, the size of the generated image (PicA') is smaller than the size of the input image (PicA). The size of the generated image is not limited to this example and may be larger than the size of the input image. Image processing or the size of the generated image per frame may be based on the performance of the NPU or GPU. This size information may also be included in the bitstream.
[0389] In the example above, a new reference picture (also called a generated image or processed image) is generated from one or more encoded or decoded reference pictures using a neural network. The method for generating the new reference picture is not limited to the example above. For example, a new reference picture may be generated using a process such as image conversion, and the generated new reference picture may be defined as the generated reference picture.
[0390] Adding these generated reference pictures to the reference picture list and making them accessible may improve encoding efficiency.
[0391] In the example above, images are generated on a picture-by-picture basis, but the generation unit is not limited to this example. Reference data may be generated in predetermined units such as blocks, slices, or tiles contained within a picture.
[0392] [Basic Flowchart of Interpretation] Figure 38 is a flowchart showing the basic flow of interpretation.
[0393] The interpretation unit 126 first generates a predicted image (steps Se_1 to Se_3). Next, the subtraction unit 104 generates the difference between the current block and the predicted image as the predicted residual (step Se_4).
[0394] Here, the interpretation unit 126 generates the predicted image by, for example, determining the motion vector (MV) of the current block (steps Se_1 and Se_2) and performing motion compensation (step Se_3). In addition, the interpretation unit 126 determines the MV by, for example, selecting a candidate motion vector (candidate MV) (step Se_1) and deriving the MV (step Se_2).
[0395] The selection of a candidate MV is performed, for example, by the inter-prediction unit 126 generating a candidate MV list and selecting at least one candidate MV from the candidate MV list. Note that previously derived MVs may be added to the candidate MV list as candidate MVs. Furthermore, in the MV derivation process, the inter-prediction unit 126 may determine the selected at least one candidate MV as the MV for the current block by selecting at least one more candidate MV from the at least one candidate MV.
[0396] Alternatively, the interpretation unit 126 may determine the MV of the current block by searching the region of the reference picture indicated by each of the selected at least one candidate MV. This search of the region of the reference picture may be called motion estimation.
[0397] Furthermore, in the above example, steps Se_1 to Se_3 are performed by the interpretation unit 126, but processing such as step Se_1 or step Se_2 may be performed by other components included in the encoding device 100.
[0398] Furthermore, a candidate MV list may be created for each process in each interpretation mode, or a common candidate MV list may be used across multiple interpretation modes. Also, the processes in steps Se_3 and Se_4 correspond to the processes in steps Sa_3 and Sa_4 shown in Figure 4B, respectively. In addition, the process in step Se_3 corresponds to the process in step Sd_1b in Figure 34.
[0399] [MV Derivation Flowchart] Figure 39 is a flowchart showing an example of MV derivation.
[0400] The interpretation unit 126 may derive the MV of the current block in a mode that encodes motion information (e.g., MV). In this case, for example, the motion information may be encoded as prediction parameters and then signaled. That is, the encoded motion information is included in the stream.
[0401] Alternatively, the interpretation unit 126 may derive MV in a mode that does not encode motion information. In this case, motion information is not included in the stream.
[0402] Here, the modes for MV derivation include the normal intermode, normal merge mode, FRUC mode, and affine mode, which will be described later. Of these modes, the modes that encode motion information include the normal intermode, normal merge mode, and affine mode (specifically, the affine intermode and affine merge mode). Note that motion information may include not only MV but also the predicted MV selection information described later. Modes that do not encode motion information include the FRUC mode, among others.
[0403] The interpretation unit 126 selects a mode from these multiple modes for deriving the MV of the current block, and uses the selected mode to derive the MV of the current block.
[0404] Figure 40 is a flowchart showing another example of MV derivation.
[0405] The interpretation unit 126 may derive the MV of the current block in a mode that encodes the differential MV. In this case, for example, the differential MV is encoded as a prediction parameter and signaled. That is, the encoded differential MV is included in the stream. This differential MV is the difference between the MV of the current block and its predicted MV. The predicted MV is the predicted motion vector.
[0406] Alternatively, the interpretation unit 126 may derive MV in a mode that does not encode the differential MV. In this case, the encoded differential MV is not included in the stream.
[0407] Here, as described above, the modes for deriving the MV include the normal inter, normal merge mode, FRUC mode, and affine mode, etc. among these modes, the modes for encoding the differential MV include the normal inter mode and the affine mode (specifically, the affine inter mode), etc. Also, the modes for not encoding the differential MV include the FRUC mode, the normal merge mode, and the affine mode (specifically, the affine merge mode), etc.
[0408] The inter prediction unit 126 selects a mode for deriving the MV of the current block from these multiple modes, and derives the MV of the current block using the selected mode.
[0409] [Mode of MV Derivation] FIGS. 41A and 41B are diagrams showing an example of the classification of each mode of MV derivation. For example, as shown in FIG. 41A, depending on whether to encode the motion information and whether to encode the differential MV, the mode of MV derivation is roughly classified into three modes. The three modes are the inter mode, the merge mode, and the FRUC (frame rate up-conversion) mode. The inter mode is a mode that performs motion search and is a mode that encodes motion information and differential MV.
[0410] For example, as shown in FIG. 41B, the inter mode includes the affine inter mode and the normal inter mode. The merge mode is a mode that does not perform motion search, selects an MV from the surrounding encoded blocks, and derives the MV of the current block using that MV. This merge mode is basically a mode that encodes motion information and does not encode differential MV.
[0411] For example, as shown in Figure 41B, the merge modes include normal merge mode (sometimes called regular merge mode), MMVD (Merge with Motion Vector Difference) mode, CIIP (Combined inter merge / intra prediction) mode, GPM mode, ATMVP mode, and affine merge mode. Here, among the modes included in the merge modes, the MMVD mode is an exception in which the differential MV is encoded.
[0412] The aforementioned affine merge mode and affine inter mode are modes included in affine mode. Affine mode is a mode that assumes an affine transformation and derives the MV of each of the multiple subblocks that make up the current block as the MV of the current block. FRUC mode is a mode that derives the MV of the current block by performing a search between encoded regions, and does not encode either motion information or differential MV. Details of each of these modes will be described later.
[0413] Note that the classification of modes shown in Figures 41A and 41B is merely an example and is not limited to this classification. For example, if the differential MV is encoded in CIIP mode, that CIIP mode is classified as an intermode. Also, the MV derivation mode may include GPM (Geometric Partitioning Mode). GPM may also be expressed as geometric shape partitioning prediction mode or GPM mode.
[0414] [MV Derivation > Normal Intermode] Normal intermode is an interpretation mode that derives the MV of the current block by finding blocks similar to the image of the current block from the region of the reference picture indicated by the candidate MV. In this normal intermode, the differential MV is also encoded.
[0415] Figure 42 is a flowchart showing an example of inter-mode prediction.
[0416] The interpretation unit 126 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks surrounding the current block in time or space (step Sg_1). In other words, the interpretation unit 126 creates a candidate MV list.
[0417] Next, the interpretation unit 126 extracts N candidate MVs (where N is an integer greater than or equal to 2) from the multiple candidate MVs obtained in step Sg_1 as predicted MV candidates, according to a predetermined priority order (step Sg_2). The priority order is predetermined for each of the N candidate MVs.
[0418] Next, the interpretation unit 126 selects one predicted MV candidate from the N predicted MV candidates as the predicted MV for the current block (step Sg_3). At this time, the interpretation unit 126 encodes predicted MV selection information into a stream to identify the selected predicted MV. In other words, the interpretation unit 126 outputs the predicted MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0419] Next, the interpretation unit 126 refers to the encoded reference picture and derives the MV of the current block (step Sg_4). At this time, the interpretation unit 126 further encodes the difference between the derived MV and the predicted MV as the difference MV into the stream. In other words, the interpretation unit 126 outputs the difference MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130. The encoded reference picture is a picture consisting of multiple blocks that have been reconstructed after encoding.
[0420] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sg_5).
[0421] Steps Sg_1 to Sg_5 are performed for each block. For example, once steps Sg_1 to Sg_5 have been performed for each of the blocks contained in a slice, the inter-mode prediction for that slice is complete. Similarly, once steps Sg_1 to Sg_5 have been performed for each of the blocks contained in a picture, the inter-mode prediction for that picture is complete.
[0422] Note that the processing in steps Sg_1 to Sg_5 is not performed for all blocks included in a slice; if it is performed for some blocks, the inter prediction using the normal inter mode for that slice may be terminated. Similarly, if the processing in steps Sg_1 to Sg_5 is performed for some blocks included in a picture, the inter prediction using the normal inter mode for that picture may be terminated.
[0423] The predicted image is the interprediction signal described above. Furthermore, information indicating the interprediction mode used to generate the predicted image (normal intermode in the example above), which is included in the encoded signal, is encoded, for example, as a prediction parameter.
[0424] The candidate MV list may be the same as the list used in other modes. Furthermore, processing related to the candidate MV list may be applied to processing related to lists used in other modes. This processing related to the candidate MV list may include, for example, extracting or selecting candidate MVs from the candidate MV list, rearranging candidate MVs, or deleting candidate MVs.
[0425] [MV Derivation > Normal Merge Mode] Normal merge mode is an interpretation mode in which an MV is derived by selecting a candidate MV from a list of candidate MVs as the MV of the current block. Note that normal merge mode is a merge mode in the narrow sense and is sometimes simply called merge mode. In this embodiment, a distinction is made between normal merge mode and merge mode, and merge mode may be used in a broad sense. Candidate MVs in merge mode are also called merge candidates or merge mode candidates.
[0426] Figure 43 is a flowchart showing an example of interpretation using normal merge mode.
[0427] The interpretation unit 126 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks surrounding the current block in time or space (step Sh_1). In other words, the interpretation unit 126 creates a candidate MV list.
[0428] Next, the interpretation unit 126 derives the MV of the current block by selecting one candidate MV from among the multiple candidate MVs obtained in step Sh_1 (step Sh_2). At this time, the interpretation unit 126 encodes MV selection information to identify the selected candidate MV into a stream. In other words, the interpretation unit 126 outputs the MV selection information as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0429] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Sh_3).
[0430] Steps Sh_1 to Sh_3 are executed for each block, for example. For instance, once steps Sh_1 to Sh_3 have been executed for each block in a slice, the inter prediction using normal merge mode for that slice is complete. Similarly, once steps Sh_1 to Sh_3 have been executed for each block in a picture, the inter prediction using normal merge mode for that picture is complete.
[0431] Note that the processing in steps Sh_1 to Sh_3 is not performed on all blocks included in a slice; if it is performed on some blocks, the inter prediction using normal merge mode for that slice may be terminated. Similarly, if the processing in steps Sh_1 to Sh_3 is performed on some blocks included in a picture, the inter prediction using normal merge mode for that picture may be terminated.
[0432] Furthermore, information indicating the inter-prediction mode used to generate the predicted image (normal merge mode in the example above), which is included in the stream, is encoded, for example, as prediction parameters.
[0433] Figure 44 is a diagram illustrating an example of the MV derivation process for the current picture using normal merge mode.
[0434] First, the interpretation unit 126 generates a candidate MV list in which candidate MVs are registered. Candidate MVs include spatially adjacent candidate MVs, which are the MVs of multiple encoded blocks located spatially around the current block; temporally adjacent candidate MVs, which are the MVs of nearby blocks projected onto the current block's position in the encoded reference picture; combined candidate MVs, which are MVs generated by combining the MV values of spatially adjacent candidate MVs and temporally adjacent candidate MVs; and zero candidate MVs, which are MVs with a value of zero.
[0435] Next, the interpretation unit 126 selects one candidate MV from among the multiple candidate MVs registered in the candidate MV list, thereby determining that one candidate MV as the MV for the current block.
[0436] Furthermore, the entropy coding unit 110 writes and encodes merge_idx, a signal indicating which candidate MV was selected, into the stream.
[0437] Note that the candidate MVs registered in the candidate MV list explained in Figure 44 are just an example, and the number of candidates may differ from the number shown in the figure, the configuration may not include some of the candidate MV types shown in the figure, or it may include candidate MV types other than those shown in the figure.
[0438] The final MV may be determined by performing DMVR (decoder motion vector refinement), described later, using the MV of the current block derived by normal merge mode. Note that in normal merge mode, the differential MV is not encoded, but in MMVD mode, the differential MV is encoded.
[0439] The MMVD mode selects one candidate MV from a list of candidate MVs, similar to the normal merge mode, but encodes a differential MV. Such MMVD may be classified as a merge mode along with the normal merge mode, as shown in Figure 41B. Note that the differential MV in MMVD mode does not have to be the same as the differential MV used in inter-mode; for example, the derivation of the differential MV in MMVD mode may be a less computationally intensive process than the derivation of the differential MV in inter-mode.
[0440] Alternatively, a Combined Inter Merge / Intra Prediction (CIIP) mode may be used, which generates a prediction image for the current block by overlaying the prediction image generated by inter prediction with the prediction image generated by intra prediction.
[0441] The candidate MV list may also be referred to as the candidate list. Furthermore, merge_idx is the MV selection information.
[0442] Furthermore, in the MV derivation process, including the merge mode, a process to correct the MV of the candidate MV may be performed.
[0443] For example, template matching may be performed on multiple candidate MVs by referencing the reconstructed image of the encoded block and the encoded reference picture, and a correction process may be performed on the candidate MV list. Next, the MV of the current block may be derived by selecting one candidate MV from the multiple corrected candidate MVs. Note that the correction process is not limited to this example and may be performed using other methods. Also, the correction process is not limited to being applied to normal merge mode.
[0444] For example, correction processing may be applied when correcting candidate MVs in other modes. Other modes may include template matching (TM) mode, bilateral matching (BM) merge mode, MMVD (Merge with Motion Vector Difference) merge mode, intrablock copy (IBC) merge mode, CIIP (Combined inter merge / intra prediction) merge mode, GPM merge mode, or affine merge mode.
[0445] Here, the intrablock copy (IBC) merge mode is a mode in which a predicted image is obtained by copying prediction blocks from the encoded or decoded peripheral region of the same picture. The CIIP (Combined inter merge / intra prediction) merge mode is a mode in which the inter-prediction image generated in the normal merge mode and the intra-prediction image generated in the planar prediction mode are combined using a weighted average.
[0446] Furthermore, correction processes such as DMVR correction and template matching correction, as described later, may be applied. The correction processes are not limited to these methods. Other correction processes may be used, or multiple correction processes may be used in combination.
[0447] Furthermore, new merge candidates may be generated and added to the merge mode candidate list in Figure 44. For example, the motion vector mvA and reference picture index refIdxA of the encoded or decoded blocks adjacent to the block to be encoded or decoded are added as spatial adjacency candidates MV. In this case, new merge candidates may be generated using mvB and refidxB of the referenced block referenced using mvA and refIdxA.
[0448] This makes it possible to generate a candidate list with highly accurate motion vectors for the reference picture, which can improve encoding efficiency.
[0449] For example, as described above, a reference block is calculated from the motion vector and reference picture index of the merge candidates included in the merge mode candidate list, and a new merge candidate is generated using the motion vector and reference picture index of that reference block. This makes it possible to generate motion vectors for previously encoded or decoded reference pictures with high accuracy while tracking the motion, thereby improving encoding efficiency.
[0450] The merge candidates calculated as described above are sometimes called CMVP (Chained motion vector prediction) merge candidates.
[0451] Furthermore, the method of generating merge candidates using CMVP is not limited to normal merge mode; it may also be applied to other merge modes, or to the generation of predicted motion vectors, etc. For example, CMVP may be applied to affine merge mode. This may improve the prediction accuracy of affine prediction and improve coding efficiency.
[0452] Furthermore, if the reference block belongs to an intrapicture or intraslice, or if the reference block is a block encoded by intraprediction, the reference block may not have a motion vector. In this case, it is not possible to generate CMVP merge candidates, and therefore, it is acceptable that no CMVP merge candidates are generated from that merge candidate.
[0453] Furthermore, if a referenced block has a block vector (bv) which is a motion vector for blocks within the same picture, CMVP merge candidates may be generated using bv.
[0454] Furthermore, the reference blocks within the generated reference picture mentioned above do not contain motion vectors or reference picture information (such as the reference picture index). Therefore, if a reference block within the generated reference picture is selected as the reference target for generating a CMVP merge candidate, it is not possible to generate a CMVP merge candidate. Thus, it is acceptable that no CMVP merge candidates are generated from that merge candidate.
[0455] Alternatively, the motion vector may be tracked and added until the reference block no longer has motion vectors or reference picture information, and a CMVP merge candidate may be generated using that motion vector and reference picture information. This may improve encoding efficiency.
[0456] [MV Derivation > HMVP Mode] Figure 45 is a diagram illustrating an example of the MV derivation process for the current picture using HMVP mode.
[0457] In normal merge mode, the MV of the current block (e.g., CU) is determined by selecting one candidate MV from a list of candidate MVs generated by referencing the encoded block (e.g., CU). Other candidate MVs may be registered in this list. This mode, where other candidate MVs are registered, is called HMVP mode.
[0458] In HMVP mode, candidate MVs are managed using a separate FIFO (First-In First-Out) buffer for HMV, distinct from the candidate MV list used in normal merge mode.
[0459] The FIFO buffer stores motion information such as MV for blocks that have been processed in the past, in reverse chronological order. In the management of this FIFO buffer, each time a block is processed, the MV of the most recent block (i.e., the most recently processed CU) is stored in the FIFO buffer, and in its place, the MV of the oldest CU in the FIFO buffer (i.e., the CU that was processed first) is deleted from the FIFO buffer. In the example shown in Figure 45, HMVP1 is the MV of the most recent block, and HMVP5 is the MV of the oldest block.
[0460] For example, the interpretation unit 126 checks each MV managed in the FIFO buffer, starting with HMVP1, whether that MV is different from all the candidate MVs already registered in the normal merge mode candidate MV list. If the interpretation unit 126 determines that it is different from all the candidate MVs, it may add the MV managed in the FIFO buffer as a candidate MV to the normal merge mode candidate MV list.
[0461] At this time, one or more candidate MVs may be registered from the FIFO buffer.
[0462] By using HMVP mode in this way, it becomes possible to include not only MVs of spatially or temporally adjacent blocks to the current block, but also MVs of previously processed blocks as candidates. As a result, the variety of candidate MVs in normal merge mode is broadened, which increases the likelihood of improving encoding efficiency.
[0463] Furthermore, the aforementioned MV may also be motion information. In other words, the information stored in the candidate MV list and FIFO buffer may include not only the MV value, but also information indicating the referenced picture, the direction and number of referenced pictures, etc. Also, the aforementioned block is, for example, a CU.
[0464] Note that the candidate MV list and FIFO buffer in Figure 45 are just examples, and the candidate MV list and FIFO buffer may be lists or buffers of different sizes than those in Figure 45, or the candidate MVs may be registered in a different order than those in Figure 45. Furthermore, the processing described here is common to both the encoding device 100 and the decoding device 200.
[0465] Furthermore, HMVP mode can be applied to modes other than normal merge mode. For example, movement information such as MV of blocks previously processed in affine mode can be stored in the FIFO buffer in chronological order from newest to oldest and used as candidate MV. A mode in which HMV mode is applied to affine mode may be called history affine mode.
[0466] [MV Derivation > FRUC Mode] Motion information may be derived on the decoding device 200 side without being signaled from the encoding device 100 side. For example, motion information may be derived by performing a motion search on the decoding device 200 side. In this case, the decoding device 200 side performs a motion search without using the pixel values of the current block. Modes for performing such a motion search on the decoding device 200 side include the FRUC (frame rate up-conversion) mode or the PMMVD (pattern matched motion vector derivative) mode.
[0467] An example of FRUC processing is shown in Figure 46. First, a list is generated that refers to the MVs of each encoded block that is spatially or temporally adjacent to the current block and indicates those MVs as candidate MVs (i.e., a candidate MV list, which may be the same as the candidate MV list for normal merge mode) (step Si_1).
[0468] Next, the best candidate MV is selected from among the multiple candidate MVs registered in the candidate MV list (step Si_2). For example, an evaluation value is calculated for each candidate MV included in the candidate MV list, and based on that evaluation value, one candidate MV is selected as the best candidate MV.
[0469] Then, based on the selected best candidate MV, the MV for the current block is derived (step Si_4). Specifically, for example, the selected best candidate MV is directly derived as the MV for the current block. Alternatively, for example, the MV for the current block may be derived by performing pattern matching in the area surrounding the position in the reference picture corresponding to the selected best candidate MV.
[0470] In other words, a search is performed in the area surrounding the best candidate MV using pattern matching and evaluation values in the reference picture. If an MV with a better evaluation value is found, the best candidate MV may be updated to that MV and made the final MV for the current block. It is not necessary to update to an MV with a better evaluation value.
[0471] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the derived MV and the encoded reference picture (step Si_5).
[0472] Steps Si_1 to Si_5 are performed for each block, for example. For instance, once steps Si_1 to Si_5 have been performed for each block in a slice, the inter prediction using FRUC mode for that slice is complete. Similarly, once steps Si_1 to Si_5 have been performed for each block in a picture, the inter prediction using FRUC mode for that picture is complete.
[0473] Note that the processes in steps Si_1 to Si_5 are not performed on all blocks included in a slice; if they are performed on some blocks, the inter prediction using FRUC mode for that slice may be terminated. Similarly, if the processes in steps Si_1 to Si_5 are performed on some blocks included in a picture, the inter prediction using FRUC mode for that picture may be terminated.
[0474] Subblocks may be processed in the same way as blocks as described above.
[0475] The evaluation value may be calculated by various methods. For example, the reconstructed image of a region in the reference picture corresponding to the MV may be compared with the reconstructed image of a predetermined region (which may be, for example, a region in another reference picture or a region in an adjacent block of the current picture, as shown below). The difference in pixel values between the two reconstructed images may then be calculated and used as the evaluation value for the MV. In addition to the difference value, other information may also be used to calculate the evaluation value.
[0476] Next, we will explain pattern matching in detail. First, one candidate MV included in the candidate MV list (also called the merge list or merge candidate list) is selected as the starting point for the search using pattern matching.
[0477] For pattern matching, either first-order pattern matching or second-order pattern matching may be used. First-order pattern matching and second-order pattern matching are sometimes called bilateral matching and template matching, respectively.
[0478] [MV Derivation > FRUC > Bilateral Matching] In the first pattern matching, pattern matching is performed between two blocks in two different reference pictures that are aligned with the motion trajectory of the current block. Therefore, in the first pattern matching, a region in the other reference picture aligned with the motion trajectory of the current block is used as a predetermined region for calculating the evaluation value of the candidate MV described above.
[0479] Figure 47 illustrates an example of first pattern matching (bilateral matching) between two blocks in two reference pictures along a motion trajectory. As shown in Figure 47, in first pattern matching, two MVs (MV0, MV1) are derived by searching for the most matching pair of two blocks in two different reference pictures (Ref0, Ref1) that are along the motion trajectory of the current block (Cur block).
[0480] Specifically, for the current block, the difference is derived between the reconstructed image at a specified position in the first encoded reference picture (Ref0) specified by the candidate MV and the reconstructed image at a specified position in the second encoded reference picture (Ref1) specified by the symmetric MV obtained by scaling the candidate MV by the display time interval. An evaluation value is then calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs should be selected as the best candidate MV.
[0481] Under the assumption of a continuous motion trajectory, the MV (MV0, MV1) pointing to two reference blocks is proportional to the temporal distance (TD0, TD1) between the current picture (Cur Pic) and the two reference pictures (Ref0, Ref1). For example, if the current picture is temporally located between the two reference pictures and the temporal distances from the current picture to the two reference pictures are equal, then the first pattern matching derives a mirror-symmetric bidirectional MV.
[0482] [MV Derivation > FRUC > Template Matching] In the second pattern matching (template matching), pattern matching is performed between the template in the current picture (blocks adjacent to the current block in the current picture (e.g., blocks above and / or to the left)) and the block in the reference picture. Therefore, in the second pattern matching, the block adjacent to the current block in the current picture is used as a predetermined area for calculating the evaluation value of the candidate MV described above.
[0483] Figure 48 illustrates an example of pattern matching (template matching) between a template in the current picture and a block in the reference picture. As shown in Figure 48, in the second pattern matching, the MV of the current block is derived by searching in the reference picture (Ref0) for the block that best matches the block adjacent to the current block (Cur block) in the current picture (Cur Pic).
[0484] Specifically, for the current block, the difference between the reconstructed image of either or both of the left-adjacent and upper-adjacent encoded regions and the reconstructed image at the equivalent position in the encoded reference picture (Ref0) specified by the candidate MV is derived, and an evaluation value is calculated using the obtained difference value. The candidate MV with the best evaluation value among multiple candidate MVs should be selected as the best candidate MV.
[0485] Information indicating whether or not to apply such a FRUC mode (e.g., called the FRUC flag) may be signaled at the CU level. Furthermore, if the FRUC mode is applied (e.g., the FRUC flag is true), information indicating the applicable pattern matching method (first pattern matching or second pattern matching) may be signaled at the CU level.
[0486] Furthermore, the signaling of this information is not limited to the CU level, but may be at other levels (e.g., sequence level, picture level, slice level, brick level, CTU level, or subblock level).
[0487] [MV Derivation > Affine Mode] The affine mode is a mode in which the MV is generated using an affine transformation. For example, the MV may be derived on a subblock basis based on the MVs of multiple adjacent blocks. This mode is sometimes called the affine motion compensation prediction mode.
[0488] Figure 49A is a diagram illustrating an example of deriving the MV of a subblock based on the MV of multiple adjacent blocks.
[0489] In Figure 49A, the current block includes, for example, 16 subblocks consisting of 4x4 pixels. Here, the motion vector v of the upper left corner control point of the current block is determined based on the MV of the adjacent blocks. 0 The following is derived, and similarly, the motion vector v of the upper right corner control point of the current block based on the MV of the adjacent subblock. 1 The following is derived. Then, by the following equation (A), the two motion vectors v 0 and v 1 Project the motion vector (v) of each subblock within the current block. x ,v y ) is derived.
[0490]
[0491] Here, x and y represent the horizontal and vertical positions of the subblock, respectively, and w represents a predetermined weighting coefficient.
[0492] Such information indicating an affine mode (for example, called an affine flag) may be signaled at the CU level. However, the signaling of this information indicating an affine mode is not limited to the CU level, and may be at other levels (for example, sequence level, picture level, slice level, brick level, CTU level, or subblock level).
[0493] Furthermore, such affine modes may include several modes with different methods for deriving the MV of the upper-left and upper-right corner control points. For example, there are two affine modes: the affine inter (also called the affine normal inter) mode and the affine merge mode.
[0494] Figure 49B illustrates an example of deriving the subblock unit MV in affine mode using three control points.
[0495] In Figure 49B, the current block includes, for example, 16 subblocks consisting of 4x4 pixels. Here, the motion vector v of the upper left corner control point of the current block is determined based on the MV of the adjacent blocks.0 is derived. Similarly, based on the MVs of adjacent blocks, the motion vector v of the upper-right control point of the current block 1 is derived, and based on the MVs of adjacent blocks, the motion vector v of the lower-left control point of the current block 2 is derived.
[0496] Then, by the following formula (B), the three motion vectors v <00(...)010>, v <00(...)011>and v <00(...)012>are projected to derive the motion vectors (v <00(...)013>, v <00(...)014>[[ID=1...]]8) of each sub-block within the current block.
[0497]
[0498] Here, x and y respectively indicate the horizontal position and vertical position of the sub-block center, and w and h indicate predetermined weight coefficients. w may indicate the width of the current block, and h may indicate the height of the current block. [[ID=...]]
[0499] The affine mode using different numbers of control points (e.g., two and three) may be switched and signaled at the CU level. Note that the information indicating the number of control points of the affine mode used at the CU level may be signaled at other levels (e.g., sequence level, picture level, slice level, block level, CTU level or sub-block level). ]...]]
[0500] Also, the affine mode having such three control points may include several modes with different derivation methods of the MVs of the upper-left, upper-right and lower-left control points. For example, the affine mode having three control points includes two modes, the affine inter-mode and the affine merge mode, similar to the affine mode having the above-mentioned two control points.
[0501] Note that in the affine mode, the size of each sub-block included in the current block is not limited to 4x4 pixels and may be other sizes. For example, the size of each sub-block may be 8×8 pixels.
[0502] [MV Derivation > Affine Mode > Control Point] Figures 50A, 50B, and 50C are conceptual diagrams illustrating an example of MV derivation for a control point in affine mode.
[0503] In affine mode, as shown in Figure 50A, for example, the predicted MV for each control point of the current block is calculated based on multiple MVs corresponding to blocks encoded in affine mode from among the encoded blocks A (left), B (top), C (upper right), D (lower left), and E (upper left) adjacent to the current block.
[0504] Specifically, these blocks are examined in the order of encoded block A (left), block B (top), block C (top right), block D (bottom left), and block E (top left), and the first valid block encoded in affine mode is identified. Based on the multiple MVs corresponding to this identified block, the MV of the control point of the current block is calculated.
[0505] For example, as shown in Figure 50B, if block A adjacent to the left of the current block is encoded in affine mode with two control points, the motion vector v projected onto the upper left and upper right corners of the encoded block containing block A is... 3 and v 4 The following is derived. And the derived motion vector v 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The result is calculated.
[0506] For example, as shown in Figure 50C, if block A adjacent to the left of the current block is encoded in affine mode with three control points, then the motion vector v projected onto the upper left, upper right, and lower left corners of the encoded block containing block A is... 3 , v 4 and v 5 The following is derived. And the derived motion vector v3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 The result is calculated.
[0507] The MV derivation method shown in Figures 50A to 50C may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.
[0508] Figures 51A and 51B are conceptual diagrams illustrating another example of the derivation of the control point MV in affine mode.
[0509] Figure 51A is a diagram illustrating an affine mode having two control points.
[0510] In this affine mode, as shown in Figure 51A, the MV selected from the respective MVs of the encoded blocks A, B, and C adjacent to the current block is the motion vector v of the upper left corner control point of the current block. 0 It is used as follows. Similarly, the MV selected from the respective MVs of the encoded blocks D and E adjacent to the current block is used as the motion vector v of the upper right corner control point of the current block. 1 It is used as such.
[0511] Figure 51B is a diagram illustrating an affine mode having three control points.
[0512] In this affine mode, as shown in Figure 51B, the MV selected from the respective MVs of the encoded blocks A, B, and C adjacent to the current block is the motion vector v of the upper left corner control point of the current block. 0 It is used as such.
[0513] Similarly, the MV selected from the respective MVs of the encoded blocks D and E adjacent to the current block is the motion vector v of the upper right corner control point of the current block. 1 It is used as follows. Furthermore, the MV selected from the respective MVs of the encoded blocks F and G adjacent to the current block is used as the motion vector v of the lower left corner control point of the current block. 2 It is used as such.
[0514] The MV derivation method shown in Figures 51A and 51B may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.
[0515] Here, for example, when switching between affine modes with different numbers of control points (e.g., two and three) at the CU level to create a signal, the number of control points may differ between the encoded block and the current block.
[0516] Figures 52A and 52B are conceptual diagrams illustrating an example of a method for deriving the MV of control points when the number of control points differs between the encoded block and the current block.
[0517] For example, as shown in Figure 52A, the current block is encoded in affine mode, having three control points at the upper left corner, upper right corner, and lower left corner, while block A, adjacent to the left of the current block, has two control points.
[0518] In this case, the motion vector v is projected onto the upper-left and upper-right corners of the encoded block containing block A. 3 and v 4 The following is derived. And the derived motion vector v 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The following is calculated. Furthermore, the derived motion vector v 0and v 1 From there, the motion vector v of the lower left corner control point. 2 This is calculated.
[0519] For example, as shown in Figure 52B, the current block is encoded in affine mode, having two control points at the upper left corner and the upper right corner, and block A adjacent to the left of the current block has three control points.
[0520] In this case, the motion vector v is projected onto the upper-left, upper-right, and lower-left corners of the encoded block containing block A. 3 , v 4 and v 5 The following is derived. And the derived motion vector v 3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 The result is calculated.
[0521] The MV derivation method shown in Figures 52A and 52B may be used to derive the MV of each control point in the current block in step Sk_1 shown in Figure 53, which will be described later, or it may be used to derive the predicted MV of each control point in the current block in step Sj_1 shown in Figure 54, which will be described later.
[0522] [MV Derivation > Affine Mode > Affine Merge Mode] Figure 53 is a flowchart showing an example of the affine merge mode.
[0523] In affine merge mode, the interpretation unit 126 first derives the MV for each control point of the current block (step Sk_1). The control points are the upper left and upper right corners of the current block, as shown in Figure 49A, or the upper left, upper right, and lower left corners of the current block, as shown in Figure 49B. At this time, the interpretation unit 126 may encode MV selection information into a stream to identify two or three of the derived MVs.
[0524] For example, when using the MV derivation method shown in Figures 50A to 50C, the interpretation unit 126 examines the encoded blocks in the order of A (left), B (top), C (upper right), D (lower left), and E (upper left), as shown in Figure 50A, and identifies the first valid block encoded in affine mode.
[0525] The interpretation unit 126 derives the MV of the control point using the first valid block encoded in the identified affine mode. For example, if block A is identified and block A has two control points, as shown in Figure 50B, the interpretation unit 126 derives the motion vectors v of the upper left and upper right corners of the encoded block containing block A. 3 and v 4 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 Calculate the result.
[0526] For example, the interpretation unit 126 generates motion vectors v of the upper left and upper right corners of the encoded block. 3 and v 4 By projecting this onto the current block, the motion vector v of the upper left corner control point of the current block is obtained. 0 And the motion vector v of the upper right corner control point 1 Calculate the result.
[0527] Alternatively, if block A is identified and block A has three control points, as shown in Figure 50C, the interpretation unit 126 predicts the motion vectors v of the upper left corner, upper right corner, and lower left corner of the encoded block containing block A. 3 , v 4 and v 5 From there, the motion vector v of the control point at the upper left corner of the current block. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 Calculate the result.
[0528] For example, the interpretation unit 126 generates motion vectors v of the upper left corner, upper right corner, and lower left corner of the encoded block. 3 , v 4 and v 5 By projecting this onto the current block, the motion vector v of the upper left corner control point of the current block is obtained. 0 And the motion vector v of the upper right corner control point 1 And the motion vector v of the lower left corner control point 2 Calculate the result.
[0529] Furthermore, as shown in Figure 52A above, if block A is identified and block A has two control points, the MV of three control points may be calculated. Alternatively, as shown in Figure 52B above, if block A is identified and block A has three control points, the MV of two control points may be calculated.
[0530] Next, the interpretation unit 126 performs motion compensation for each of the multiple subblocks included in the current block. That is, the interpretation unit 126 calculates two motion vectors v for each of the multiple subblocks. 0 and v 1 Using the above equation (A), or the three motion vectors v 0 , v 1 and v 2 Using the above formula (B), the MV of the subblock is calculated as the affine MV (step Sk_2).
[0531] Then, the interpretation unit 126 performs motion compensation on the subblock using the affine MV and encoded reference picture (step Sk_3). Once steps Sk_2 and Sk_3 have been executed for each of the subblocks contained in the current block, the process of generating a predicted image using the affine merge mode for the current block is completed. In other words, motion compensation is performed on the current block, and a predicted image for the current block is generated.
[0532] In step Sk_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived for each control point using multiple MV derivation methods. The multiple MV derivation methods may be any combination of the MV derivation methods shown in Figures 50A to 50C, the MV derivation methods shown in Figures 51A and 51B, the MV derivation methods shown in Figures 52A and 52B, and other MV derivation methods.
[0533] The candidate MV list may also include candidate MVs for modes other than affine mode, where prediction is performed on a sub-block basis.
[0534] Furthermore, the candidate MV list may include, for example, a candidate MV for an affine merge mode having two control points and a candidate MV for an affine merge mode having three control points.
[0535] Alternatively, a candidate MV list may be generated containing candidate MVs for an affine merge mode having two control points, and a candidate MV list may be generated containing candidate MVs for an affine merge mode having three control points. Alternatively, a candidate MV list may be generated containing candidate MVs for one of the modes: an affine merge mode having two control points or an affine merge mode having three control points.
[0536] Candidate MVs may be, for example, the MVs of encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or they may be the MVs of any valid block among those blocks.
[0537] Alternatively, you may send an index indicating which candidate MV from the candidate MV list is being selected as MV selection information.
[0538] [MV Derivation > Affine Mode > Affine Intermode] Figure 54 is a flowchart showing an example of an affine intermode.
[0539] In affine intermode, first, the inter prediction unit 126 predicts the MV (v) of each of the two or three control points of the current block. 0 ,v 1 ) or (v 0 ,v 1 ,v 2 Derive the formula (step Sj_1). The control points are the upper left corner, upper right corner, or lower left corner of the current block, as shown in Figure 49A or Figure 49B.
[0540] For example, when using the MV derivation method shown in Figures 51A and 51B, the interpretation unit 126 selects the MV of any of the encoded blocks near each control point of the current block shown in Figure 51A or Figure 51B, thereby predicting the MV (v) of the control point of the current block. 0 ,v 1 ) or (v 0 ,v 1 ,v 2 The following is derived: At this time, the interpretation unit 126 encodes prediction MV selection information into a stream to identify two or three selected prediction MVs.
[0541] For example, the interpretation unit 126 may determine which block's MV from the encoded blocks adjacent to the current block to select as the predicted MV for the control point using cost evaluation or the like, and write a flag indicating which predicted MV was selected to the bitstream. In other words, the interpretation unit 126 outputs the predicted MV selection information, such as a flag, as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0542] Next, the interpretation unit 126 performs motion search (steps Sj_3 and Sj_4) while updating the predicted MV selected or derived in step Sj_1 (step Sj_2).
[0543] In other words, the interpretation unit 126 calculates the MV of each subblock corresponding to the updated predicted MV as an affine MV using the above-described equation (A) or equation (B) (step Sj_3). Then, the interpretation unit 126 performs motion compensation for each subblock using these affine MVs and encoded reference pictures (step Sj_4). The processes in steps Sj_3 and Sj_4 are performed for all blocks in the current block each time the predicted MV is updated in step Sj_2.
[0544] As a result, the interpretation unit 126 determines, for example, the predicted MV that yields the smallest cost in the motion search loop as the MV of the control point (step Sj_5). At this time, the interpretation unit 126 further encodes the difference between the determined MV and the predicted MV as the difference MV into the stream. In other words, the interpretation unit 126 outputs the difference MV as a prediction parameter to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0545] Finally, the interpretation unit 126 generates a predicted image of the current block by performing motion compensation on the current block using the determined MV and the encoded reference picture (step Sj_6).
[0546] In step Sj_1, the above-described candidate MV list may be generated. The candidate MV list may be, for example, a list containing candidate MVs derived for each control point using multiple MV derivation methods. The multiple MV derivation methods may be any combination of the MV derivation methods shown in Figures 50A to 50C, the MV derivation methods shown in Figures 51A and 51B, the MV derivation methods shown in Figures 52A and 52B, and other MV derivation methods.
[0547] The candidate MV list may also include candidate MVs for modes other than affine mode, where prediction is performed on a sub-block basis.
[0548] Furthermore, a candidate MV list may be generated that includes candidate MVs for affine intermodes having two control points and candidate MVs for affine intermodes having three control points.
[0549] Alternatively, a candidate MV list may be generated containing candidate MVs for affine intermodes having two control points, and a candidate MV list may be generated containing candidate MVs for affine intermodes having three control points. Alternatively, a candidate MV list may be generated containing candidate MVs for one of the modes: an affine intermode with two control points or an affine intermode with three control points.
[0550] Candidate MVs may be, for example, the MVs of encoded blocks A (left), B (top), C (top right), D (bottom left), and E (top left), or they may be the MVs of any valid block among those blocks.
[0551] Additionally, an index indicating which candidate MV from the candidate MV list is being sent as predicted MV selection information.
[0552] [MV Derivation > GPM] In the above example, the interpretation unit 126 generates a prediction image for a rectangular current block. However, motion compensation may be performed for two regions defined by diagonal lines within the current block. These regions are also called partitions. This operating mode is called GPM (Geometric Partitioning Mode). GPM may also be expressed as geometric shape partitioning prediction mode or GPM mode, etc.
[0553] Figure 55A shows examples of two regions where a predicted image is generated by GPM. As shown in Figure 55A, GPM makes it possible to efficiently predict images with oblique object boundaries while maintaining a large CU size.
[0554] Figure 55B is a conceptual diagram showing the pattern of two regions defined in the GPM. The pattern of the two regions (first partition and second partition) defined in the current block in the GPM is indicated by an index that shows the angle of the lines and the distance from the center. There are 64 possible patterns, for example, as shown in Figure 55B. In addition, the shape of the partition may be a triangle, trapezoid, pentagon, or rectangle, etc.
[0555] The interpretation unit 126 specifies one single prediction motion vector each for the first partition and the second partition (the first MV and the second MV in Figure 55A).
[0556] Specifically, candidates are selected by index from a list of merge candidates for dual predictions, similar to the normal merge mode. If the index is even, a motion vector candidate from the L0 list is selected. If the index is odd, a motion vector candidate from the L1 list is selected. If no motion vector candidate to be selected exists in either the L0 or L1 list, a motion vector candidate from the other list is selected.
[0557] Furthermore, when deriving motion vectors for each of the two partitions, the operation is restricted to prevent the same motion vector candidate from being selected. Finally, motion compensation is performed on the first and second partitions using the first and second selected motion vectors, respectively.
[0558] The interpretation unit 126 then generates a predicted image for the current block by weighting near the boundary between the first and second partitions. Specifically, as shown in Figure 55C, it generates a predicted image for the current block by performing a weighted average so that the pixel values of the two partitions are gradually mixed at the boundary.
[0559] Figure 55C is a conceptual diagram showing the weighted average of pixel values at region boundaries. For example, when generating a predicted image of the current block, the pixel labeled "8" in Figure 55C may use only the pixel value based on the first partition, and the pixel labeled "0" may use only the pixel value from the second partition. In that case, for pixels "1" through "7" in Figure 55C, the pixel values from the two partitions are mixed while varying the weight for each pixel position.
[0560] Furthermore, the two motion vectors used in GPM are stored for each subblock, which is determined by dividing the current block into 4x4 subblocks. Specifically, depending on whether the 4x4 subblock belongs to partition 1 only, partition 2 only, or both partition 1 and partition 2 (i.e., near the boundary), one or two motion vectors corresponding to each partition are stored.
[0561] Furthermore, in the case of a subblock located near a boundary, MV1 and MV2 are stored as dual predictive motion vectors. In this case, if both MV1 and MV2 are motion vectors selected from the same list, they are not stored as dual predictive motion vectors, and only one of them is stored as the motion vector for that subblock.
[0562] Figure 56 is a flowchart showing an example of the GPM mode.
[0563] In GPM mode, the interpretation unit 126 first divides the current block into a first partition and a second partition (step Sx_1). At this time, the interpretation unit 126 may encode partition information, which is information regarding the division into each partition, into the stream as prediction parameters. In other words, the interpretation unit 126 may output partition information as prediction parameters to the entropy coding unit 110 via the prediction parameter generation unit 130.
[0564] Next, the interpretation unit 126 first obtains multiple candidate MVs for the current block based on information such as the MVs of multiple encoded blocks surrounding the current block in time or space (step Sx_2). In other words, the interpretation unit 126 creates a candidate MV list.
[0565] Then, the interpretation unit 126 selects the candidate MV for the first partition and the candidate MV for the second partition from among the multiple candidate MVs obtained in step Sx_2 as the first MV and the second MV, respectively (step Sx_3).
[0566] In this case, the interpretation unit 126 may encode MV selection information for identifying the selected candidate MV as prediction parameters into the stream. In other words, the interpretation unit 126 may output MV selection information as prediction parameters to the entropy encoding unit 110 via the prediction parameter generation unit 130.
[0567] Next, the interpretation unit 126 generates a first predicted image by performing motion compensation using the selected first MV and the encoded reference picture (step Sx_4). Similarly, the interpretation unit 126 generates a second predicted image by performing motion compensation using the selected second MV and the encoded reference picture (step Sx_5).
[0568] Finally, the interpretation unit 126 generates a predicted image of the current block by weighting and adding the first predicted image and the second predicted image (step Sx_6).
[0569] [MV Derivation > ATMVP Mode] Figure 57 shows an example of ATMVP mode in which MV is derived on a subblock basis.
[0570] ATMVP mode is a mode classified as a merge mode. For example, in ATMVP mode, candidate MVs at the subblock level are registered in the candidate MV list used in normal merge mode.
[0571] Specifically, in ATMVP mode, first, as shown in Figure 57, the time MV reference block associated with the current block is identified in the encoded reference picture specified by the MV (MV0) of the block adjacent to the lower left of the current block. Next, for each subblock within the current block, the MV used when encoding the area corresponding to that subblock within the time MV reference block is identified.
[0572] The identified motion video (MV) is then included in the candidate MV list as a candidate MV for the subblock of the current block. When a candidate MV for a subblock is selected from the candidate MV list, motion compensation using that candidate MV is performed for that subblock. This generates a predicted image for each subblock.
[0573] In the example shown in Figure 57, the block adjacent to the lower left of the current block was used as the peripheral MV reference block, but other blocks may be used. Also, the size of the subblock may be 4x4 pixels, 8x8 pixels, or other sizes. The size of the subblock may be switched in units such as slices, bricks, or pictures.
[0574] [Motion Search > DMVR] Figure 58 shows the relationship between merge mode and DMVR.
[0575] The interpretation unit 126 derives the MV of the current block in merge mode (step Sl_1). Next, the interpretation unit 126 determines whether or not to perform an MV search, i.e., a motion search (step Sl_2). If the interpretation unit 126 determines not to perform a motion search (No. in step Sl_2), it determines the MV derived in step Sl_1 as the final MV for the current block (step Sl_4). In other words, in this case, the MV of the current block is determined in merge mode.
[0576] On the other hand, if it is determined in step Sl_2 to perform a motion search (Yes in step Sl_2), the interpretation unit 126 derives the final MV for the current block by searching the surrounding region of the reference picture indicated by the MV derived in step Sl_1 (step Sl_3). In other words, in this case, the MV of the current block is determined by the DMVR.
[0577] Figure 59 is a conceptual diagram illustrating an example of a DMVR for determining MV.
[0578] First, for example in merge mode, candidate MVs (L0 and L1) are selected for the current block. Then, according to the candidate MV (L0), reference pixels are identified from the first reference picture (L0), which is an encoded picture in the L0 list. Similarly, according to the candidate MV (L1), reference pixels are identified from the second reference picture (L1), which is an encoded picture in the L1 list. A template is generated by taking the average of these reference pixels.
[0579] Next, using the template, the surrounding regions of candidate MVs for the first reference picture (L0) and the second reference picture (L1) are searched, and the MV with the minimum cost is determined as the final MV for the current block. The cost may be calculated, for example, using the difference between each pixel value in the template and each pixel value in the search region, as well as the candidate MV value.
[0580] Any process that can explore the vicinity of candidate MVs and derive the final MV can be used, even if it is not the exact process described here.
[0581] Figure 60 is a conceptual diagram illustrating another example of a DMVR for determining MV. Unlike the example of a DMVR shown in Figure 59, this example in Figure 60 calculates costs without generating a template.
[0582] First, the interpretation unit 126 searches around the reference blocks contained in the reference pictures in the L0 list and L1 list, respectively, based on the initial MV, which is a candidate MV obtained from the candidate MV list. For example, as shown in Figure 60, the initial MV corresponding to the reference block in the L0 list is InitMV_L0, and the initial MV corresponding to the reference block in the L1 list is InitMV_L1.
[0583] In motion search, the interpretation unit 126 first sets a search position relative to the reference picture in the L0 list. The difference vector indicating the set search position, specifically the difference vector from the position indicated by the initial MV (i.e., InitMV_L0) to that search position, is MVd_L0.
[0584] The interpretation unit 126 then determines the search position in the reference picture of the L1 list. This search position is indicated by a difference vector from the position indicated by the initial MV (i.e., InitMV_L1) to the search position. Specifically, the interpretation unit 126 determines the difference vector as MVd_L1 by mirroring MVd_L0. In other words, the interpretation unit 126 sets the search position to a position that is symmetrical to the position indicated by the initial MV in the reference pictures of the L0 list and the L1 list.
[0585] The interpretation unit 126 calculates a cost for each search location, such as the sum of the absolute differences in pixel values within the block at that search location (SAD), and finds the search location that minimizes this cost.
[0586] Figure 61A is a diagram showing an example of motion search in a DMVR, and Figure 61B is a flowchart showing that same motion search example.
[0587] First, in Step 1, the interpretation unit 126 calculates the cost of the search position indicated by the initial MV (also called the starting point) and the eight surrounding search positions. Then, the interpretation unit 126 determines whether the cost of the search positions other than the starting point is the minimum. If the interpretation unit 126 determines that the cost of the search positions other than the starting point is the minimum, it moves to the search position with the minimum cost and proceeds to Step 2. On the other hand, if the cost of the starting point is the minimum, the interpretation unit 126 skips Step 2 and proceeds to Step 3.
[0588] In Step 2, the interpretation unit 126 uses the search position moved according to the processing result of Step 1 as a new starting point and performs a search similar to the processing in Step 1. The interpretation unit 126 then determines whether the cost of the search positions other than the starting point is the minimum. If the interpretation unit 126 determines that the cost of the search positions other than the starting point is the minimum, it proceeds to Step 4. On the other hand, if the interpretation unit 126 determines that the cost of the starting point is the minimum, it proceeds to Step 3.
[0589] In Step 4, the interpretation unit 126 treats the starting point's search position as the final search position, and determines the difference between the position indicated by the initial MV and the final search position as the difference vector.
[0590] In Step 3, the interpretation unit 126 determines the pixel position with the minimum cost based on the costs at the four points above, below, to the left and right of the starting point of Step 1 or Step 2, and sets that pixel position as the final search position. This decimal-precision pixel position is determined by weighting the vector of the four points above, below, to the left and right ((0,1), (0,-1), (-1,0), (1,0)) with the cost at each of the four search positions as the weight. The interpretation unit 126 then determines the difference between the position indicated by the initial MV and the final search position as the difference vector.
[0591] [Motion Compensation > BIO / OBMC / LIC] Motion compensation includes modes that generate a predictive image and then correct that predictive image. These modes include, for example, BIO, OBMC, and LIC, which are described below.
[0592] Figure 62 is a flowchart showing an example of predictive image generation.
[0593] The interpretation unit 126 generates a predicted image (step Sm_1) and corrects the predicted image according to one of the above modes (step Sm_2).
[0594] Figure 63 is a flowchart showing another example of predictive image generation.
[0595] The interpretation unit 126 derives the MV of the current block (step Sn_1). Next, the interpretation unit 126 generates a predicted image using the MV (step Sn_2) and determines whether or not to perform correction processing (step Sn_3). If the interpretation unit 126 determines that correction processing should be performed (Yes in step Sn_3), it generates the final predicted image by correcting the predicted image (step Sn_4).
[0596] In the LIC described later, brightness and color difference may be corrected in step Sn_4. On the other hand, if the interpretation unit 126 determines that no correction processing is necessary (No in step Sn_3), it outputs the predicted image as the final predicted image without correction (step Sn_5).
[0597] [Motion Compensation > OBMC] Interpretation images may be generated using not only the motion information of the current block obtained by motion search, but also the motion information of adjacent blocks. Specifically, interpretation images may be generated on a subblock basis within the current block by weighting and adding together a prediction image based on motion information obtained by motion search (within the reference picture) and a prediction image based on motion information of adjacent blocks (within the current picture).
[0598] Such interpretation (motion compensation) is sometimes called OBMC (overlapped block motion compensation) or OBMC mode.
[0599] In OBMC mode, information indicating the size of the subblock for OBMC (e.g., called the OBMC block size) may be signaled at the sequence level. Furthermore, information indicating whether or not to apply OBMC mode (e.g., called the OBMC flag) may be signaled at the CU level. Note that the signaling levels of this information are not limited to the sequence level and CU level, but may be other levels (e.g., picture level, slice level, brick level, CTU level, or subblock level).
[0600] The OBMC mode will be explained in more detail. Figures 64 and 65 are flowcharts and conceptual diagrams illustrating the overview of the predictive image correction process using OBMC.
[0601] First, as shown in Figure 65, a predicted image (Pred) is obtained using normal motion compensation with the MV assigned to the current block. In Figure 65, the arrow "MV" points to the reference picture, indicating what the current block of the current picture is referencing in order to obtain the predicted image.
[0602] Next, the previously derived MV (MV_L) for the encoded left-adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_L). The MV (MV_L) is indicated by an arrow “MV_L” pointing from the current block to the reference picture. Then, the first correction of the predicted image is performed by superimposing the two predicted images, Pred and Pred_L. This has the effect of blending the boundaries between adjacent blocks.
[0603] Similarly, the previously derived MV (MV_U) for the encoded upper adjacent block is applied (reused) to the current block to obtain the predicted image (Pred_U). The MV (MV_U) is indicated by an arrow “MV_U” pointing from the current block to the reference picture. Then, the predicted image Pred_U is superimposed onto the predicted image that has undergone the first correction (for example, Pred and Pred_L) to perform a second correction of the predicted image.
[0604] This has the effect of blending the boundaries between adjacent blocks. The predicted image obtained by the second correction is the final predicted image of the current block, with the boundaries with adjacent blocks blended (smoothed).
[0605] The above example is a two-pass correction method using left-adjacent and top-adjacent blocks, but the correction method may also be a three-pass or more-pass correction method using right-adjacent and / or bottom-adjacent blocks.
[0606] Furthermore, the area to be superimposed does not have to be the entire pixel area of the block, but rather only a portion of the area near the block boundary.
[0607] This section describes OBMC's predictive image correction process for obtaining a single predictive image Pred by superimposing additional predictive images Pred_L and Pred_U onto a single reference picture.
[0608] However, if the predicted image is corrected based on multiple reference images, the same process may be applied to each of the multiple reference pictures. In such cases, by performing OBMC image correction based on multiple reference pictures, a corrected predicted image is obtained from each reference picture, and then the multiple corrected predicted images obtained are further superimposed to obtain the final predicted image.
[0609] In OBMC, the unit of a current block may be a PU unit, or it may be a subblock unit obtained by further dividing a PU.
[0610] One method for determining whether or not to apply OBMC is to use the obmc_flag signal, which indicates whether or not OBMC should be applied.
[0611] As a specific example, the encoding device 100 may determine whether the current block belongs to a region with complex motion. When the encoding device 100 determines that the current block belongs to a region with complex motion, it sets the value 1 as obmc_flag and applies OBMC to perform encoding. When the current block does not belong to a region with complex motion, it sets the value 0 as obmc_flag and performs block encoding without applying OBMC.
[0612] On the other hand, in the decoding device 200, by decoding the obmc_flag described in the stream, decoding is performed by switching whether to apply OBMC according to the value.
[0613] [Motion Compensation > BIO] Next, a method for deriving the MV will be described. First, a mode for deriving the MV based on a model assuming a constant linear motion will be described. This mode is sometimes called the BIO (bi-directional optical flow) mode. Also, this bi-directional optical flow may be denoted as BDOF instead of BIO.
[0614] FIG. 66 is a diagram for explaining a model assuming a constant linear motion. In FIG. 66, (v x , v y ) indicates the velocity vector, and τ 0 , τ 1 respectively indicate the temporal distances between the current picture (Cur Pic) and two reference pictures (Ref 0 , Ref 1 ). (MVx 0 , MVy 0 ) indicates the MV corresponding to the reference picture Ref 0 , and (MVx 1 , MVy 1 ) indicates the MV corresponding to the reference picture Ref 1 . <00C1387>At this time, under the assumption of the constant linear motion of the velocity vector (v x , v y ), (MVx 0 , MVy 0 ) and (MVx 1, MVy 1 ) are, respectively, (v x τ 0 ,v y τ 0 ) and (-v x τ 1 , -v y τ 1 This can be expressed as follows, and the following optical flow equation holds:
[0616]
[0617] Here, I(k) represents the luminance value of the reference image k (k=0,1) after motion compensation. This optical flow equation shows that the sum of (i) the time derivative of the luminance value, (ii) the product of the horizontal velocity and the horizontal component of the spatial gradient of the reference image, and (iii) the product of the vertical velocity and the vertical component of the spatial gradient of the reference image is equal to zero. Based on this optical flow equation and Hermitian interpolation, block-level motion vectors obtained from a candidate MV list, etc., may be corrected on a pixel-by-pixel basis.
[0618] Furthermore, the motion vector (MV) may be derived by the decoding device 200 using a method different from that used to derive the motion vector based on a model assuming uniform linear motion. For example, the motion vector may be derived on a sub-block basis based on the MVs of multiple adjacent blocks.
[0619] Figure 67 is a flowchart illustrating an example of inter prediction according to BIO. Figure 68 is a diagram illustrating an example of the configuration of the inter prediction unit 126 that performs the inter prediction according to BIO.
[0620] As shown in Figure 68, the interpretation unit 126 includes, for example, a memory 126a, an interpolation image derivation unit 126b, a gradient image derivation unit 126c, an optical flow derivation unit 126d, a correction value derivation unit 126e, and a prediction image correction unit 126f. Note that the memory 126a may be a frame memory 122.
[0621] The interpretation unit 126 uses two different reference pictures (Ref) from the picture (Cur Pic) that contains the current block. 0 ,Ref1 Using this, two motion vectors (M0, M1) are derived. Then, the interpretation unit 126 uses these two motion vectors (M0, M1) to derive a predicted image of the current block (step Sy_1).
[0622] Note that the motion vector M0 is the reference picture Ref 0 The corresponding motion vector (MVx 0 , MVy 0 ) and the motion vector M1 is the reference picture Ref 1 The corresponding motion vector (MVx 1 , MVy 1 )
[0623] Next, the interpolated image derivation unit 126b refers to the memory 126a and uses the motion vector M0 and the reference picture L0 to create an interpolated image I of the current block. 0 The interpolated image derive unit 126b also references memory 126a and uses motion vector M1 and reference picture L1 to derive the interpolated image I of the current block. 1 Derive the following (step Sy_2).
[0624] Here, interpolated image I 0 This is derived for the current block, the reference picture Ref 0 The image included is interpolated image I 1 This is derived for the current block, the reference picture Ref 1 This is an image included in [the file / location].
[0625] Interpolated image I 0 and interpolated image I 1 Each of these may be the same size as the current block. Alternatively, interpolated image I 0 and interpolated image I 1 Each of these may be a larger image than the current block in order to properly derive the gradient image described later. Furthermore, interpolated image I 0 and I 1 This may include a motion vector (M0, M1) and a reference picture (L0, L1), and a predicted image derived by applying a motion compensation filter.
[0626] Furthermore, the gradient image derivation unit 126c generates an interpolated image I 0 and interpolated image I 1 From there, the gradient image of the current block (Ix 0 , Ix 1 , Iy 0 , Iy 1 ) is derived (step Sy_3). Note that the horizontal gradient image is (Ix 0 , Ix 1 ) and the vertical gradient image is (Iy 0 , Iy 1 The gradient image derivation unit 126c may derive the gradient image by, for example, applying a gradient filter to the interpolated image. The gradient image only needs to show the amount of spatial change in pixel values along the horizontal or vertical direction.
[0627] Next, the optical flow derivation unit 126d interpolates the image (I) in units of multiple subblocks that constitute the current block. 0 , I 1 ) and gradient image (Ix 0 , Ix 1 , Iy 0 , Iy 1 Using the above velocity vector, optical flow (v x ,v y Derive the following (step Sy_4).
[0628] Optical flow is a coefficient that corrects the spatial displacement of pixels, and may also be called a local motion estimate, corrected motion vector, or corrected weight vector. For example, a subblock may be a 4x4 pixel subCU. Note that the derivation of optical flow may be performed at other units, such as pixel units, rather than subblock units.
[0629] Next, the interpretation unit 126 calculates the optical flow (v x ,v y The predicted image of the current block is corrected using ). For example, the correction value derivation unit 126e uses optical flow (v x ,v yThe correction value for the pixel values included in the current block is derived using (step Sy_5). The predicted image correction unit 126f may then correct the predicted image of the current block using the correction value (step Sy_6). The correction value may be derived for each pixel, or for multiple pixels or subblocks.
[0630] Note that the BIO processing flow is not limited to the processing disclosed in Figure 67. Only a portion of the processing disclosed in Figure 67 may be performed, different processing may be added or replaced, or the processing may be executed in a different order.
[0631] [Motion Compensation > LIC] Next, we will explain an example of a mode that generates a predicted image (prediction) using LIC (local illumination compensation).
[0632] Figure 69A is a diagram illustrating an example of a predictive image generation method using brightness correction processing by an LIC (Luminance Injector). Figure 69B is a flowchart illustrating an example of the predictive image generation method using the LIC.
[0633] First, the interpretation unit 126 derives MV from the encoded reference picture and obtains the reference image corresponding to the current block (step Sz_1).
[0634] Next, the interpretation unit 126 extracts information for the current block indicating how the luminance values have changed between the reference picture and the current picture (step Sz_2). This extraction is performed based on the luminance pixel values of the encoded left adjacent reference region (peripheral reference region) and the encoded upper adjacent reference region (peripheral reference region) in the current picture, and the luminance pixel values at the equivalent positions in the reference picture specified by the derived MV.
[0635] Then, the interpretation unit 126 calculates a brightness correction parameter using information indicating how the brightness value has changed (step Sz_3).
[0636] The interpretation unit 126 generates a predicted image for the current block by performing a brightness correction process that applies its brightness correction parameter to the reference image in the reference picture specified by MV (step Sz_4).
[0637] In other words, the predicted image, which is a reference image within the reference picture specified in MV, is corrected based on luminance correction parameters. This correction may involve correcting luminance or color difference. Specifically, color difference correction parameters may be calculated using information indicating how the color difference has changed, and the color difference correction process may be performed.
[0638] Note that the shape of the peripheral reference region in Figure 69A is just one example, and other shapes may be used.
[0639] Furthermore, although this explanation describes the process of generating a predicted image from a single reference picture, the process is similar when generating predicted images from multiple reference pictures. Alternatively, the brightness correction process may be applied to each reference picture obtained from the reference picture in the same manner as described above before generating the predicted image.
[0640] One method for determining whether to apply LIC is to use a signal called lic_flag, which indicates whether or not to apply LIC. As a specific example, the encoding device 100 determines whether the current block belongs to a region where a brightness change is occurring. If it belongs to a region where a brightness change is occurring, it sets lic_flag to a value of 1 and applies LIC to perform encoding. If it does not belong to a region where a brightness change is occurring, it sets lic_flag to a value of 0 and performs encoding without applying LIC.
[0641] On the other hand, the decoding device 200 may decode the lic_flag described in the stream and then switch whether or not to apply LIC depending on its value before performing the decoding.
[0642] Another way to determine whether to apply LIC is, for example, by checking whether LIC has been applied in surrounding blocks.
[0643] As a specific example, when the current block is being processed in merge mode, the interpretation unit 126 determines whether the surrounding encoded blocks selected during the MV derivation in merge mode were encoded with LIC applied. The interpretation unit 126 then switches whether to apply LIC and performs encoding based on the result. In this example as well, the same process is applied to the decoding device 200.
[0644] The LIC (Luminance Correction Processing) was explained using Figures 69A and 69B, but its details will be explained below.
[0645] First, the interpretation unit 126 derives an MV for obtaining the reference image corresponding to the current block from the reference picture, which is an encoded picture.
[0646] Next, the interpretation unit 126 extracts information indicating how the luminance values have changed between the reference picture and the current picture, using the luminance pixel values of the left-adjacent and upper-adjacent encoded peripheral reference regions and the luminance pixel values at equivalent positions in the reference picture specified by MV, and calculates a luminance correction parameter.
[0647] For example, let p0 be the luminance pixel value of a pixel in the peripheral reference area of the current picture, and let p1 be the luminance pixel value of a pixel in the peripheral reference area of the reference picture at the same position as that pixel. The interpretation unit 126 calculates coefficients A and B as luminance correction parameters to optimize A × p1 + B = p0 for multiple pixels in the peripheral reference area.
[0648] Next, the interpretation unit 126 generates a predicted image for the current block by performing a brightness correction process on the reference image in the reference picture specified by MV using a brightness correction parameter. For example, let p2 be the brightness pixel value in the reference image, and p3 be the brightness pixel value of the predicted image after brightness correction processing. The interpretation unit 126 generates the predicted image after brightness correction processing by calculating A × p2 + B = p3 for each pixel in the reference image.
[0649] Furthermore, a portion of the peripheral reference region shown in Figure 69A may be used. For example, a region containing a predetermined number of pixels obtained by thinning out the upper adjacent pixels and the left adjacent pixels may be used as the peripheral reference region. Also, the peripheral reference region is not limited to the region adjacent to the current block, but may also be a region not adjacent to the current block.
[0650] Furthermore, in the example shown in Figure 69A, the peripheral reference area within the reference picture is the area specified by the MV of the current picture, relative to the peripheral reference area within the current picture, but it may also be an area specified by another MV. For example, this other MV may be the MV of the peripheral reference area within the current picture.
[0651] Although the operation of the encoding device 100 has been described here, the operation of the decoding device 200 is similar.
[0652] Furthermore, LIC may be applied not only to luminance but also to color difference. In this case, correction parameters may be derived individually for each of Y, Cb, and Cr, or a common correction parameter may be used for any of them.
[0653] Furthermore, LIC processing may be applied on a subblock basis. For example, correction parameters may be derived using the surrounding reference region of the current subblock and the surrounding reference region of the reference subblock within the reference picture specified by the MV of the current subblock.
[0654] [Generated Pictures] Generated pictures are images obtained through a generation process, unlike regular reference images (in other words, images decoded by intra / inter prediction). Here, as an example, we show how to generate a new reference picture by inputting one or more pictures into a neural network, but the method of generating a new reference picture is not limited to this.
[0655] The generation process may support the application of NN loop filter (NN LF) processing, which includes NN (Neural Network) processing, or it may support transformation processing (scaling, rotation, mirroring, shifting, etc.). Furthermore, the generation process may be a combination of multiple processes (for example, a process that generates a predicted image using an NN such as NN LF, a transformation process, and other processes).
[0656] The new reference picture generated by the generation process is also referred to as a generated reference picture, a generated picture, or a processed picture. Furthermore, such a new reference picture or its image may also be referred to as a generated reference image, a generated image, or a processed image. Adding the generated picture to the reference picture list and making it accessible may improve encoding efficiency.
[0657] First, we will explain the generated reference picture that is generated using the NN. The encoding device 100 and the decoding device 200 input one or more images to the NN and generate a generated picture.
[0658] Figure 70A is a conceptual diagram illustrating an example of a generated picture. As illustrated in Figure 70A, a new image (PicA') generated by inputting an image (PicA) into the NN may be used as a reference picture. For example, a reference picture generated by NN LF is an example of this.
[0659] Figure 70B is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 70B, a new image (PicAB') generated by inputting two images (PicA and PicB) into the NN may be used as the reference picture. Note that there are not limited to two images input into the NN; more than two images can be input.
[0660] Figure 70C is a conceptual diagram showing another example of generating a generated picture. As illustrated in Figure 70C, if NN processing such as NN LF is not fast enough, a new image (PicA') smaller in size than the input image (PicA) may be used as a reference picture. In other words, the size of the input image and the size of the generated image may be different. Also, information about the size of the generated image may be included in the bitstream.
[0661] While this example shows generation at the picture level, the generation unit is not limited to this. The generated reference picture may be generated in predetermined units such as blocks, slices, or tiles contained within the picture.
[0662] Furthermore, the NN may be implemented using an NPU (Neural Processing Unit), a GPU (Graphics Processing Unit), or a TPU (Tensor Processing Unit).
[0663] Furthermore, as mentioned above, the newly generated image (PicA' or PicAB') may be used as a generated picture for inter prediction. Inter prediction using this generated picture may be called neural network intercoding (NN inter).
[0664] Furthermore, multiple neural networks (NNs) may be combined in the generation of the generated picture. In other words, for example, PicA and PicB may be input to NN1 to generate PicAB', and PicAB' may be input to another NN2 to generate PicAB''.
[0665] In typical image encoding on a CPU, a conventional pipeline is used, as shown in Figure 71A. Specifically, the entire encoding process is divided into several stages (0 to n), and processing is performed in block units within each stage. Furthermore, multiple stages are processed in parallel with each other.
[0666] In contrast, for generated pictures produced using NN processing, as can be seen from the configuration examples in Figures 71B and 71C, the image is first processed by the CPU, and then the data is passed to the NPU / GPU / TPU, etc. Then, NN processing is executed on the NPU / GPU / TPU side using a different pipeline than the CPU. In addition, processing may be executed on the NPU / GPU / TPU in units different from those of the CPU.
[0667] Due to overhead caused by data transfer and other processes, it may take some time for the generated picture produced by the NN processing to become available as a reference image for interpretation.
[0668] Patches may be defined to generate the generated picture by NN processing. A patch refers to a predetermined pixel region containing a predetermined number of pixels, as shown in Figure 72A. The shape of the patch may be a square, a non-square rectangle, or any other shape, as shown in Figure 72A. For example, a patch may have a non-rectangular shape such as a triangle or a rhombus.
[0669] Figure 72B is a conceptual diagram showing a patch containing an extended region. A patch containing an extended region includes an extended region in the vertical and / or horizontal directions relative to the overall size of the patch.
[0670] In the example shown in Figure 72B, the extension area has equal size in the vertical and horizontal directions, resulting in a square patch. The sizes of the vertical and horizontal extension areas are not limited to the example in Figure 72B and may be different. Also, the extension areas do not have to be equal, for example, at the top and bottom, or left and right. Furthermore, information regarding the shape or size of the patch may be added to the header area of the bitstream, or it may be added to the bitstream as metadata (also called meta information).
[0671] Figure 72C is a conceptual diagram showing an image containing patches 1A and 1B that have overlapping regions. Patches 1A and 1B overlap each other across multiple pixels (=set of samples), as shown by the dotted lines in Figure 72C. Information indicating the shape (width and height) or size (number of pixels included in the overlapping region) of this overlapping region may be stored in the bitstream.
[0672] Furthermore, in the example in Figure 72C, each patch contains a 256x256 block. Therefore, each patch can also be described as a patch with an extended area relative to the 256x256 block in Figure 72B. In the example in Figure 72C, the 256x256 block in patch 1A and the 256x256 block in patch 1B are defined so as not to overlap.
[0673] Furthermore, while Figure 72C shows an example of two patches defined within a single picture, in an NN interface, for example, two patches may be defined in each of two pictures. In that case, the positions of the patches defined in the first picture and the second picture (for example, patch 1A in the first picture and patch 2A in the second picture, which is not shown) may be the same or different.
[0674] Figure 72D is a conceptual diagram showing the patches defined in each picture. Specifically, it shows multiple patches defined in an example such as an NN interface, including the first and third patches in the first picture, the second and fourth patches in the second picture, and the fifth and sixth patches in the third picture. The juxtaposed points of the third, fourth, and sixth patches are indicated by an "x" mark.
[0675] Figure 72E is a conceptual diagram showing an example of a method for generating the sixth patch. The sixth patch is obtained by cropping the seventh patch, which is generated by inputting the third and fourth patches into an NN generator. Similarly, the fifth patch may be obtained by cropping an image (the eighth patch, not shown) generated by inputting the first and second patches into an NN generator.
[0676] Furthermore, the first and second pictures may be ordinary reference pictures. In contrast, the third picture, which includes the fifth and sixth patches, may be a generated picture. In one embodiment, the encoding order and display order may be first picture → second picture → third picture. In another embodiment, the display order may be first picture → third picture → second picture.
[0677] Next, the generated picture produced using the transformation process will be described. The encoding device 100 and the decoding device 200 generate a generated picture by performing a transformation process on a single image. The single image input to the transformation process may be a normal reference image (in other words, an image decoded by intra / inter prediction) or an image generated by an NN. Furthermore, the image generated by the transformation process may be used as input to the NN.
[0678] Figure 73A is a conceptual diagram illustrating another example of a generated picture. As illustrated in Figure 73A, a new image (PicA') generated by scaling (enlarging or reducing) an image (PicA) may be used as the reference picture.
[0679] In the case of upscaling (enlarging), a new image PicA' is generated by enlarging the image and then cropping the resulting enlarged image. In the case of downscaling (reducing), a new image PicA' is generated by reducing the image and then padding the resulting reduced image. Note that enlargement can also be expressed as expansion.
[0680] Figure 73B is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73B, a new image (PicA') generated by rotating an image (PicA) may be used as a reference picture. For example, a new image PicA' is generated by rotating the image by an angle θ in a specified direction, and then cropping and padding the rotated image. The rotation direction may always be clockwise or always counterclockwise, and the rotation direction may also be specified as a parameter included in the bitstream along with the rotation angle.
[0681] Figure 73C is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73C, a new image (PicA') generated by mirroring an image (PicA) may be used as the reference picture. In Figure 73C, the left-hand sample of the input image (PicA) is mirrored horizontally.
[0682] Mirroring is not limited to the example above; the sample on the right may be mirrored to the left, or vertical mirroring may be performed (mirroring the upper sample to the lower side, or mirroring the lower sample to the upper side). In this example, when mirroring, half of the width x of the input image (PicA) (or half of the height y in the case of vertical mirroring) is always mirrored, and the width / height of the mirroring is uniformly defined, but the width / height of the mirroring may also be specified by parameters.
[0683] Figure 73D is a conceptual diagram illustrating another example of generating a generated picture. As illustrated in Figure 73D, a new image (PicA') generated by shifting an image (PicA) may be used as the reference picture. In Figure 73D, the input image (PicA) is shifted to the right. In other words, the left side of the input image is padded and the right side is cropped. The shift is not limited to this example; the right side of the sample may be shifted to the left, or a vertical shift (shifted upwards or downwards) may be performed.
[0684] In the above, shifting the image to the right corresponds to shifting the image sample, i.e., the image content, to the right, and to shifting the processing range of the image to the left. Similarly, shifting the image to the left, up, and down corresponds to shifting the image sample, i.e., the image content, to the left, up, and down, respectively, and to shifting the processing range of the image to the right, down, and up, respectively.
[0685] Furthermore, multiple transformations may be combined. For example, an image (PicA) may be scaled to produce an image (Scaled PicA), and then a shifted image (Shifted PicA) may be generated.
[0686] [NN Loop Filter] The NN Loop Filter (NN LF: Neural Network Loop Filter) corresponds to a filtering process performed on an input image (such as a PicA) using a neural network (NN). The image generated by the NN LF may be used as a reference image for interpretation, or it may be used as the display image as is.
[0687] [NN Inter Prediction] NN Inter Prediction (Neural Network Inter Prediction) corresponds to inter prediction performed using a neural network (NN). Specifically, in NN inter prediction, one or more images (PicA and / or PicB, etc.) are used as input to the NN, and an image is generated as a generated image (PicAB', etc.). The generated image is then used as a reference image for inter prediction. Since the image is generated using an NN, the aforementioned NN processing may also be used.
[0688] The generated reference image used in NN interpretation is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).
[0689] Note that while this example shows the use of newly generated images using a neural network as an example of NN interpretation, the NN interpretation process is not limited to this example. For example, images to which NN LF processing has been applied may also be used.
[0690] [Translation Inter Prediction] Translation Inter Prediction corresponds to an inter prediction performed using a transformation process. Specifically, in translation inter prediction, a transformation process is applied to one or more images (PicA and / or PicB, etc.) to generate an image (PicAB', etc.). This generated image is then used as the reference image for inter prediction. The transformation process described above may also be used as the transformation process.
[0691] The generated reference image referenced in the conversion interface prediction is not an encoded or decoded image, and therefore may not contain motion information, reference image index, encoding information, or any combination thereof. Furthermore, since the generated prediction image is generated by predicting the current picture, it may have the same POC (Proof of Concept) as the current picture (CurrentPic).
[0692] [Form for generating another reference image using a reference image (first form)] This section describes an example of performing inter prediction using a reference image (another reference image) that has been additionally generated using the original reference image. In the following, the original reference image will be referred to as the reconstructed reference image, but it may also be referred to as the first reference image or the ungenerated reference image. The reconstructed reference image is an encoded image. In the following, another reference image that has been additionally generated using the reconstructed reference image will be referred to as the generated reference image, but it may also be referred to as the second reference image. Furthermore, the reconstructed reference image is the reference image to be displayed, while the generated reference image is the reference image that is not to be displayed.
[0693] First, a method for generating a generated reference image from a reconstructed reference image will be explained with reference to Figures 74 to 81B. Figure 74 is a flowchart showing an example of an encoding process that includes the generation of a generated reference image by the encoding device 100 according to the embodiment.
[0694] As shown in Figure 74, the encoding device 100 encodes the parameter set and the first image into a bitstream (S1001). The encoding device 100 generates a bitstream containing the encoded parameters and the first image by encoding the parameter set and the first image, for example. The parameter set includes information about the process for generating a generated reference image from a reconstructed reference image (e.g., one or more parameters), and includes information indicating a generation method for generating a generated reference image using the reconstructed first image (e.g., NN, transformation method, etc.). The parameter set may also include, for example, RPL (Reference Picture List) parameters, which will be described later.
[0695] When generating a generated reference image from a reconstructed reference image using a neural network (NN) that transforms and outputs an input image, the parameter set includes information about the NN used to generate the generated reference image from the reconstructed reference image. Details about the NN will be described later. Information about the NN may also be included in the parameters.
[0696] Furthermore, when a generated reference image is generated by performing a predetermined transformation process on the reconstructed reference image, the parameter set includes information about the predetermined transformation process used to generate the generated reference image from the reconstructed reference image. Details regarding the predetermined transformation process will be described later. Information about the predetermined transformation process may be included in the parameters.
[0697] The details of the predetermined conversion process will be described later, but it includes at least one of the following: scaling (scaling process) to enlarge or reduce the reconstructed reference image; shifting (shifting process) to shift the reconstructed reference image in one direction; rotation (rotation process) to rotate the reconstructed reference image; padding (padding process) to replace the pixel values in the reconstructed reference image; and mirroring (mirroring process) to invert the reconstructed reference image with respect to predetermined axis coordinates.
[0698] The first image is the original image received from an external source, but it may also be a predicted image generated by intra-prediction or inter-prediction.
[0699] The parameter set may also include syntax elements to indicate the order of the reconstructed reference image and the generated reference image (for example, the syntax elements shown in Figures 84 to 85B, 92A, 92B, and 100 described later), and syntax elements to indicate the generation method for generating the generated reference image (for example, the syntax elements shown in Figures 76, 78A to 78C described later).
[0700] Next, the encoding device 100 derives a reconstructed reference image by reconstructing the first image (S1002). The encoding device 100 may also derive a reconstructed reference image by applying a loop filter to the reconstructed image obtained by reconstructing the first image. The derived reconstructed reference image is added to the RPL. The reference index assigned to the reconstructed reference image will be explained in the second embodiment. The process in step S1002 is the same as the process in steps Sa_8 to Sa_11 shown in Figure 4B.
[0701] Next, the encoding device 100 generates a generated reference image from the reconstructed reference image using a parameter set (S1003). The process of generating a generated reference image from the reconstructed reference image is performed on a picture-by-picture basis, but it may also be performed on a block-by-block basis or the like.
[0702] The method for generating the generated reference image will be described later using Figures 76 to 81B. The generated reference image is added to the RPL. The reference index assigned to the reconstructed reference image and the generated reference image will be explained in the second form.
[0703] Next, the encoding device 100 encodes the second image into a bitstream using the generated reference image (S1004). In step S1004, for example, an interpretation process is performed in which the generated reference image is used as a reference to predict the second image. The original image, the second image, is encoded using the predicted image predicted using the generated reference image.
[0704] In step S1004, the second image may be encoded into a bitstream using a reference image other than the generated reference image.
[0705] Figure 75 is a flowchart showing an example of a decoding process including the generation of a reference image by the decoding device 200 according to the embodiment.
[0706] As shown in Figure 75, the decoding device 200 receives the bitstream transmitted from the encoding device 100, decodes the parameter set from the bitstream (S1101), and derives a reconstructed reference image by decoding the first image from the bitstream (S1102).
[0707] Next, the decoding device 200 generates a generated reference image from the reconstructed reference image using a parameter set (S1103). The decoding device 200 generates a generated reference image from the reconstructed reference image based on information regarding the process for generating a generated reference image from the reconstructed reference image. The generated reference image generated in step S1103 and the generated reference image generated in step S1003 shown in Figure 74 are the same image.
[0708] Next, the decoding device 200 decodes the second image from the bitstream using the generated reference image (S1104). In step S1104, for example, an interpretation process is performed in which the generated reference image is used as a reference to predict the second image. The original image, the second image, is decoded using the predicted image predicted using the generated reference image.
[0709] In step S1104, the second image may be decoded from the bitstream using a reference image other than the generated reference image.
[0710] Next, the generation method of the generated reference image performed by the encoding device 100 and the decoding device 200 will be explained with reference to Figures 77A, 77B, and 79 to 81B, and each syntax structure will be explained with reference to Figures 76 and 78A to 78C. The generation method of the generated reference image can be broadly classified into (generation method a) a generation method that generates a generated reference image by NN processing (NN filtering) using one input image, (generation method b) a generation method that generates a generated reference image by NN processing (NN filtering) using multiple input images, (generation method c) a generation method that generates a generated reference image by image transformation processing, and (generation method d) a generation method that generates a generated reference image using bidirectional prediction. In generation method a, the NN processes one input image to generate one output image as the generated reference image, and in generation method b, the NN processes two or more input images to generate one output image as the generated reference image.
[0711] All generated reference images may be generated using only one of the four generation methods, or each reconstructed reference image may be generated using a different generation method (i.e., a generation method determined individually for each reconstructed reference image). Information indicating which generation method is used to generate the generated reference images may be encoded as a parameter in the bitstream.
[0712] Figure 76 shows an example of a syntax structure used in the method for generating a generated reference image. Information about each syntax element included in the syntax structure is added to the bitstream and transmitted to the decoding device 200. The same applies to subsequent figures showing syntax structures.
[0713] In Figure 76, the syntax structure includes generated_flag and generated_method as syntax elements.
[0714] `generated_flag` is a flag that indicates whether or not a generated reference image exists that was generated from a reconstructed reference image of a given POC. `generated_flag=1` indicates that a generated reference image exists, and `generated_flag=0` indicates that a generated reference image does not exist, but this is not limited to these two cases. Note that `generated_flag` may be assigned to each reference image in the RPL. For example, `generated_flag=0` may be assigned to a reference image that is not a generated reference image (e.g., a reconstructed reference image), and `generated_flag=1` may be assigned to a generated reference image.
[0715] This makes it easy to determine whether a reference image in the RPL is a reconstructed reference image or a generated reference image by checking whether the value of generated_flag is 0 or not.
[0716] `generated_method` indicates a generation method for generating a generated reference image from a reconstructed reference image, when such a generated reference image exists. If `generated_method = 0`, it indicates that the above-described (generation method a) is used to generate the generated reference image. If `generated_method = 1`, it may indicate that the above-described (generation method b) is used to generate the generated reference image. Furthermore, if `generated_method = 2`, it indicates that the above-described (generation method c) is used to generate the generated reference image. If `generated_method = 3`, it indicates that the above-described (generation method d) is used to generate the generated reference image, but it is not limited to these. Note that if `generated_method = 1`, it may be possible to generate a higher quality generated reference image than if `generated_method = 0`.
[0717] Furthermore, the encoding device 100 may add generated_method to the bitstream (for example, the header area of the bitstream) when generated_flag = 1. This reduces the number of bits because generated_method is not added to the bitstream when generated_flag = 0.
[0718] Such flags and information indicating the method for generating the generated reference image may be added to the bitstream as a parameter set. Furthermore, this information is added to the bitstream according to the RPL syntax, but may also be added according to the VPS, SPS, PPS, SEI, or SH syntax. The information added to the bitstream is transmitted from the encoding device 100 to the decoding device 200, where it is decoded.
[0719] Furthermore, generated_method may be assigned to each generated reference image in the RPL. For example, if the reference image with refIdx=0 is a generated reference image, a generated_method for refIdx=0 may be added to the bitstream, and if the reference image with refIdx=2 is a generated reference image, a generated_method for refIdx=2 may be added to the bitstream. In this case, different values may be assigned as generated_method for refIdx=0 and refIdx=2. This allows generated reference images generated by different methods, such as generated_method=0 to 4, to be included in the reference images, thereby improving encoding efficiency.
[0720] Figure 77A is a conceptual diagram showing an example of a method for generating a generated reference image when there is one input image. Figure 77B is a conceptual diagram showing an example of a method for generating a generated reference image when there are multiple input images. Figures 77A and 77B show examples of NN interpretation. Although Figure 77B shows an example with two input images, the number of input images is not particularly limited as long as it is two or more. Figure 77A illustrates (generation method a), and Figure 77B illustrates (generation method b). Images with dot hunting are generated reference images.
[0721] As shown in Figure 77A, a generated reference image (PicA, which is POC_A) is generated using a single reference image as input to the NN. The generated reference image is then used as the reference image for interpretation. The single input reference image is a reconstructed reference image. Note that the image input to the NN may or may not be assigned a reference index such as a POC number. In other words, the image input to the NN may or may not be a reference image.
[0722] As shown in Figure 77B, two reference images (PicA, which is POC_A, and PicB, which is POC_B) are used as input to the NN to generate one generated reference image (PicAB', which is POC_C). Each of the two input reference images is a reconstructed reference image. PicB is a different image from PicA, and for example, PicA and PicB constitute a moving image. PicA is an example of a first reference image, and PicB is an example of a third reference image. PicB is a reconstructed reference image derived by reconstructing an image (encoded image) different from PicA.
[0723] Furthermore, if generate_method = 0 or 1 (i.e., if the generated reference image is generated using an NN), the encoding device 100 may add information for generating the generated reference image using the NN to the bitstream as information about the NN. The information about the NN includes the type of filter, the filtering method, the number of layers, the model size, the dataset used, URI information for downloading the NN model, the POC numbers of one or more input images used as input to the NN, the POC numbers of the images output from the NN, an SEI (neural network post-filter activation SEI) indicating the activation of the post-processing filter by the NN, an SEI (neural network post-filter characteristic SEI) indicating the characteristics of the post-processing filter by the NN, and at least one of other SEIs related to the NN. Information about the NN including at least one of these is added to the bitstream.
[0724] Since the NN information includes URI information, the decoding device 200 can download the NN model using the URI information attached to the bitstream and generate a reference image using the same NN model as the encoding device 100, thereby enabling proper decoding of the bitstream.
[0725] Furthermore, the information regarding the NN may include different information for each generated reference image. For example, the encoding device 100 may include information indicating the use of a different NN model for each generated reference image, such as URL information for downloading NN model A for generated reference image A, and URL information for downloading NN model B for generated reference image B. This allows for the generation of one or more different generated reference images using one or more different NN models and their addition to the RPL, thereby improving prediction efficiency and, consequently, coding efficiency.
[0726] Next, "(Generation Method c) Generation Method Using Image Transformation Processing" will be described. The encoding device 100 may extract camera parameters and convert them into image transformation parameters (parameters used when performing scaling, padding, shifting, etc.), and add the transformation parameters to the bitstream. In this case, the encoding device 100 and the decoding device 200 generate a generated reference image by performing scaling, padding, shifting, etc., on the reconstructed image using the camera parameters. A parameter set including transformation parameters may be used when generating a generated reference image from a reconstructed reference image.
[0727] Here, examples of transformation methods include scaling, rotation, mirroring, shifting, and padding. Furthermore, a generated reference image may be produced by combining at least two of these transformation methods. The parameters (an example of a parameter set) for each transformation method are described below. Additionally, an example where the image to be transformed is a reconstructed reference image is described below. Note that the image to be transformed may or may not have a reference index assigned to it, such as a POC number. In other words, the image to be transformed may or may not be a reference image.
[0728] Figure 78A shows an example of the syntax structure used when generating a generated reference image using the conversion method. Figure 78B shows another example of the syntax structure used when generating a generated reference image using the conversion method. Figure 78C is a table showing the relationship between the conversion method and the value.
[0729] The generated_flag shown in Figures 78A and 78B is a flag, similar to Figure 76, that indicates whether or not a generated reference image exists that was generated from a reconstructed reference image of a given POC. If generated_flag = 1, a generated reference image exists, and if generated_flag = 0, a generated reference image does not exist, but this is not limited to these two cases.
[0730] The `transform_method` parameter shown in Figure 78A indicates the type of transformation method used when generating a generated reference image from a reconstructed reference image. As shown in Figure 78C, different values are pre-set for each type of transformation method, and the value corresponding to the type of transformation method used is set in `transform_method`.
[0731] Furthermore, as shown in Figure 78B, when multiple transformation methods are used, num_transform and transform_method[i] are set. num_transform indicates the total number of transformation processes required to generate a generated reference image from a reconstructed reference image. Index i represents the order of each transformation process performed on the reconstructed image to generate the generated reference image.
[0732] Furthermore, if generated_method=2 or 3, a set of conversion parameters indicating the generation method for generating the generated reference image may be signaled to the decoding device 200. Also, generated_flag may be replaced with generated_method=2 or 3. In other words, the syntax element does not need to include generated_flag.
[0733] Figure 79 is a conceptual diagram showing another example of generating a generated picture. The hatched areas (margins) in Figure 79 indicate areas to be padded.
[0734] As shown in Figure 79, a new image (PicA') generated by scaling (reducing) the image (PicA) may be used as the reference image. In the case of downscaling (reducing), a new image PicA' is generated by padding the reduced image (Scaled PicA) generated by reducing the original image. Image (PicA) is an example of a reconstructed reference image, and image PicA' is an example of a generated reference image.
[0735] Downscaling can be performed in the parallel direction, the vertical direction, or both.
[0736] A schematic diagram of the case where the conversion method is upscaling (enlargement) is shown in Figure 73A above. Upscaling can be horizontal upscaling, vertical upscaling, or both, and the hatched areas (margins) are cropped.
[0737] Furthermore, if transform_method = 0 (i.e., the generated reference image is generated using scaling), the encoding device 100 may add information for generating the generated reference image using scaling to the bitstream as information related to the predetermined transformation process described above.
[0738] Information regarding a predetermined conversion process includes each scale factor. Each scale factor includes, for example, horizontal_scaling_num, which represents the horizontal scale factor (numerator); horizontal_scaling_dem, which represents the horizontal scale factor (denominator) (minimum value is 1); vertical_scaling_num, which represents the vertical scale factor (numerator); and vertical_scaling_dem, which represents the vertical scale factor (denominator) (minimum value is 1). The horizontal scale is calculated as horizontal_scaling_num / horizontal_scaling_dem, and the vertical scale is calculated as vertical_scaling_num / vertical_scaling_dem, but is not limited to these.
[0739] Upscaling is performed when the scale > 1, and downscaling is performed when the scale < 1. The values of scaling_num and scaling_dem may be converted to values equivalent to powers of 2 so that division is replaced by a right shift operation.
[0740] Next, a schematic diagram of the case where the transformation method is rotation is shown in Figure 73B above.
[0741] Furthermore, if transform_method = 3 (i.e., the generated reference image is generated using rotation), the encoding device 100 may add information for generating the generated reference image using rotation to the bitstream as information related to the predetermined transformation process described above.
[0742] Information regarding the predetermined conversion process includes conversion parameters. These conversion parameters include rotation_angle and rotation_clockwise.
[0743] `rotation_angle` indicates the rotation angle when rotating the reconstructed reference image, and corresponds to the "angle θ" shown in Figure 73B. `rotation_angle` is, for example, an integer value.
[0744] `rotation_clockwise` indicates the direction of rotation when rotating the reconstructed reference image, and can be either clockwise or counterclockwise. `rotation_clockwise = 0` indicates counterclockwise rotation, and `rotation_clockwise = 1` indicates clockwise rotation, but is not limited to these two.
[0745] Furthermore, the encoding device 100 may include the following conversion parameters in the information regarding the predetermined conversion process, as another example of conversion parameters. The conversion parameters include at least rotation_angle. If rotation_angle < 0, it indicates counterclockwise rotation and its rotation angle, and if rotation_angle > 0, it indicates clockwise rotation and its rotation angle.
[0746] Next, a schematic diagram of the case where the conversion method is mirroring is shown in Figure 73C above.
[0747] Furthermore, if transform_method = 4 (i.e., the generated reference image is generated using mirroring), the encoding device 100 may add information for generating the generated reference image using mirroring to the bitstream as information related to the predetermined transformation process described above.
[0748] Information regarding a predetermined conversion process includes conversion parameters. These conversion parameters include mirror_position, mirror_direction, and mirror_forward.
[0749] `mirror_position` indicates the pixel position of the line used as the axis for mirroring (inverting) the generated reference image.
[0750] `mirror_direction` indicates whether the mirroring is vertical or horizontal. `mirror_direction = 0` indicates horizontal mirroring along the y-axis, while `mirror_direction = 1` may indicate vertical mirroring along the x-axis.
[0751] `mirror_forward` indicates which sample (which region in the reconstructed reference image) to mirror in which direction. When `mirror_forward` = 0, it means that in horizontal mirroring, the sample on the right is mirrored to the left, and in vertical mirroring, the sample on the lower side is mirrored to the upper side. When `mirror_forward` = 1, it may indicate that in horizontal mirroring, the sample on the left is mirrored to the right, and in vertical mirroring, the sample on the upper side is mirrored to the lower side.
[0752] The mirroring method may be one of the above pre-configured, or it may be set for each reconstructed reference image.
[0753] Furthermore, the encoding device 100 may include the following conversion parameters in the information regarding the predetermined conversion process, as another example of conversion parameters. The conversion parameters include mirror_position and mirror_direction. mirror_position indicates the pixel position of the line to be used as the axis for mirroring, as described above.
[0754] mirror_direction=0 indicates horizontal mirroring along the y-axis (mirroring the right sample to the left), mirror_direction=1 indicates horizontal mirroring along the y-axis (mirroring the left sample to the right), mirror_direction=2 indicates vertical mirroring along the y-axis (mirroring the lower sample to the upper), and mirror_direction=3 may indicate vertical mirroring along the y-axis (mirroring the upper sample to the lower).
[0755] Next, schematic diagrams of the case where the conversion method is shift are shown in Figure 73D above and Figure 80 below. Shifting is performed, for example, horizontally (right or left), vertically (up or down), or a combination of both.
[0756] Figure 80 is a conceptual diagram showing another example of generating a generated picture. The hatched areas in Figure 80 indicate areas to be padded.
[0757] Figure 80 shows a conceptual diagram of shifting an image to the left by a specified number of pixels. Specifically, Figure 80 shows a conceptual diagram when the shift direction is horizontal, the shift movement is 0 (moving to the left), and the number of pixels to be shifted is 10 pixels.
[0758] Furthermore, if transform_method = 1 (i.e., the generated reference image is generated using shift), the encoding device 100 may add information for generating the generated reference image using shift as information related to the predetermined transformation process described above to the bitstream.
[0759] Information regarding a predetermined conversion process includes conversion parameters. These conversion parameters include shift_pixels, shift_direction, and shift_movement.
[0760] `shift_pixels` indicates the amount of shift (number of pixels) to shift the image by the specified pixel value, and is expressed as an integer.
[0761] `shift_direction` indicates the shift direction; `shift_direction = 0` may mean the image is shifted horizontally, and `shift_direction = 1` may mean the image is shifted vertically.
[0762] If shift_direction=0 and shift_movement=0, it indicates that the image will be shifted to the left by shift_pixels. Also, if shift_direction=0 and shift_movement=1, it indicates that the image will be shifted to the right by shift_pixels. Furthermore, if shift_direction=1 and shift_movement=0, it indicates that the image will be shifted upward by shift_pixels. Also, if shift_direction=1 and shift_movement=1, it indicates that the image will be shifted downward by shift_pixels.
[0763] Furthermore, the example in Figure 80 shows an image generated by shifting the image to the left by 10 pixels when shift_direction=0, shift_movement=0, and shift_pixels=10.
[0764] Furthermore, the encoding device 100 may include the following conversion parameters in the information regarding a predetermined conversion process, as another example of conversion parameters. The conversion parameters include Horizontal_shift_pixel and Vertical_shift_pixel.
[0765] Horizontal_shift_pixel indicates the direction of the shift, for example, a positive value indicates a shift to the right, and a negative value indicates a shift to the left.
[0766] Vertical_shift_pixel indicates the direction of the shift, for example, a positive value indicates an upward shift, and a negative value indicates a downward shift.
[0767] Furthermore, shift_pixels may be further subdivided by pixel precision (size_pixel_accuracy). Examples of pixel units include, but are not limited to, 4 pixels, 8 pixels, and 128 pixels. When size_pixel_accuracy = 128, the actual pixel shift amount is calculated as Horizontal_shift_pixel × size_pixel_accuracy.
[0768] Next, schematic diagrams illustrating the case where the conversion method is padding are shown in Figures 73A and 73D above. Padding is used in conjunction with scaling, shifting, etc., but it may also be used alone.
[0769] Furthermore, if transform_method = 2 (i.e., the generated reference image is generated using padding), the encoding device 100 may add information for generating the generated reference image using padding to the bitstream as information related to the predetermined transformation process described above.
[0770] Information regarding a predetermined conversion process includes conversion parameters. These conversion parameters include the pad_method.
[0771] pad_method indicates the padding method. If pad_method=0, it indicates duplicating the pixels of the last row / column; if pad_method=1, it indicates padding with 0; if pad_method=2, it indicates padding with the maximum value (e.g., 1023 for 10 bits); and if pad_method=3, it may indicate padding with a predefined value. The padding method may be fixed in advance to one of the above, or it may be determined for each conversion process.
[0772] Next, an example of generating a generated reference image from images generated using different conversion methods will be explained with reference to Figures 81A and 81B. Figures 81A and 81B are conceptual diagrams showing another example of generating a generated picture. Figures 81A and 81B illustrate an example of generating a generated reference image using two different conversion methods, but the generated reference image may be generated using three or more different conversion methods.
[0773] Figure 81A shows a conceptual diagram of a case where a new reference image (generated reference image) is generated by applying bidirectional prediction to two reference images (reconstructed reference images) that have undergone different scaling processes. Specifically, it shows a conceptual diagram of a case where a new reference image (POC1') is generated based on an image (POC0') generated by upscaling a reference image (POC0) and an image (POC2') generated by downscaling a reference image (POC2). Reference images (POC0) and (POC2) are examples of reconstructed reference images, and reference image (POC1') is an example of a generated reference image. Furthermore, reference image (POC0) is an example of a first reference image, and reference image (POC2) is an example of a third reference image.
[0774] Figure 81B shows a conceptual diagram of generating a new reference image (generated reference image) by applying bidirectional prediction to two reference images (reconstructed reference images) that have undergone different shift processing. Specifically, Figure 81B shows a conceptual diagram of generating a new reference image (POC1') based on an image (POC0') generated by shifting a reference image (POC0) to the right and an image (POC2') generated by shifting a reference image (POC2) to the left. Reference images (POC0) and (POC2) are examples of reconstructed reference images, and reference image (POC1') is an example of a generated reference image.
[0775] By using bidirectional prediction to generate the generated reference image in this way, prediction accuracy can be improved in areas where prediction is difficult and prediction errors are likely to occur with unidirectional prediction.
[0776] Furthermore, if generated_method=3, it may indicate that bidirectional prediction will be performed. Each reference image may be appended with information indicating the transformation method, such as transform_method[x], where x may indicate the number of transformation operations for each reference image.
[0777] Furthermore, there are no particular limitations on the combination of transformation methods used when performing bidirectional prediction; at least two transformation methods must be selected from scaling (upscaling / downscaling), rotation, mirroring, shifting, and padding. For example, scaling and mirroring may be selected, or rotation and shifting may be selected.
[0778] Thus, by using a generated reference image created from a reconstructed reference image, it is possible to generate a predicted image that is closer to the original image, which can improve coding efficiency.
[0779] [Method for managing other reference images generated using a reference image (second form)] Next, a method for managing other reference images (generated reference images) generated using a reference image will be explained with reference to Figures 82 to 100. The reconstructed reference image and the generated reference image are added to the RPL and stored and managed in memory. The reconstructed reference image and the generated reference image are stored and managed, for example, in the frame memory 122 of the encoding device 100, the frame memory 214 of the decoding device 200, etc. The frame memory 214 of the decoding device 200 is also referred to as DPB (Decoded Picture Buffer).
[0780] Reconstructed reference images have MV (Motion Value), but generated reference images may not. Therefore, the following section also describes how to manage reconstructed and generated reference images separately.
[0781] Figure 82 is a flowchart showing an example of an encoding process including a method for managing generated reference images by an encoding device according to an embodiment. The generated reference image may also be referred to as another reference image.
[0782] As shown in Figure 82, the encoding device 100 encodes the first image into a bitstream (S2001). The encoding device 100 may also encode the first image and the corresponding RPL parameter set into a bitstream. The RPL parameter set includes information on whether the reference image used when encoding the first image is another reference image generated using the reference image (i.e., a generated reference image), or information that distinguishes between the reference image and another reference image (i.e., a generated reference image).
[0783] Furthermore, if the first image is encoded using a generated reference image, the encoding device 100 may encode the generated reference image and add it to the bitstream. This allows the decoding device 200 to decode the generated reference image from the bitstream, thus eliminating the need for the decoding device 200 to generate the generated reference image. Therefore, since the decoding device 200 does not need to perform processing to generate the generated reference image, such as NN filtering or conversion processing, the processing load or circuit size of the decoding device 200 can be reduced.
[0784] Next, the encoding device 100 derives a reconstructed reference image by reconstructing the first image (S2002). The process in step S2002 is the same as the process in steps Sa_8 to Sa_11 shown in Figure 4B.
[0785] Next, the encoding device 100 generates a generated reference image from the reconstructed reference image (S2003). The encoding device 100 directly generates the generated reference image from the reconstructed reference image using, for example, the same process as in step S1003 in Figure 74. Direct generation means generating the generated reference image by performing filtering, transformation, etc., on the reconstructed reference image.
[0786] Next, the encoding device 100 assigns a plurality of different reference indices to a plurality of reference images, including the reconstructed reference image and the generated reference image, based on the RPL parameter set (S2004). The reference index is identification information for identifying the reference image used for image prediction. The encoding device 100 assigns, for example, reference indices that can distinguish between the reconstructed reference image and the generated reference image. The assigned reference indices will be described later, but examples include POC numbers, index numbers, or combinations thereof.
[0787] Next, the encoding device 100 encodes the second image into a bitstream using the reference index (S2005). The encoding device 100 performs interpretation using an image selected from among multiple reference images, including the reconstructed reference image and the generated reference image, using the reference index, as the reference image for predicting the second image. The RPL includes the reconstructed reference image and the generated reference image.
[0788] Figure 83 is a flowchart showing an example of a decoding process including a method for managing generated reference images by the decoding device 200 according to the embodiment.
[0789] As shown in Figure 83, the decoding device 200 decodes the first image from the bitstream (S2101). The decoding device 200 may also decode the first image and the corresponding RPL parameter set from the bitstream.
[0790] Next, the decoding device 200 derives a reconstructed reference image by reconstructing the first image (S2102). The process in step S2102 is the same as the process in steps Sp_1 to Sp_5 shown in Figure 7B.
[0791] Next, the decoding device 200 generates a generated reference image from the reconstructed reference image (S2103). The decoding device 200 directly generates a generated reference image from the reconstructed reference image using, for example, the same process as in step S1103 in Figure 75.
[0792] Furthermore, if the decoding device 200 has received a bitstream in which the generated reference image is encoded, it may decode the generated reference image from the bitstream. This eliminates the need for the decoding device 200 to decode the generated reference image from the bitstream, thereby reducing the decoding processing load on the decoding device 200.
[0793] Next, the decoding device 200 assigns a plurality of different reference indices to a plurality of reference images, including the reconstructed reference image and the generated reference image, based on the RPL parameter set (S2104). The decoding device 200 assigns, for example, reference indices that can distinguish between the reconstructed reference image and the generated reference image. The assigned reference indices will be described later, but examples include POC numbers, index numbers, or combinations thereof.
[0794] Furthermore, if the bitstream contains information indicating whether a reference image included in multiple reference images is a reference image generated from other reference images, the decoding device 200 may decode such information from the bitstream.
[0795] Next, the decoding device 200 decodes the second image from the bitstream using the reference index (S2105). The decoding device 200 performs interpretation using an image selected from among a plurality of reference images, including the reconstructed reference image and the generated reference image, using the reference index, as the reference image for predicting the second image.
[0796] Here, we will explain the reallocation of reference indices when a generated reference image is generated using a reconstructed reference image, as performed by the encoding device 100 and the decoding device 200, respectively. When a reconstructed reference image is generated, a reference index is assigned to it. However, if a generated reference image is generated afterward, a reallocation of reference indices is performed for each reference image in the RPL. We will now explain the case where the reference index is a POC number.
[0797] Figures 84 to 85B show an example of a syntax structure used when reallocating a reference index. The bold text in Figures 84 to 85B indicates a syntax element in the syntax structure. Here, the syntax structure includes at least one of the following syntax elements: generated_pic_flag, generated_poc_scale_minus1, and generated_pic_idx. The RPL parameter set may include a value corresponding to at least one of each syntax element.
[0798] The generated_pic_flag shown in Figure 84 is a flag indicating whether a reference image is a generated reference image or not. For example, if generated_pic_flag = 0, the generated reference screen is not an image generated using a reconstructed reference image, so the POC number is not reassigned. On the other hand, if generated_pic_flag = 1, the generated reference screen is generated using a reconstructed reference image, so the POC is reassigned to all reference images in the RPL and DPB. By assigning a new POC (POC') to all reference images, including the generated reference image, the reference images in the RPL and DPB can be arranged in any order, increasing the system's flexibility and the potential for improved coding efficiency.
[0799] `generated_poc_scale_minus1` represents the POC difference between two adjacent images in display order and is a value greater than or equal to 0. The value of `generated_poc_scale_minus1` is appended to the bitstream. In the example in Figure 84, `generated_poc_scale_minus1` is set when the reference image is a generated reference image. In the example in Figure 85A, `generated_poc_scale_minus1` is set regardless of whether the reference image is a generated reference image or not.
[0800] `generated_poc_scale_minus1` is expressed as a power of n (n >= 1), and integer values can be replaced by shift operations. For example, if n = 2, the number of generated images will be a power of 2 (2, 4, 8, ...), but it is not limited to this.
[0801] For example, generated_poc_scale_minus1 may represent the value of the exponent of a power of 2. For example, generated_poc_scale_minus1=0 may represent 2 to the power of 0, generated_poc_scale_minus1=1 may represent 2 to the power of 1, and generated_poc_scale_minus1=N may represent 2 to the power of N. This allows, for example, multiplication operations using generated_poc_scale_minus1 to be replaced with left bit shift operations, and division operations to be replaced with right bit shift operations, thereby reducing the amount of processing required.
[0802] If there is one generated reference image, the encoding device 100 and the decoding device 200 convert the POC to POC' using the following formula 1, and reassign the POC number to the reconstructed reference image using formula 2. The encoding device 100 and the decoding device 200 may each store formulas 1 and 2 in advance. In other words, the encoding device 100 and the decoding device 200 may each store the same formula in advance.
[0803] POC'=[(generated_poc_scale_minus1+1)×POC+(generated_pic_idx_minus1+1)] ... (Formula 1)
[0804] POC' = [(generated_poc_scale_minus1+1)×POC] ... (Formula 2)
[0805] The encoding device 100 and the decoding device 200 may either add a PictureOutputFlag to the bitstream as information indicating whether or not to display the decoded picture, or, for the generated picture (i.e., the generated reference image), they may add a value to the bitstream indicating that it should not be displayed, such as the PictureOutputFlag value. This allows the decoding device 200 to determine whether or not to display the generated reference image. The information indicating whether or not to display is an example of a reference picture.
[0806] Furthermore, if the value of the PictureOutputFlag attached to the generated reference image is 1, which indicates that it should be displayed, the decoding device 200 may determine that this constitutes a compliance violation of the standard. For example, the standard may specify a bitstream conformance requirement that the PictureOutputFlag of the generated reference image must be 0. This prevents the PictureOutputFlag of the generated reference image from being set to 1.
[0807] Furthermore, even if the PictureOutputFlag value of the generated reference image is 1, the decoding device 200 may treat it as 0 and not display the generated reference image. This prevents the generated reference image from being displayed. Also, if the encoding device 100 does not add the PictureOutputFlag value of the generated reference image to the bitstream, the decoding device 200 may estimate that value to be 0 and proceed with processing. This reduces the amount of bits used.
[0808] Furthermore, if multiple generated reference images are generated from a single reconstructed reference image, generated_poc_scale_minus1 and generated_pic_idx_minus1 may be set for each generated generated reference image. Also, the POC number may be reassigned using generated_pic_idx as shown in Figure 85B. For the reconstructed reference image and the generated reference image, the decoding device 200 may convert the POC to POC' using the following formula. The same applies to the encoding device 100.
[0809] POC'=[(generated_poc_scale_minus1+1)×POC+generated_pic_idx] (Formula 3)
[0810] Note that generated_pic_idx may be assigned to each reference picture in the RPL. For example, generated_pic_idx = 0 may be assigned to the reconstructed reference image, and generated_pic_idx >= 1 may be assigned to the generated reference image. This allows you to determine whether a reference image in the RPL is a reconstructed reference image or a generated reference image by checking whether the value of generated_pic_idx is 0 or not. Furthermore, by assigning the value of generated_pic_idx as described above, the value of POC' can be calculated using a common formula (formula 3) regardless of whether it is a reconstructed reference image or a generated reference image.
[0811] Furthermore, generated_pic_flag may be assigned to each reference image within the RPL. For example, generated_pic_flag=0 may be assigned to the reconstructed reference image and generated_pic_flag=1 to the generated reference image. This allows determining whether a reference image within the RPL is a reconstructed reference image or a generated reference image by checking whether the value of generated_pic_flag is 0 or not.
[0812] Furthermore, the encoding device 100 may add generated_pic_idx to the bitstream when generated_pic_flag = 1. This ensures that generated_pic_idx is not added to the bitstream when generated_pic_flag = 0, thereby reducing the number of bits.
[0813] Furthermore, if generated_pic_flag = 0 and generated_pic_idx is not attached to the bitstream, the decoding device 200 may estimate the value of generated_pic_idx to be 0. This allows the decoding device 200 to assign the value 0 to generated_pic_idx of the reconstructed reference image.
[0814] Furthermore, if multiple generated reference images are generated from a single reconstructed reference image, generated_poc_scale_minus1 and generated_pic_idx are set for each generated reference image.
[0815] The following explains the specific reassignment of reference indices. generated_pic_flag=0 indicates a reconstructed reference image, and generated_pic_flag=1 in...
Claims
1. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit uses the memory to encode a parameter set into a bitstream, encode a first image into the bitstream, derive a first reference image by reconstructing the first image, and generate a second reference image from the first reference image using the parameter set.
2. The encoding device according to claim 1, wherein the second reference image is generated by inputting the first reference image into a neural network that transforms and outputs an input image, and the parameter set includes information indicating the neural network.
3. The encoding device according to claim 1, wherein the second reference image is generated by performing a predetermined transformation process on the first reference image, and the parameter set includes information indicating the predetermined transformation process.
4. The encoding apparatus according to claim 3, wherein the predetermined conversion process includes at least one of the following: a scaling process for enlarging or reducing the first reference image; a shifting process for shifting the first reference image in one direction; a rotation process for rotating the first reference image; a padding process for replacing the pixel values in the first reference image; and a mirroring process for inverting the first reference image with respect to predetermined axis coordinates.
5. The encoding device according to any one of claims 1 to 4, wherein the process of generating the second reference image from the first reference image is performed on a picture-by-picture basis.
6. An encoding device according to any one of claims 1 to 4, which generates the second reference image based on the first reference image and a third reference image derived by reconstructing a second image different from the first image.
7. The encoding device according to any one of claims 1 to 4, wherein the parameter set includes information indicating a generation method for generating the second reference image from the first reference image.
8. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit uses the memory to decode a parameter set from a bitstream, derives a first reference image by decoding a first image from the bitstream, and generates a second reference image from the first reference image using the parameter set.
9. The decoding device according to claim 8, wherein the second reference image is generated by inputting the first reference image into a neural network that transforms and outputs an input image, and the parameter set includes information indicating the neural network.
10. The decoding device according to claim 8, wherein the second reference image is generated by performing a predetermined transformation process on the first reference image, and the parameter set includes information indicating the predetermined transformation process.
11. The decoding device according to claim 10, wherein the predetermined conversion process includes at least one of the following: a scaling process for enlarging or reducing the first reference image; a shifting process for shifting the first reference image in one direction; a rotation process for rotating the first reference image; a padding process for replacing the pixel values in the first reference image; and a mirroring process for inverting the first reference image with respect to predetermined axis coordinates.
12. A decoding device according to any one of claims 8 to 11, wherein the process of generating the second reference image from the first reference image is performed on a picture-by-picture basis.
13. A decoding device according to any one of claims 8 to 11, which generates a second reference image based on the first reference image and a third reference image derived by reconstructing a second image different from the first image.
14. The decoding device according to any one of claims 8 to 11, wherein the parameter set includes information indicating a generation method for generating the second reference image from the first reference image.
15. An encoding method comprising encoding a parameter set into a bitstream, encoding a first image into the bitstream, deriving a first reference image by reconstructing the first image, and generating a second reference image from the first reference image using the parameter set.
16. A decoding method comprising decoding a parameter set from a bitstream, deriving a first reference image by decoding a first image from the bitstream, and generating a second reference image from the first reference image using the parameter set.