Image decoding device, image encoding device, image decoding method, and image encoding method
The image decoding and encoding apparatus addresses the ambiguity in video processing by using a neural network filter unit to explicitly define the input and output formats of the neural network filter, ensuring efficient and clear processing across various color formats and subsampling schemes.
Patent Information
- Application Number
- JP2024161428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-01-14
AI Technical Summary
Existing video encoding and decoding technologies, such as those described in Non-Patent Documents 1 and 2, lack explicit definitions for the input and output of post-filter processing, particularly in terms of color space and subsampling, leading to ambiguity and inefficiency in processing and output generation.
An image decoding and encoding apparatus that includes a header decoding unit to derive filter strength and bit depth for filter processing, and a neural network filter unit that specifies the format of an input tensor of a neural network model, allowing for explicit derivation of input image information, filter processing, and output tensor generation.
This solution enables clear and efficient processing by explicitly defining the input and output formats of the neural network filter, facilitating the determination of processing abilities and supporting various color formats and subsampling schemes.
Smart Images

Figure 0007697122000001 
Figure 0007697122000002 
Figure 0007697122000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to an image decoding apparatus, an image encoding apparatus, an image decoding method, and an image encoding method.
Background Art
[0002] In order to efficiently transmit or record a moving image, a moving image encoding apparatus that generates encoded data by encoding the moving image, and a moving image decoding apparatus that generates a decoded image by decoding the encoded data are used.
[0003] Specific moving image encoding methods include, for example, the H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding) methods.
[0004] In such a moving image encoding method, an image (picture) constituting the moving image is managed by a hierarchical structure composed of a slice obtained by dividing the image, a coding tree unit (CTU) obtained by dividing the slice, a coding unit (sometimes called a Coding Unit: CU) obtained by dividing the coding tree unit, and a transform unit (TU) obtained by dividing the coding unit, and is encoded / decoded for each CU. and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the prediction image from the input image (original image) is encoded. Examples of the method for generating the prediction image include inter-picture prediction (inter prediction) and intra-picture prediction (intra prediction).
[0005] Also, in such a moving image encoding method, usually, a prediction image is generated based on a local decoded image obtained by encoding / decoding the input image, and
[0006] In addition, Non-Patent Document 1 can be cited as a technology for recent video encoding and decoding.
[0007] H.274 stipulates an additional extension information SEI for transmitting information such as the nature of an image, display method, and timing simultaneously with encoded data.
[0008] In Non-Patent Document 1, Non-Patent Document 2, and Non-Patent Document 3, an SEI for transmitting the topology and parameters of a neural network filter used as a post-filter is disclosed in a method of explicitly stipulating it and a method of indirectly stipulating it as reference information.
Prior Art Documents
Non-Patent Documents
[0009]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0010] However, both Non-Patent Document 1 and Non-Patent Document 2 have the problem that the input and output of the post-filter processing are not explicitly defined. That is, there is a problem that it is not explicitly defined.
[0011] In Non-Patent Document 2, the output color space and color difference subsampling are not specified, and there is a problem that it is impossible to specify what kind of output can be obtained from the additional information.
[0012] In Non-Patent Document 1 and Non-Patent Document 2, although the types of the input and output tensors can be analyzed from the topology of the neural network, the relationship between the channels of the tensor and the color components cannot be specified. For example, it is not defined how to set the luminance channel and the color difference channel in the input tensor, and in what color space the processed output tensor is output. Therefore, there is a problem that the process cannot be specified and executed. Also 0:0, 4:2:0, 4:2:2, 4:4:4 color difference subsampling, the width and height of the luminance component and the color difference component of the image are different, but it is impossible to specify how to process to derive the input tensor. Also, it is impossible to specify how to generate an image from the output tensor according to the color difference subsampling.
[0013] In Non-Patent Document 3, the relationship between the channels of the tensor and the color components is defined only when the input and output are 4:2:0 color difference subsampling, the input tensor is 10 channels and the output tensor is 6 channels. However, there is a problem that other formats cannot execute the process. Also, there is a problem that it is impossible to support a model that does not require the input of additional information. That is, there is also a problem that it cannot support a model that does not require the input of additional information.
Means for Solving the Problems
[0014] An image decoding apparatus according to an aspect of the present invention is an image decoding apparatus that decodes an image from encoded data, and includes a header decoding unit that derives a filter strength and a bit depth for filter processing, and a neural network filter unit that derives input image information that specifies the format of an input tensor of a neural network model. The neural network filter unit derives a part of the input tensor using an image based on the input image information, sets additional information in another part of the input tensor, derives an output tensor by performing filter processing of the neural network using the input tensor, derives a filtered image from the output tensor, and the additional information is derived using the filter strength and the bit depth.
[0015] An image encoding apparatus according to an aspect of the present invention is an image encoding apparatus that encodes an image, and includes a header encoding unit that derives a filter strength and a bit depth for filter processing, and a neural network filter unit that derives input image information that specifies the format of an input tensor of a neural network model. The neural network filter unit derives a part of the input tensor using an image based on the input image information, sets additional information in another part of the input tensor, derives an output tensor by performing filter processing of the neural network using the input tensor, derives a filtered image from the output tensor, and the additional information is derived using the filter strength and the bit depth.
[0016] An image decoding method according to an aspect of the present invention is an image decoding method for decoding an image from encoded data, which includes deriving a filter strength and a bit depth for filter processing, deriving input image information that specifies the format of an input tensor of a neural network model, deriving a part of the input tensor using an image based on the input image information, setting additional information in another part of the input tensor, deriving an output tensor by performing filter processing of the neural network using the input tensor, and deriving a filtered image from the output tensor.
[0017] An image encoding method according to an aspect of the present invention is an image encoding apparatus for encoding an image, which derives a filter strength and a bit depth for filter processing, derives input image information for specifying a format of an input tensor of a neural network model, derives a part of the input tensor using the image based on the input image information, sets additional information in another part of the input tensor, derives an output tensor by performing filter processing of the neural network using the input tensor, derives a filtered image from the output tensor, and is characterized in that the additional information is derived using the filter strength and the bit depth.
Advantages of the Invention
[0018] With such a configuration, it is possible to refer to the complexity of the neural network model specified by the NN filter without analyzing the neural network model specified by the URI. This has the effect of facilitating the determination of whether a moving image decoding apparatus has the processing ability of a post-filter using a neural network filter.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0020] (First Embodiment) Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0021] FIG. 1 is a schematic diagram showing the configuration of a moving image transmission system according to this embodiment.
[0022] The moving image transmission system 1 encodes encoded data of images with different resolutions whose resolutions are converted, transmits the encoded data, decodes the transmitted encoded data, and inverse-converts the image to the original resolution for display. The moving image transmission system 1 includes a moving image encoding device 10, a network 21, a moving image decoding device 30, and an image display device 41. The moving image encoding device 10 is composed of a resolution conversion device (resolution conversion unit) 51, an image encoding device (image encoding unit) 11, an inverse conversion information creation device (inverse conversion information creation unit) 71, and an inverse conversion information encoding device (inverse conversion information encoding unit) 81.
[0023] The moving image decoding device 30 is composed of an image decoding device (image decoding unit) 31, a resolution inverse conversion device (resolution inverse conversion unit) 61, and an inverse conversion information decoding device (inverse conversion information decoding unit) 91.
[0024] The resolution conversion device 51 converts the resolution of the image T included in the moving image, and supplies a variable resolution moving image T2 including images with different resolutions to the image encoding device 11. Further, the resolution conversion device 51 supplies inverse conversion information indicating the presence or absence of image resolution conversion to the image encoding device 11. When the information indicates resolution conversion, the moving image encoding device 10 sets the resolution conversion information ref_pic_resampling_enabled_flag, which will be described later, to 1, and includes it in the sequence parameter set SPS (Sequence Parameter Set) of the encoded data Te for encoding.
[0025]
[0026] The inverse conversion information creation device 71 creates inverse conversion information based on the image T1 included in the moving image. The inverse conversion information is derived or selected from the relationship between the input image T1 before resolution conversion and the image Td1 after resolution conversion, encoding, and decoding. The additional information is information indicating what to select.
[0027] The inverse conversion information encoding device 81 receives the inverse conversion information. The inverse conversion information encoding device 81 encodes the inverse conversion information to generate encoded inverse conversion information and sends it to the network 21.
[0028] The variable resolution image T2 is input to the image encoding device 11. The image encoding device 11 encodes the image size information of the input image in units of PPS using the RPR (Reference Picture Resampling) framework and sends it to the image decoding device 31.
[0029] In FIG. 1, the inverse conversion information encoding device 81 is not connected to the image encoding device 11, but the inverse conversion information encoding device 81 and the image encoding device 11 may communicate the necessary information as appropriate.
[0030] The network 21 transmits the encoded inverse conversion information and the encoded data Te to the image decoding device 31. Part or all of the encoded inverse conversion information may be included in the encoded data Te as additional extension information SEI. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. The network 21 is not necessarily limited to a two-way communication network and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Also, the network 21 may be replaced by a storage medium that records the encoded data Te such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark). )
[0031] The image decoding device 31 decodes each of the encoded data Te transmitted by the network 21 to generate a variable resolution decoded image Td 1 and supplies it to the resolution inverse conversion device 61 .
[0032] The inverse conversion information decoding device 91 decodes the coded inverse conversion information transmitted by the network 21 to generate inverse conversion information, and supplies the generated inverse conversion information to the resolution inverse conversion device 61 .
[0033] 1, the inverse transformation information decoding device 91 is illustrated separately from the image decoding device 31, but the inverse transformation information decoding device 91 may be included in the image decoding device 31. For example, the inverse transformation information decoding device 91 may be included in the image decoding device 31 separately from each functional unit of the image decoding device 31. Also, although not connected to the image decoding device 31 in FIG. 1, the inverse transformation information decoding device 91 and the image decoding device 31 may communicate necessary information as appropriate.
[0034] When the resolution conversion information indicates resolution conversion, the resolution inverse conversion device 61 generates a decoded image of the original size by inversely converting the resolution-converted image through super-resolution processing using a neural network based on the image size information included in the encoded data.
[0035] The image display device 41 receives one or more decoded images Td2 from the resolution inverse conversion device 61. The image display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. When the display device has a high processing power, it displays a high quality image, and when the display device has a lower processing power, it displays an image that does not require a high processing power or display power.
[0036] FIG. 3 is a conceptual diagram of an image to be processed in the video transmission system shown in FIG. It is a diagram showing the change in the resolution of the image over time. However, in FIG. 3 does not distinguish whether the image is encoded or not. FIG. 3 shows an example of reducing the resolution and transmitting the image to the image decoder 31 in the processing process of the moving image transmission system. As shown in FIG. 3, usually, the resolution conversion device 51 performs a conversion to reduce the resolution of the image in order to reduce the amount of information of the transmitted information.
[0037] <Operator> The operators used in this specification are described below.
[0038] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR , |= is an OR assignment operator, and || represents a logical OR.
[0039] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).
[0040] Clip3(a, b, c) is a function that clips c to a value between a and b. If c < a, it returns a. If c > b, it returns b, and in other cases, it returns c (however, a <= b).
[0041] abs(a) is a function that returns the absolute value of a.
[0042] Int(a) is a function that returns the integer value of a.
[0043] floor(a) is a function that returns the largest integer less than or equal to a.
[0044] ceil(a) is a function that returns the smallest integer greater than or equal to a.
[0045] a / d represents the division of a by d (rounding down the decimal part).
[0046] a^b represents power(a, b). When a = 2, it is equal to 1<<b.
[0047] <Structure of Encoded Data Te> Prior to the detailed description of the image encoding device 11 and the image decoding device 31 according to the present embodiment, the data structure of the encoded data Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be described.
[0048] FIG. 2 is a diagram showing the hierarchical structure of data in the encoded data Te. The encoded data Te exemplarily includes a sequence and a plurality of pictures constituting the sequence. FIG. 2 shows an encoded video sequence that defines a sequence SEQ, an encoded pi cture that defines a picture PICT, an encoded slice that defines a slice S, an encoded slice slice data that defines slice data, an encoded tree unit included in the encoded slice data, and an encoded unit included in the encoded tree unit.
[0049] (Encoded Video Sequence) In the encoded video sequence, a set of data that the image decoding device 31 refers to in order to decode the sequence SEQ to be processed is defined. As shown in FIG. 2, the sequence SEQ includes a video parameter set VPS (Video Parameter Set), a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an Adaptation Parameter Set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).
[0050] In the video parameter set VPS, in a moving image composed of a plurality of layers, A set of encoding parameters common to a plurality of moving images, a plurality of layers included in the moving image, and a set of encoding parameters related to individual layers are defined.
[0051] In the sequence parameter set SPS, a set of encoding parameters that the image decoder 31 refers to for decoding the target sequence is defined. For example, the width and height of a picture are defined. Note that there may be multiple SPSs. In that case, one of the multiple SPSs is selected from the PPS. Select.
[0052] Here, the sequence parameter set SPS includes the following syntax elements. · ref_pic_resampling_enabled_flag: A flag that defines whether to use a function (resampling) that makes the resolution variable when decoding each image included in a single sequence that refers to the target SPS. From another aspect, this flag indicates that the size of the reference picture referred to in the generation of the predicted image changes among the images indicated by a single sequence. When the value of this flag is 1, the above resampling is applied, and when it is 0, it is not applied. · pic_width_max_in_luma_samples: A syntax element that specifies the width of the image with the maximum width among the images in a single sequence in luminance block units. Also, the value of this syntax element is required to be not 0 and an integer multiple of Max(8, MinCbSizeY). Here, MinCbSizeY is a value determined by the minimum size of the luminance block. is required. Here, MinCbSizeY is a value determined by the minimum size of the luminance block. ·pic_height_max_in_luma_samples: A syntax element that specifies, in luminance block units, the height of the image with the maximum height among the images in a single sequence. Also, the value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY). It is required. ·sps_temporal_mvp_enabled_flag: A flag that specifies whether to use temporal motion vector prediction when decoding the target sequence. If the value of this flag is 1, temporal motion vector prediction is used; if the value is 0, temporal motion vector prediction is not used. Also, by specifying this flag, it is possible to prevent the reference coordinate position from shifting when referring to reference pictures with different resolutions, etc. When the value of this flag is 1, temporal motion vector prediction is used; if the value is 0, temporal motion vector prediction is not used. Also, by specifying this flag, it is possible to prevent the reference coordinate position from shifting when referring to reference pictures with different resolutions, etc. When the value of this flag is 1, temporal motion vector prediction is used; if the value is 0, temporal motion vector prediction is not used. Also, by specifying this flag, it is possible to prevent the reference coordinate position from shifting when referring to reference pictures with different resolutions, etc. When the value of this flag is 1, temporal motion vector prediction is used; if the value is 0, temporal motion vector prediction is not used. Also, by specifying this flag, it is possible to prevent the reference coordinate position from shifting when referring to reference pictures with different resolutions, etc.
[0053] In the picture parameter set PPS, a set of encoding parameters that the image decoding device 31 refers to in order to decode each picture in the target sequence is defined. For example, it includes the reference value of the quantization width (pic_init_qp_minus26) used for picture decoding and the flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence. In the picture parameter set PPS, a set of encoding parameters that the image decoding device 31 refers to in order to decode each picture in the target sequence is defined. For example, it includes the reference value of the quantization width (pic_init_qp_minus26) used for picture decoding and the flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence. In the picture parameter set PPS, a set of encoding parameters that the image decoding device 31 refers to in order to decode each picture in the target sequence is defined. For example, it includes the reference value of the quantization width (pic_init_qp_minus26) used for picture decoding and the flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected from each picture in the target sequence. In that case, one of the multiple PPSs is selected from each picture in the target sequence.
[0054] Here, the picture parameter set PPS includes the following syntax elements. ·pic_width_in_luma_samples: A syntax element that specifies the width of the target picture. It is required that the value of the syntax element is not 0, is an integer multiple of Max(8, MinCbSizeY), and is a value less than or equal to pic_width_max_in_luma_samples. · pic_height_in_luma_samples: A syntax element that specifies the height of the target picture. It is required that the value of the syntax element is not 0, is an integer multiple of Max(8, MinCbSizeY), and is a value less than or equal to pic_height_max_in_luma_samples. · conformance_window_flag: Conformance (cropping) window offset A flag indicating whether subsequent parameters are notified, and a flag indicating the location where the conformance window is displayed. When this flag is 1, the parameter is notified, and when it is 0, it indicates that there is no conformance window offset parameter. This is indicated. · conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, conf_win_bottom_offset: Offset values for specifying the left, right, top, and bottom positions of the picture output in the decoding process for a rectangular area specified in the output picture coordinates. Also, when the value of conformance_window_flag is 0, the values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are presumed to be 0. ·scaling_window_flag: A flag indicating whether the scaling window offset parameter exists in the target PPS, and it is a flag related to the definition of the output image size. If this flag is 1, it indicates that the parameter exists in the PPS; if this flag is 0, it indicates that the parameter does not exist in the PPS. Also, when the value of ref_pic_resampling_enabled_flag is 0, the value of scaling_window_flag is also required to be 0. ·scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, scaling_win_bottom_offset: Syntax elements that specify, in luminance sample units, the offsets applied to the image size for calculating the scaling ratio for the left, right, top, and bottom positions of the target picture, respectively. Also, when the value of scaling_window_flag is 0, the values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset are presumed to be 0. Also, the value of scaling_win_left_offset + scaling_win_right_offset is required to be less than pic_width_in_luma_samples, and the value of scaling_win_top_offset + scaling_win_bottom_offset is required to be less than pic_height_in_luma_samples.
[0055] The width PicOutputWidthL and height PicOutputHeightL of the output picture are derived as follows.
[0056] PicOutputWidthL = pic_width_in_luma_samples - (scaling_win_right_offset + scaling_win_left_offset) PicOutputHeightL = pic_height_in_pic_size_units - (scaling_win_bottom_offset +scaling_win_top_offset) (Sub-picture) The picture may be further divided into rectangular sub-pictures. The size of the sub-picture may be a multiple of the CTU. The sub-picture is defined as a set of vertically and horizontally continuous integer tiles. That is, the picture is divided into rectangular tiles, and the sub-picture is defined as a set of rectangular tiles. The sub-picture may be defined using the ID of the upper-left tile and the ID of the lower-right tile of the sub-picture. Also, the slice header may include sh_subpic_id indicating the ID of the sub-picture which may be included.
[0057] (Encoded Picture) In the encoded picture, a set of data that the image decoding device 31 refers to in order to decode the picture PICT to be processed is defined. As shown in FIG. 2, the picture PICT includes a picture header PH, slices 0 to NS - 1 (NS is the total number of slices included in the picture PICT) .
[0058] Hereinafter, when it is not necessary to distinguish each of slices 0 to NS - 1, the subscript of the code may be omitted in the description. The same applies to the data included in the encoded data Te described below and other data with subscripts .
[0059] The picture header includes the following syntax elements ·pic_temporal_mvp_enabled_flag: A flag that specifies whether to use temporal motion vector prediction for inter prediction of the slice associated with the picture header. When the value of the flag is 0, the syntax elements of the slice associated with the picture header is restricted so that temporal motion vector prediction is not used in decoding the slice. When the value of the flag is 1, decoding of the slice associated with the picture header indicates that temporal motion vector prediction is used. Also, when the flag is not defined, it is presumed that the value is 0.
[0060] (Encoded slice) In the encoded slice, a set of data that the image decoding device 31 refers to in order to decode the slice S to be processed is defined. As shown in FIG. 2, a slice includes a slice header, and also includes slice data.
[0061] The slice header includes a set of encoding parameters that the image decoding device 31 refers to in order to determine the decoding method of the target slice. The slice type designation information (slice_type) that designates the slice type is an example of the encoding parameters included in the slice header.
[0062] Slice types that can be specified by the slice type designation information include: (1) I slice that uses only intra prediction during encoding, (2) P slice that uses single prediction (L0 prediction) or intra prediction during encoding, and (3) B slice that uses single prediction (L0 prediction or L1 prediction), bi-prediction, or intra prediction during encoding. Note that inter prediction is not limited to single prediction and bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P and B slices, it means a slice including a block that can use inter prediction.
[0063] Note that the slice header may include a reference to the picture parameter set PPS (pic_parameter_set_id).
[0064] (Encoded slice data) In the symbolized slice data, a set of data that the image decoding device 31 refers to in order to decode the slice data to be processed is defined. The slice data includes the coded slice in FIG. 2 As shown in the header, it contains CTUs. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice, and is sometimes also called the largest coding unit (LCU).
[0065] (Coding Tree Unit) In FIG. 2, a set of data that the image decoding device 31 refers to in order to decode the CTU to be processed is defined. A CTU is divided into coding units (CUs), which are the basic units of coding processing, by recursive quadtree partitioning (QT (Quad Tree) partitioning), binary tree partitioning (BT (Binary Tree) partitioning), or ternary tree partitioning (TT (Ternary Tree) partitioning). The combination of BT partitioning and TT partitioning is called multi-tree partitioning (MT (Multi Tree) partitioning). A node of the tree structure obtained by recursive quadtree partitioning is called a coding node. The intermediate nodes of the quadtree, binary tree, and ternary tree are coding nodes, and the CTU itself is also defined as the topmost coding node.
[0066] CT, as CT information, includes a CU split flag (split_cu_flag) indicating whether to perform CT splitting, a QT split flag (qt_split_cu_flag) indicating whether to perform QT splitting, an MT split direction (mtt_split_cu_vertical_flag) indicating the split direction of MT splitting, and an MT split type (mtt_split_cu_binary_flag) indicating the split type of MT splitting. split_cu_flag, qt_split_cu_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary_flag are transmitted for each coding node.
[0067] Trees that differ in terms of luminance and color difference may be used. The type of tree is indicated by treeType. For example, when using a common tree for both luminance (Y, cIdx = 0) and color difference (Cb / Cr, cIdx = 1, 2), the common single tree is indicated by treeType = SINGLE_TREE. When using two different trees (DUAL trees) for luminance and color difference, the luminance tree is indicated by treeType = DUAL_TREE_LUMA, and the color difference tree is indicated by treeType = DUAL_TREE_CHROMA.
[0068] (Coding Unit) FIG. 2 shows the data referred to by the image decoding apparatus 31 for decoding the coding unit to be processed. Specifically, the CU is composed of a CU header CUH, prediction parameters, transform parameters, quantized transform coefficients, etc. The prediction mode, etc. are defined in the CU header.
[0069] The prediction process may be performed in units of CUs or in units of sub-CUs obtained by further dividing a CU. When the sizes of the CU and the sub-CU are equal, there is one sub-CU in the CU. When the CU is larger than the size of the sub-CU, the CU is divided into sub-CUs. For example, when the CU is 8x8 and the sub-CU is 4x4, the CU is divided into four sub-CUs consisting of two horizontal divisions and two vertical divisions.
[0070] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is prediction within the same picture, and inter prediction refers to prediction processing performed between different pictures (for example, between display times, between layer images).
[0071] The transform and quantization processing is performed in units of CUs, but the quantized transform coefficients may be entropy-coded in units of sub-blocks such as 4x4.
[0072] (Prediction Parameters) The predicted image is derived by prediction parameters associated with blocks. The prediction parameters include prediction parameters for intra prediction and inter prediction.
[0073] Hereinafter, the prediction parameters for inter prediction will be described. The inter prediction parameters are composed of a prediction list usage flag predFlagL0 and predFlagL1, a reference picture index refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether a reference picture list (L0 list, L1 list) is used. When the value is 1, the corresponding reference picture list is used. In this specification, when it is described as "a flag indicating whether XX", if the flag is other than 0 (for example, 1) when XX, and 0 when not XX, 1 is treated as true and 0 is treated as false in logical negation, logical product, etc. (the same applies hereinafter). However, in an actual apparatus or method, other values can also be used as true values and false values.
[0074] Syntax elements for deriving inter prediction parameters include, for example, an affine flag affine_flag, a merge flag merge_flag, a merge index merge_idx, an MMVD flag mmvd_flag, and an inter prediction identifier inter_pred_idc for selecting a reference picture used in the AMVP mode, a reference picture index refIdxLX, a prediction vector index mvp_LX_idx for deriving a motion vector, a differential vector mvdLX, and a motion vector precision mode amvr_mode.
[0075] (Reference Picture List) The reference picture list is a list composed of reference pictures stored in the reference picture memory 306. FIG. 4 is a conceptual diagram showing an example of a reference picture and a reference picture list. In the conceptual diagram showing an example of the reference picture in FIG. 4, the rectangle is a picture, and the arrow is the Reference relationship, the horizontal axis is time, and I, P, and B in the rectangle are intra picture, uni-prediction picture, and bi-prediction picture, respectively. The numbers in the predicted pictures and rectangles indicate the decoding order. As shown in the figure, the decoding order of the pictures is I0, P1, B2, B3, B4, and the display order is I0, B3, B2, B4, P1. In FIG. 4, picture B3 13 shows an example of a reference picture list of a target picture (a target picture). A reference picture list is a list indicating candidates for a reference picture, and one picture (slice) may have one or more reference picture lists. In the example shown in the figure, a target picture B3 has L0 lists RefPicList0 and Each CU has two reference picture lists: L1 list RefPicList1 and L2 list RefPicList2. refIdxLX indicates which picture in the picture list RefPicListX (X=0 or 1) is actually referenced. The figure shows an example where refIdxL0=2 and refIdxL1=0. Note that LX is a notation method used when there is no distinction between L0 prediction and L1 prediction, and hereafter, parameters for the L0 list and parameters for the L1 list will be distinguished by replacing LX with L0 and L1.
[0076] (Merge prediction and AMVP prediction) The decoding (encoding) method of prediction parameters includes a merge prediction mode and an AMVP (Advanced Motion Vector Prediction) mode, and merge_flag is a flag for identifying these. The merge prediction mode is a mode that derives from the prediction parameters of neighboring blocks that have already been processed without including the prediction list utilization flag predFlagLX, the reference picture index refIdxLX, and the motion vector mvLX in the encoded data. The AMVP mode is a mode that includes inter_pred_idc, refIdxLX, and mvLX in the encoded data. Note that mvLX is encoded as mvp_LX_idx that identifies the prediction vector mvpLX and the difference vector mvdLX. In addition to the merge prediction mode, there may be an affine prediction mode and an MMVD prediction mode.
[0077] inter_pred_idc is a value indicating the type and number of reference pictures, and takes any value of PRED_L0, PRED_L1, or PRED_BI. PRED_L0 and PRED_L1 indicate single prediction using one reference picture managed in the L0 list and the L1 list, respectively. PRED_BI indicates bi-prediction using two reference pictures managed in the L0 list and the L1 list.
[0078] merge_idx is an index indicating which prediction parameter among the prediction parameter candidates (merge candidates) derived from the blocks for which the processing has been completed is to be used as the prediction parameter of the target block.
[0079] (Motion Vector) mvLX indicates the shift amount between blocks on two different pictures. The prediction vector and the difference vector regarding mvLX are called mvpLX and mvdLX, respectively.
[0080] (Inter Prediction Identifier inter_pred_idc and Prediction List Utilization Flag predFlagLX) The relationship between inter_pred_idc, predFlagL0, and predFlagL1 is as follows and they are mutually convertible: inter_pred_idc = (predFlagL1<<1)+predFlagL0 predFlagL0 = inter_pred_idc & 1 predFlagL1 = inter_pred_idc >> 1 Note that for the inter-prediction parameters, the prediction list utilization flag may be used, or the inter-prediction identifier may be used. Also, the determination using the prediction list utilization flag may be replaced with the determination using the inter-prediction identifier. Conversely, the determination using the inter-prediction identifier may be replaced with the determination using the prediction list utilization flag.
[0081] (Configuration of the Image Decoding Apparatus) The configuration of the image decoding apparatus 31 (Fig. 5) according to this embodiment will be described.
[0082] The image decoding apparatus 31 includes an entropy decoding unit 301, a parameter decoding unit (predicted image decoding apparatus) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a predicted image generation unit (predicted image generation apparatus) 308, an inverse quantization / inverse transform unit 311, and an addition unit 312, and a prediction parameter derivation unit 320. Note that, in accordance with the image encoding apparatus 11 described later, there is also a configuration in which the image decoding apparatus 31 does not include the loop filter 305.
[0083] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding It includes a section 3022 (prediction mode decoding section), and the CU decoding section 3022 further includes a TU decoding section 3024. These may be collectively referred to as a decoding module. The header decoding section 3020 decodes parameter set information such as VPS, SPS, PPS, APS, and slice header (slice information) from the encoded data. The CT information decoding section 3021 decodes CT from the encoded data. The CU decoding section 3022 decodes CU from the encoded data. When the TU contains prediction error, the TU decoding section 3024 decodes QP update information (quantization correction value) and quantized prediction error (residual_coding) from the encoded data. Decode.
[0084] When it is not in the skip mode (skip_mode == 0), the TU decoding section 3024 decodes QP update information and quantized prediction error from the encoded data. More specifically, when skip_mode == 0, the TU decoding section 3024 decodes a flag cu_cbp indicating whether the target block contains quantized prediction error. When cu_cbp is 1, it decodes the quantized prediction error. If cu_cbp does not exist in the encoded data it is derived as 0.
[0085] The TU decoding section 3024 decodes an index mts_idx indicating the conversion basis from the encoded data. Also, the TU decoding section 3024 decodes an index stIdx indicating the use of secondary conversion and the conversion basis from the encoded data. When stIdx is 0, it indicates non - application of secondary conversion. When stIdx is 1, it indicates one of the conversions in the set (pair) of secondary conversion bases. When stIdx is 2, it indicates the other conversion in the above pair.
[0086] The predicted image generation section 308 is composed of an inter - predicted image generation section 309 and an intra - predicted image generation section 310. composed of.
[0087] The prediction parameter derivation section 320 is composed of an inter - prediction parameter derivation section 303 and an intra - prediction parameter derivation section 304.
[0088] The entropy decoding unit 301 performs entropy decoding on the encoded data Te input from the outside to decode individual codes (syntax elements). For entropy encoding, there are a method of performing variable-length encoding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and the surrounding situation, and a method of performing variable-length encoding of syntax elements using a predefined table or calculation formula. The former, CABAC (Context Adaptive Binary Arithmetic Coding), stores in memory the CABAC state (dominant symbol type (0 or 1) and probability state index pStateIdx for specifying the probability) of the context. The entropy decoding unit 301 initializes all CABAC states at the start of a segment (tile, CTU row, slice). The entropy decoding unit 301 converts the syntax element into a binary string (Bin String) and decodes each bit of the Bin String. When using a context, a context index ctxInc is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the used context is updated. Bits not using a context are decoded with equal probability (EP, bypass), and derivation of ctxInc and CABAC state updates are omitted. The decoded syntax elements include prediction information for generating a predicted image, prediction errors for generating a differential image, and so on. The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code is, for example, the prediction mode predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. Control of which code to decode is performed based on the instruction of the parameter decoding unit 302.
[0089]
[0090] (Basic Flow) FIG. 6 is a flowchart for explaining the schematic operation of the image decoding apparatus 31.
[0091] (S1100: Decoding Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as VPS, SPS, and PPS from the encoded data.
[0092] (S1200: Decoding Slice Information) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.
[0093] Hereinafter, the image decoding apparatus 31 derives a decoded image for each CTU included in the target picture by repeating the processes from S1300 to S5000 for each CTU.
[0094] (S1300: Decoding CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.
[0095] (S1400: Decoding CT Information) The CT information decoding unit 3021 decodes the CT from the encoded data.
[0096] (S1500: Decoding CU) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data and.
[0097] (S1510: Decoding CU Information) The CU decoding unit 3022 decodes CU information, prediction information, TU split flag split_transform_flag, CU residual flags cbf_cb, cbf_cr, cbf_luma, etc. from the encoded data.
[0098] (S1520: Decoding TU Information) When the TU contains prediction error, the TU decoding unit 3024 performs encoding Decode the QP update information, quantization prediction error, and transform index mts_idx from the data. Note that the QP update information is the difference value from the quantization parameter prediction value qPpred, which is the predicted value of the quantization parameter QP. It is a difference value.
[0099] (S2000: Prediction Image Generation) The prediction image generation unit 308 generates a prediction image for each block included in the target CU based on the prediction information.
[0100] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transformation unit 311 performs inverse quantization and inverse transformation processing for each TU included in the target CU.
[0101] (S4000: Decoded Image Generation) The addition unit 312 adds the prediction image supplied from the prediction image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transformation unit 311 to generate the decoded image of the target CU.
[0102] (S5000: Loop Filter) The loop filter 305 applies loop filters such as a deblocking filter, SAO, and ALF to the decoded image to generate the decoded image.
[0103] (Configuration of Inter-Prediction Parameter Derivation Unit) The inter-prediction parameter derivation unit 303 (motion vector derivation device) derives inter-prediction parameters by referring to the prediction parameters stored in the prediction parameter memory 307 based on the syntax elements input from the parameter decoding unit 302. Also, the inter-prediction parameters are output to the inter-prediction image generation unit 309 and the prediction parameter memory 307. The inter-prediction parameter derivation unit 303 and its internal elements, namely the AMVP prediction parameter derivation unit 3032, merge prediction parameter derivation unit 3036, affine prediction unit 30372, MMVD prediction unit 30373, GPM unit 30377, DMVR unit 30537, and MV addition unit 3038, are common means in the image encoding device and the image decoding device. Therefore, These may be collectively referred to as a motion vector derivation unit (motion vector derivation device).
[0104] The scale parameter derivation unit 30378 provided in the header decoding unit 3020 and the header encoding unit 1110 derives the horizontal scaling ratio RefPicScale[i][j][0] of the reference picture, the vertical scaling ratio RefPicScale[i][j][1] of the reference picture, and RefPicIsScaled[i][j] indicating whether the reference picture is scaled. Here, i indicates whether the reference picture list is the L0 list (i = 0) or the L1 list (i = 1), and j is the value (reference picture) of the L0 reference picture list or the L1 reference picture list, and is derived as follows: RefPicScale[i][j][0] = ((fRefWidth << 14)+(PicOutputWidthL >> 1)) / PicOutputWidthL RefPicScale[i][j][1] = ((fRefHeight << 14)+(PicOutputHeightL >> 1)) / PicOutputHeightL RefPicIsScaled[i][j] = (RefPicScale[i][j][0] != (1<<14)) || (RefPicScale[i][j][1] != (1<<14)) Here, the variable PicOutputWidthL is the value when calculating the horizontal scaling ratio when the coded picture is referenced, and the value obtained by subtracting the left and right offset values from the number of horizontal pixels of the luminance of the coded picture is used. The variable PicOutputHeightL is the value when calculating the vertical scaling ratio when the coded picture is referenced, and the value obtained by subtracting the upper and lower offset values from the number of vertical pixels of the luminance of the coded picture is used. The variable fRefWidth Here, the variable PicOutputWidthL is the value when calculating the horizontal scaling ratio when the coded picture is referenced, and the value obtained by subtracting the left and right offset values from the number of horizontal pixels of the luminance of the coded picture is used. The variable PicOutputHeightL is the value when calculating the vertical scaling ratio when the coded picture is referenced, and the value obtained by subtracting the upper and lower offset values from the number of vertical pixels of the luminance of the coded picture is used. The variable fRefWidth is the horizontal width of the reference picture, and fRefHeight is the vertical height of the reference picture. Let the value of PicOutputWidthL be the reference list value j of list i, and let the variable fRefHight be the value of PicOutputHeightL of the reference picture list value j of list i.
[0105] (MV Addition Unit) The MV addition unit 3038 adds the input mvpLX from the AMVP prediction parameter derivation unit 3032 and the decoded mvdLX to calculate mvLX. The addition unit 3038 outputs the calculated mvLX to the inter prediction image generation unit 309 and the prediction parameter memory 307: mvLX[0] = mvpLX[0]+mvdLX[0] mvLX[1] = mvpLX[1]+mvdLX[1] The loop filter 305 is a filter provided within the encoding loop, and is a filter that removes block distortion and ringing distortion and improves the image quality. The loop filter 305 performs filters such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) on the decoded image of the CU generated by the addition unit 312.
[0106] The DF unit 601 includes a bS derivation unit 602 that derives the strength bS of the deblocking filter in units of pixels, boundaries, and line segments, and a DF filter unit 602 that performs deblocking filter processing to reduce block noise.
[0107] The DF unit 601 derives the edge degree edgeIdc indicating whether there are partition division boundaries, prediction block boundaries, and transform block boundaries in the input image resPicture before the NN processing (processing of the NN filter unit 601) and the maximum filter length maxFilterLength of the deblocking filter. Further, from the edgeIdc, the transform block boundary, and the encoding parameters, the strength bS of the deblocking filter is Derive. The encoding parameters are, for example, the prediction mode CuPredMode, the BDPCM prediction mode intra_bdpcm_luma_flag, a flag indicating whether it is in the IBC prediction mode, the motion vector, the reference picture, Flags such as tu_y_coded_flag and tu_u_coded_flag indicating whether non-zero coefficients exist in the transform block. edgeIdc and bS may take values of 0, 1, 2 or other values.
[0108] The reference picture memory 306 stores the decoded image of the CU at a predetermined position for each target picture and target CU.
[0109] The prediction parameter memory 307 stores the prediction parameters at a predetermined position for each CTU or CU. Specifically, the prediction parameter memory 307 stores the parameters decoded by the parameter decoder 302, the parameters derived by the prediction parameter derivation unit 320, etc.
[0110] The parameters derived by the prediction parameter derivation unit 320 are input to the prediction image generation unit 308. Also, the prediction image generation unit 308 reads the reference picture from the reference picture memory 306. The prediction image generation unit 308 generates a predicted image of a block or sub-block in the prediction mode indicated by predMode, using the parameters and the reference picture (reference pic ture block). Here, the reference picture block is a set of pixels on the reference picture (usually a rectangle, so it is called a block), and is the area referred to for generating the predicted image.
[0111] When predMode indicates the inter prediction mode, the inter prediction image generation unit 309 uses the inter prediction parameters input from the inter prediction parameter derivation unit 303 and the reference picture to generate a predicted image of a block or sub-block by inter prediction.
[0112] (Motion Compensation) The motion compensation unit 3091 (interpolation image generation unit 3091) receives an input from the inter-prediction parameter derivation unit 303 and reads a reference block from the reference picture memory 306 based on the inter-prediction parameters (predFlagLX, refIdxLX, mvLX) to generate an interpolation image (motion-compensated image). The reference block is a block at a position shifted by mvLX from the position of the target block on the reference picture RefPicLX specified by refIdxLX. Here, when mvLX does not have integer precision, a filter for generating pixels at fractional positions, called a motion compensation filter, is applied to generate the interpolation image. First, the motion compensation unit 3091 derives the integer position (xInt, yInt) and phase (xFrac, yFrac) corresponding to the coordinates (x, y) within the prediction block using the following equations:
[0113] xInt = xPb+(mvLX[0]>>(log2(MVPREC)))+x xFrac = mvLX[0]&(MVPREC-1) yInt = yPb+(mvLX[1]>>(log2(MVPREC)))+y yFrac = mvLX[1]&(MVPREC-1) Here, (xPb, yPb) is the upper-left coordinate of a block of size bW*bH, where x = 0…bW-1 and y = 0…bH-1, and MVPREC indicates the precision of mvLX (1 / MVPREC pixel precision). For example, MVPREC = 16.
[0114] The motion compensation unit 3091 performs horizontal interpolation processing on the reference picture refImg using an interpolation filter to derive a temporary image temp[][]. The following Σ is the sum over k from k = 0..NTAP-1, shift1 is a normalization parameter for adjusting the value range, and offset1 = 1<<(shift1-1): temp[x][y] = (ΣmcFilter[xFrac][k]*refImg[xInt+k-NTAP / 2+1][yInt]+offset1)>>shift1 Subsequently, the motion compensation unit 3091 performs vertical interpolation processing on the temporary image temp[][] to derive an interpolated image Pred [][]. The following Σ is the sum with respect to k where k = 0..NTAP-1, and shift2 is used to adjust the value range range normalization parameter, and offset2 = 1<<(shift2-1): Pred[x][y] = (ΣmcFilter[yFrac][k]*temp[x][y+k-NTAP / 2+1]+offset2)>>shift2 In the case of dual prediction, the above Pred[][] is derived for each of the L0 list and the L1 list (referred to as the interpolated images PredL0[][] and PredL1[]), and the interpolated image Pred[][] is generated from PredL0[][] and PredL1[][].
[0115] Note that the motion compensation unit 3091 has a function of scaling the interpolated image according to the horizontal scaling ratio RefPicScale[i][j][0] of the reference picture derived by the scale parameter derivation unit 30378 and the vertical scaling ratio RefPicScale[i][j][1] of the reference picture
[0116] When predMode indicates the intra prediction mode, the intra prediction image generation unit 310 performs intra prediction using the intra prediction parameters input from the intra prediction parameter derivation unit 304 and the reference pixels read from the reference picture memory 306
[0117] The inverse quantization and inverse transformation unit 311 (residual decoding unit) inverse quantizes the quantized transform coefficients input from the parameter decoding unit 302 to obtain the transform coefficients
[0118] The inverse quantization and inverse transformation unit 311 receives the quantized transform coefficients qd[] input from the entropy decoding unit 301 and scales (inverse quantizes) them by the scaling unit 31111 to obtain the transform coefficients d[] .
[0119] The scaling unit 31111 scales the transform coefficients decoded by the TU decoding unit using the quantization parameter and the scaling factor derived in the parameter decoding unit 302, with the weight in coefficient units
[0120] Here, the quantization parameter qP is derived as follows using the color component cIdx of the target transform coefficient and the joint chrominance residual coding flag tu_joint_cbcr_flag
[0121] qP = qPY (cIdx==0) qP = qPCb (cIdx==1 && tu_joint_cbcr_flag==0) qP = qPCr (cIdx==2 && tu_joint_cbcr_flag==0) qP = qPCbCr (tu_joint_cbcr_flag!=0) The scaling unit 31111 derives a value rectNonTsFlag related to the size or shape from the size (nTbW, nTbH) of the target TU
[0122] rectNonTsFlag = (((Log2(nTbW)+Log2(nTbH)) & 1)==1 && transform_skip_flag[xTbY] [yTbY]==0) The transform_skip_flag is a flag indicating whether to skip the transformation
[0123] The scaling unit 31111 performs the following process using the ScalingFactor[][] derived in the scaling list decoding unit 3026 (not shown) .
[0124] When the scaling list is not valid (scaling_list_enabled_flag == 0), or when transform skip is used (transform_skip_flag == 1), the scaling unit 31111 sets m[x][y] = 16. That is, uniform quantization is performed. scaling_list_enabled_flag is a flag indicating whether the scaling list is valid or not.
[0125] In other cases (i.e., when scaling_list_enabled_flag == 1 and transform_skip_flag == 0), the scaling unit 31111 uses the scaling list. Here, m[][] is set as follows.
[0126] m[x][y] = ScalingFactor[Log2(nTbW)][Log2(nTbH)][matrixId][x][y] Here, matrixId is set according to the prediction mode (CuPredMode) of the target TU, the color component index (cIdx), and the applicability of non-separable transform (lfnst_idx).
[0127] The scaling unit 31111 derives the scaling factor ls[x][y] by the following formula when sh_dep_quant_used_flag is 1 in the case.
[0128] ls[x][y] = (m[x][y]*quantScale[rectNonTsFlag][(qP+1)%6]) << ((qP+1) / 6) In other cases (sh_dep_quant_used_flag = 0), it may be derived by the following formula.
[0129] ls[x][y] = (m[x][y] * quantScale[rectNonTsFlag][qP % 6]) << (qP / 6) Here, quantScale[] = {{40, 45, 51, 57, 64, 72}, {57, 64, 72, 80, 90, 102}}. sh_dep_quant_used_flag is a flag that is set to 1 when dependent quantization is performed and 0 when it is not. The value of quantScale is derived by the following formula based on the value of x (x = 0..6).
[0130] quantScale[x] = RoundInt(2 ^ (6 / (x - qsoffset))) qsoffset = rectNonTsFlag == 0? 4 : 2 When the value of qP is 4, quantScale is 64. Here, RoundInt is a function that rounds by adding a rounding constant (e.g., 0. 5) and then truncating the decimal part to obtain an integer.
[0131] Scaling unit 31111 performs inverse quantization by deriving dnc[][] from the product of the scaling factor ls[][] and the decoded transform coefficient TransCoeffLevel.
[0132] dnc[x][y] = (TransCoeffLevel[xTbY][yTbY][cIdx][x][y] * ls[x][y] + bdOffset1) >> bdShift1 Here, bdOffset1 = 1 << (bdShift1 - 1) Finally, scaling unit 31111 clips the inverse quantized transform coefficient to derive d[x][y].
[0133] d[x][y] = Clip3(CoeffMin, CoeffMax, dnc[x][y]) (Equation CLIP - 1) CoeffMin and CoeffMax are the minimum and maximum values for clipping, and are derived by the following formula.
[0134] CoeffMin = -(1 << log2TransformRange) CoeffMax = (1 << log2TransformRange) - 1 Here, log2TransformRange is a value indicating the range of transform coefficients derived by the method described later.
[0135] d[x][y] is transmitted to the inverse core transform unit 31123 or the inverse non-separable transform unit 31121. The inverse non-separable transform unit 31121 applies an inverse non-separable transform to the transform coefficients d[][] after inverse quantization and before core transform.
[0136] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.
[0137] The inverse quantization / inverse transform unit 311 inverse quantizes the quantized transform coefficients input from the parameter decoding unit 302 to obtain the transform coefficients.
[0138] The adder 312 adds the predicted image of the block input from the predicted image generation unit 308 and the prediction error input from the inverse quantization / inverse transform unit 311 for each pixel to generate the decoded image of the block. The adder 312 stores the decoded image of the block in the reference picture memory 306 and also outputs it to the loop filter 305.
[0139] (Configuration example of the NN filter unit 611) FIG. 14 is a diagram showing a configuration example of an interpolation filter, a loop filter, and a post filter using a neural network filter unit (NN filter unit 611). Hereinafter, an example of the post filter will be described, but an interpolation filter or a loop filter may also be used. Hereinafter, an example of the post filter will be described, but an interpolation filter or a loop filter may also be used.
[0140] The post-processing unit 61 after the moving image decoding device includes an NN filter unit 611. When outputting the image of the reference picture memory 306, the NN filter unit 611 performs filtering processing and outputs it externally. The output image may be displayed, written to a file, re-encoded (transcoded), transmitted, etc. The NN filter unit 611 is means for performing filtering processing on the input image by means of a neural network model. At the same time, reduction / enlargement by equal magnification or rational magnification may be performed.
[0141] Here, the neural network model (hereinafter, NN model) means the elements and connection relationships (topology) of the neural network and the parameters (weights, biases) of the neural network. Note that the neural network model may switch only the parameters while fixing the topology.
[0142] (Details of the NN filter unit 611) The NN filter unit uses the input image inSamples and input parameters (for example, QP, bS, etc.) to perform filtering processing by means of a neural network model. The input image may be an image for each component, or an image having a plurality of components as channels respectively. Also, the input parameters may be assigned to channels different from the image.
[0143] The NN filter unit may repeatedly apply the following processing.
[0144] The NN filter unit convolves the inSamples with the kernel k[m][i][j] (conv, convolution) and derives the output image outSamples to which the bias is added. Here, nn = 0..n - 1, xx = 0..width - 1, yy = 0..height - 1.
[0145] outSamples[nn][xx][yy] = ΣΣΣ(k[mm][i][j] * inSamples[mm][xx + i - of][yy + j - of] + bias [nn]) In the case of 1x1 Conv, Σ represents the sum for each mm = 0..m - 1, i = 0, j = 0. At this time, set of = 0 to do. In the case of 3x3 Conv, Σ represents the sum for each mm = 0..m - 1, i = 0..2, j = 0..2. At this time, set of = 1. n is the number of channels of outSamples, m is the number of channels of inSamples, width is the width of inSamples and outSamples, and height is the height of inSamples and outSamples. of is the padding area provided around inSamples to make the sizes of inSamples and outSamples the same size. Hereinafter, when the output of the NN filter unit is not an image but a value (correction value), corrNN is used to represent the output instead of outSamples
[0146] Note that it is equivalent to the following processing when described using inTensor and outTensor in CHW format instead of inSamples and outSamples in CWH format
[0147] outTensor[nn][yy][xx] = ΣΣΣ(k[mm][i][j] * inTensor[mm][yy + j - of][xx + i - of] + bias[nn]) Also, the following process shown by the formula called Depth wise Conv may be performed. Here, nn = 0..n - 1, xx = 0..width - 1, yy = 0..height - 1
[0148] outSamples[nn][xx][yy] = ΣΣ(k[nn][i][j] * inSamples[nn][xx + i - of][yy + j - of] + bias[nn]) Σ represents the sum for each i and j. n is the number of channels of outSamples and inSamples, width is the width of inSamples and outSamples, and height is the height of inSamples and outSamples.
[0149] Also, a non-linear process called Activate, such as ReLU, may be used. ReLU(x) = x >= 0? x : 0 Also, leakyReLU shown in the following equation may be used.
[0150] leakyReLU(x) = x >= 0? x : a * x Here, a is a predetermined value, for example, 0.1 or 0.125. Also, in order to perform integer operations, all the values of k, bias, and a above may be made integers, and a right shift may be performed after conv.
[0151] In ReLU, values less than 0 are always 0, and for values greater than or equal to 0, the input value is output as it is. On the other hand, in leakyReLU, for values less than 0, linear processing is performed with the gradient set by a. In ReLU, the gradient for values less than 0 disappears, so learning may become difficult to proceed. In leakyReLU, the gradient for values less than 0 remains, making the above problem less likely to occur. Also, among the above leakyReLU(x), PReLU that uses the value of a parametrically may be used.
[0152] (SEI for reference to neural network model complexity) Figure 9 is a diagram showing the configuration of the syntax table of the NN filter SEI of the present embodiment. This SEI includes information on the complexity of the neural network model. ·nnrpf_id: It is the identification number of the NN filter. ·nnrpf_mode_idc: The index of the mode indicating how to specify the neural network model to be used for the NN filter. When the value is 0, it indicates that the NN filter associated with nnrpf_id is not specified in this SEI message. When the value is 1, it indicates that the NN filter associated with nnrpf_id is a neural network model identified by a predetermined URI (Uniform Resource Identifier). The URI is a string for identification indicating a logical or physical resource . Note that the actual data does not need to exist at the location indicated by the URI , and it is sufficient that the string can identify the resource. When the value is 2, the NN filter associated with nnrpf_id is a neural network model represented by the ISO / IEC 15938-17 bit stream included in this SEI message . When the value is 3, it indicates that the NN filter associated with nnrpf_id is a neural network model identified by the NN filter SEI message used in the previous decoding and updated by the ISO / IEC 15938-17 bit stream included in this SEI message. ·nnrpf_persistence_flag: A flag that specifies the persistence of this SEI message for the current layer . When the value is 0, it indicates that this SEI message is applied only to the currently decoded picture. When the value is 1, it indicates that it is applied in output order to the currently decoded picture and subsequent pictures. ·nnrpf_uri[i]: The reference URI of the neural network model to be used as the NN filter is a string to store. i is the i-th byte of a NULL-terminated UTF-8 string. When nnrpf_mode_idc == 1, the header encoding unit 1110 and the header decoding unit 3020 decode the nnrpf_uri which is a URI indicating the neural network model to be used as an NN filter. The neural network model corresponding to the string indicated by the moving image encoding device or the moving image is read from the memory provided in the decoding device, or read from the outside via the network. ·nnrpf_payload_byte[i]: Indicates the i-th byte of the bitstream conforming to ISO / IEC 15938-17.
[0153] The NN filter SEI includes the following syntax elements as neural network model complexity information (network model complex ity information). ·nnrpf_parameter_type_idc: An index indicating the variable type included in the parameters of the NN model. When the value is 0, the NN model uses only integer types. When the value is 1, the NN model uses floating point types or integer types. ·nnrpf_num_parameters_idc: An index indicating the number of parameters of the NN model used by the post-filter. When the value is 0, it indicates that the number of parameters of the NN model is not defined. When the value is not 0, the following processing is performed using nnrpf_num_parameters_idc to derive the number of pa rameters of the NN model.
[0154] The header encoding unit 1110 and the header decoding unit 3020 derive the maximum value MaxNNParameters of the number of parameters of the NN model as follows based on nnrpf_num_parameters_idc, and may encode and decode the network model complexity information.
[0155] MaxNNParameters = (UNITPARAM << nnrpf_num_parameters_idc) - 1 Here, UNITPARAM is a predetermined constant, and UNITPARAM = 2048 = 2^11 may be used. Note that the shift operation is equivalent to an exponent, and the following may also be used.
[0156] MaxNNParameters = 2 ^(nnrpf_num_parameters_idc+11) - 1 Also, the unit of the number of parameters may not be a double unit, but a combination of a double unit and a 1.5-fold unit as follows. It may be combined.
[0157] MaxNNParameters = (nnrpf_num_parameters_idc & 1)? (UNITPARAM2 << nnrpf_num_parameters_idc)-1 : (UNITPARAM << nnrpf_num_parameters_idc)-1 Here, UNITPARAM2 may be a predetermined constant such that UNITPARAM2 = UNITPARAM * 1.5. For example, when UNITPARAM = 2048, UNITPARAM2 = 3072. Also, the following may be used.
[0158] MaxNNParameters = (nnrpf_num_parameters_idc & 1)? 2^(nnrpf_num_parameters_idc+11)*1.5-1 : 2^(nnrpf_num_parameters_idc+11)-1 That is, in the header encoding unit 1110, the value of nnrpf_num_parameter_idc is set to a value such that the number of parameters of the actual NN model is equal to or less than MaxNNParameters. The header decoding unit 3020 decodes the encoded data set as described above. The header decoding unit 3020 decodes the encoded data set as described above. decodes the encoded data set set as above.
[0159] Note that a linear expression may be used for the derivation of MaxNNParameters.
[0160] MaxNNParameters = (UNITPARAM * nnrpf_num_parameters_idc) - 1 At this time, UNITPARAM may be 10000. UNITPARAM preferably has a value of 1000 or more, and is preferably a multiple of 10. ·nnrpf_num_kmac_operations_idc: A value indicating the scale of the number of operations required for the processing of the post-filter. The header encoding unit 1110 and the header decoding unit 3020 obtain MaxNNOperations as follows based on nnrpf_num_kmac_operations_idc. MaxNNOperations is the maximum value of the number of operations required for the processing of the post-filter.
[0161] MaxNNOperations = nnrpf_num_kmac_operations_idc * 1000 * pictureWidth * pictureHeight Here, pictureWidth and pictureHeight are the width and height of the picture input to the post-filter.
[0162] That is, the moving image encoding device sets the value of nnrpf_num_kmac_operations_idc according to the number of operations required for the processing of the post-filter.
[0163] As described above, by transmitting, encoding, or decoding the syntax of the network model complexity information related to the processing amount defined with a predetermined constant as a unit, there is an effect that the complexity can be transmitted simply. Further, using a multiple of 10 has the effect of making the value easy for humans to understand. Also, if the value is 1000 or more, it is possible to suitably represent the scale of the model with a small number of quantization levels and transmit it efficiently.
[0164] That is, the syntax indicating the above network model complexity information indicates the upper limit of the number of parameters or the number of operations, and the above number of parameters or the number of operations is defined with a power of 2 as a unit. Alternatively, the above number of parameters or the number of operations may be defined with a power of 2 or 1.5 times the power of 2 as a unit. Also, the above number of parameters or the number of operations may be defined with a multiple of 10 as a unit.
[0165] As described above, further, by transmitting, encoding, or decoding the syntax of the network model complexity information related to the processing amount defined by the shift representation or the exponent representation, there is an effect that the complexity can be transmitted efficiently with a shorter code. ·nnrpf_alignment_zero_bit: Bits for byte alignment. The header encoding unit 1110 and the header decoding unit 3020 encode and decode the code "0" bit by bit until the bit position reaches the byte boundary.
[0166] nnrpf_operation_type_idc: An index indicating the limitation of elements or the topology limitation used in the NN model of the post-filter. Depending on the value of the index, for example, the following processing may be performed.
[0167] When the value is 3, the element or topology is limited as follows. The maximum size of the kernel is 5x5, the maximum number of channels is 32, and only leaky ReLU or ReLU can be used as the activation function. The maximum level of the branch is 3 (excluding skip connections).
[0168] When the value is 2, in addition to the above restrictions, the use of leaky Relu in the activation function is prohibited, and branches other than skip connections are prohibited (for example, U-Net or grouped convolutions are not performed). When the value is 1, in addition to the above restrictions, mapping from space to channel (for example, Pixel Shuffler) is prohibited, and global average pooling is prohibited.
[0169] When the value is 0, the element or topology is not restricted.
[0170] When the value is 0, the element or topology is not restricted. <Another Configuration Example 1> The parameter nnrpf_parameter_type_idc indicating the network model complexity information may be defined as follows. ·nnrpf_parameter_type_idc: An index indicating the parameter type of the neural network model. For example, the parameter type may be determined as follows according to the parameter value. When the value is 0, it is defined as an 8-bit unsigned integer type; when the value is 1, it is defined as a 16-bit unsigned integer type; when the value is 2, it is defined as a 32-bit unsigned integer type; when the value is 3, it is defined as a 16-bit floating-point type (bfloat16); when the value is 4, it is defined as a 16-bit floating-point type (half-precision); when the value is 5, it is defined as a 32-bit floating-point type (single-precision). <Another Configuration Example 2> FIG. 10 is a diagram showing the structure of the syntax table of the NN filter SEI having network model complexity information. In this example, the definition of the following syntax information in the SEI of FIG. 9 is changed.
[0171] In this configuration example, among the SEIs in FIG. 9, the definition of the following syntax information is changed. ·nnrpf_parameter_type_idc: An index indicating the numerical type of the neural network model. For example, the numerical type may be determined as follows according to the value of the parameter. When the value is 0, it is defined as an integer type, and when the value is 1, it is defined as a floating-point type.
[0172] In addition to the syntax of the SEI in FIG. 9, the following syntax information is included. ·nnrpf_parameter_bit_width_idc: An index indicating the bit width of the neural network model. For example, the bit width may be determined as follows according to the value of the parameter. When the value is 0, it is defined as 8 bits, when the value is 1, it is defined as 16 bits, and when the value is 2, it is defined as 32 bits. <Another configuration example 3> FIG. 11 is a diagram showing the configuration of a syntax table of an NN filter SEI having network model complexity information. Here, the bit width of the parameter type of the neural network model is defined in logarithmic representation. In this configuration example, instead of nnrpf_parameter_bit_width_idc, the following syntax information is included.
[0173] ·nnrpf_log2_parameter_bit_width_minus3: A value indicating the bit width of the parameter of the neural network model in logarithmic representation to the base 2. Based on nnrpf_log2_parameter_bit_width_minus3, the bit width parameterBitWidth of the parameter is obtained as follows. parameterBitWidth = 1 << ( nnrpf_log2_parameter_bit_width_minus3 + 3 ) (Decoding of SEI and post-filter processing) The header decoding unit 3020 decodes the network model complexity information from the SEI messages defined in FIGS. 9 to 11. SEI is additional information for processes related to decoding, display, etc.
[0174] FIG. 12 is a diagram showing a flowchart of the processing of the NN filter unit 611. The NN filter unit 611 performs the following processing according to the parameters of the above SEI message. S6001: Read the processing amount and accuracy from the network model complexity information of the SEI. S6002: End if the complexity exceeds the complexity that the NN filter unit 611 can process. If not, proceed to S6003. S6003: End if the accuracy exceeds the accuracy that the NN filter unit 611 can process. If not, proceed to S6004. S6004: Identify the network model from the SEI and set the topology of the NN filter unit 611 . S6005: Derive the parameters of the network model from the update information of the SEI. S6006: Read the derived parameters of the network model into the NN filter unit 611. S6007: Execute the filter processing of the NN filter unit 611 and output it externally. However, the SEI is not necessarily required for constructing luminance samples and color difference samples in the decoding process to be.
[0175] (SEI for reference to neural network model data format) FIG. 16 is a diagram showing another configuration of the syntax table of the NN filter SEI of the present embodiment. This SEI includes information on the data format of the neural network model (NN model). The same syntax elements as the SEI including the neural network model complexity information already described are omitted from the description ·nnrpf_input_format_idc: Input tensor identification parameter. NN used in the NN filter This indicates the format of the input data (input tensor) of the model. The header decoding unit 3020 derives the format of the input data (number of channels (NumInChannels), data format) based on the value of nnrpf_input_format_idc, as shown in FIG.
[0176] If nnrpf_input_format_idc==0, the input data format is 1-channel (luminance) 3D data. The luminance channel of the decoded image is used as input data for the NN filter. In this embodiment, the three dimensions of the three-dimensional data are (C, H, W). Although this is defined as an order, the order of dimensions is not limited to this. For example, the order of dimensions is (H, W, C). In this case, since there is only one channel, the data is two-dimensional (H,W) and You may do so.
[0177] If nnrpf_input_format_idc==1, the input data format is 2-channel (color difference) 3D data. The two chrominance channels (U and V) of the decoded image are the inputs of the NN filter. This indicates that it is used as force data.
[0178] If nnrpf_input_format_idc==2, the input data format is 3-channel (luminance and 2 This shows that the luma and two chroma channels of the decoded image are used as input data to the NN filter in 4:4:4 format.
[0179] When nnrpf_input_format_idc==3, the input data format is 6-channel (4 luminance and 2 chrominance) 3D data (3D tensor). It indicates that 4 channels derived from the luminance channel of the decoded image in 4:2:0 format and 2 chrominance channels are used as input data for the NN filter.
[0180] ·nnrpf_output_format_idc: Output tensor identification parameter. It indicates the format of the output data (NN output data, output tensor) of the NN model used in the NN filter. Header decoding unit 3020 As shown in FIG. 18, based on the value of nnrpf_output_format_idc, the format of the output data (channel number (NumOutChannels), data format) is derived. (NumOutChannels), data format) is derived.
[0181] When nnrpf_output_format_idc == 0, the format of the NN output data is 3D data (3D tensor) of 1 channel (luminance). It indicates that the output data of the NN filter is used as the luminance channel of the output image.
[0182] When nnrpf_output_format_idc == 1, the format of the NN output data is 3D data (3D tensor) of 2 channels (color difference). It indicates that the output data of the NN filter is used as the two color difference channels (U and V) of the output image.
[0183] When nnrpf_output_format_idc == 2, the format of the NN output data is 3D data (3D tensor) of 3 channels (luminance and two color differences). It indicates that the output data of the NN filter is used in 4:4:4 format as the luminance and two color difference channels of the output image. as the luminance and two color difference channels in 4:4:4 format.
[0184] When nnrpf_output_format_idc == 3, the format of the NN output data is 3D data (3D tensor) of 6 channels (4 luminances and 2 color differences). The output data of the NN filter is used in 4:2:0 format as one luminance channel and two color difference channels derived by integrating 4 channels out of 6 channels. channels derived by integrating 4 channels out of 6 channels. and two color difference channels in 4:2:0 format.
[0185] (Processing of Post-filter SEI) In the processing of the post-filter SEI, the header encoding unit 1110 and the header decoding unit 3020 may set the syntax value and variables of the image decoder to variables for filter processing. PicWidthInLumaSamples = pps_pic_width_in_luma_samples PicHeightInLumaSamples = pps_pic_height_in_luma_samples ChromaFormatIdc = sps_chroma_format_idc BitDepthY = BitDepthC = BitDepth ComponentSample[cIdx] is a two-dimensional array storing the decoded pixel values of the cIdx-th component of the decoded image. Here, pps_pic_width_in_luma_samples, pps_pic_height_in_luma_samples, sps_chroma_format_idc are syntax values indicating the image width, height, and subsampling of color components, and BitDepth is the bit depth of the image.
[0186] The header encoding unit 1110 and the header decoding unit 3020 derive the following variables according to ChromaFormatIdc as follows. SubWidthC = 1, SubHeghtC = 1 (ChromaFormatIdc == 0) SubWidthC = 2, SubHeghtC = 2 (ChromaFormatIdc == 1) SubWidthC = 2, SubHeghtC = 1 (ChromaFormatIdc == 2) SubWidthC = 1, SubHeghtC = 1 (ChromaFormatIdc == 3) The header encoding unit 1110 and the header decoding unit 3020 derive the width and height of the luminance image to be filtered, and the width and height of the color difference image using the following variables. LumaWidth = PicWidthInLumaSamples LumaHeight = PicHeightInLumaSamples ChromaWidth = PicWidthInLumaSamples / SubWidthC ChromaHeight = PicHeightInLumaSamples / SubHeithtC SW = SubWidthC SH = SubHeightC SubWidthC (=SW) and SubHeightC (=SH) indicate the subsampling of the color component. Here it is a variable representing the ratio of the resolution of the color difference to the luminance.
[0187] Also, the header encoding unit 1110 and the header decoding unit 3020 may derive the output image width outWidth and height outHeight according to the scale values indicating the ratio as follows. outLumaWidth = LumaWidth * scale outLumaHeight = LumaHeight * scale outChromaWidth = LumaWidth * scale / outSW outChromaHeight = LumaHeight * scale / outSH outSW and outSH are the color difference subsampling values of the output image.
[0188] (Conversion to the NN input data of the post-filter) When inputting image data into the NN filter unit 611, based on the value of nnrpf_input_format_idc, the decoded image is converted into the NN input data inputTensor[][][], which is a three-dimensional array, as shown in FIG. 19 and below. Hereinafter, x and y represent the coordinates of the luminance pixels. For example, in ComponentSample, the ranges of x and y are x = 0..LumaWidth - 1 and y = LumaHeight - 1, respectively. cx and cy represent the coordinates of the chrominance pixels, and the ranges of cx and cy are cx = 0..ChromaWidth - 1 and cy = ChromaHeight - 1, respectively. The following NN filter unit 611 processes this range. When nnrpf_input_format_idc is 0 (pfp_component_idc == 0), the NN filter unit 611 derives inputTensor as follows.
[0189] inputTensor[0][y][x] = ComponentSample[0][x][y] When nnrpf_input_format_idc is 1 (pfp_component_idc == 1), inputTensor is derived as follows: inputTensor[0][cy][cx] = ComponentSample[1][cx][cy] inputTensor[1][cy][cx] = ComponentSample[2][cx][cy] Or the following may also be used: inputTensor[0][y / SH][x / SW] = ComponentSample[1][x / SW][y / SH] inputTensor[1][y / SH][x / SW] = ComponentSample[2][x / SW][y / SH] When nnrpf_input_format_idc is 2 (pfp_component_idc == 2), inputTensor is derived as follows: inputTensor[0][y][x] = ComponentSample[0][x][y] ChromaOffset = 1<<(BitDepthC-1) inputTensor[1][y][x] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[1][x / SW][y / SH] inputTensor[2][y][x] = ChromaFormatIdc==0? ChromaOffset : ComponentSample [2][x / SW][y / SH] When nnrpf_input_format_idc is 3, inputTensor is derived as follows: inputTensor[0][cy][cx] = ComponentSample[0][cx*2 ][cy*2 ] inputTensor[1][cy][cx] = ComponentSample[0][cx*2+1][cy*2 ] inputTensor[2][cy][cx] = ComponentSample[0][cx*2 ][cy*2+1] inputTensor[3][cy][cx] = ComponentSample[0][cx*2+1][cy*2+1] ChromaOffset = 1<<(BitDepthC-1) inputTensor[4][cy][cx] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[1][cx][cy] inputTensor[5][cy][cx] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[2][cx][cy] When ChromaFormatIdc is 0, ComponentSample is an image of only the luminance channel. In this case, the NN filter unit 611 may set a constant ChromaOffset derived from the bit depth to the color difference data part of the inputTensor. ChromaOffset may be other values such as 0. Alternatively, as shown in the parentheses, the NN input data may be derived according to pfp_component_idc described later.
[0190] Figure 22 is a diagram showing the relationship between the number of channels of the tensor and the conversion process.
[0191] Furthermore, the NN filter unit 611 may derive the NN input data as follows according to the number of channels NumInChannels(numTensors) of the input tensor. When NumInChannels == 1, the inputTensor is derived as follows.
[0192] inputTensor[0][y][x] = ComponentSample[0][x][y] When NumInChannels == 2, the inputTensor is derived as follows.
[0193] inputTensor[0][cy][cx] = ComponentSample[1][cx][cy] inputTensor[1][cy][cx] = ComponentSample[2][cx][cy] Or the following may also be used: inputTensor[0][y / SH][x / SW] = ComponentSample[1][x / SW][y / SH] inputTensor[1][y / SH][x / SW] = ComponentSample[2][x / SW][y / SH] When NumInChannels == 3, the inputTensor is derived as follows.
[0194] inputTensor[0][y][x] = ComponentSample[0][x][y] ChromaOffset = 1<<(BitDepthC - 1) inputTensor[1][y][x] = ChromaFormatIdc == 0? ChromaOffset : ComponentSample[1][x / SW][y / SH] inputTensor[2][y][x] = ChromaFormatIdc == 0? ChromaOffset : ComponentSample[2][x / SW][y / SH] When NumInChannels == 6, derive inputTensor as follows.
[0195] inputTensor[0][cy][cx] = ComponentSample[0][cx * 2][cy * 2] inputTensor[1][cy][cx] = ComponentSample[0][cx * 2 + 1][cy * 2] inputTensor[2][cy][cx] = ComponentSample[0][cx * 2][cy * 2 + 1] inputTensor[3][cy][cx] = ComponentSample[0][cx * 2 + 1][cy * 2 + 1] ChromaOffset = 1<<(BitDepthC - 1) inputTensor[4][cy][cx] = ChromaFormatIdc == 0? ChromaOffset : ComponentSample[1][cx][cy] inputTensor[5][cy][cx] = ChromaFormatIdc == 0? ChromaOffset : ComponentSample[2][cx][cy] As described above, there is a moving image decoding apparatus including a predicted image derivation unit that decodes a predicted image and a residual decoding unit that decodes a residual. From parameters that specify the number of channels of the input tensor and the output tensor of the neural network model, by deriving the input tensor or deriving an image from the output tensor, there is an effect that the input tensor can be surely and easily derived.
[0196] The NN filter unit 611 separates a one-channel luminance image into four channels according to the pixel positions and converts it into input data. Also, the NN filter unit 611 may be derived as follows. In the following , in order to absorb the differences of 4:2:0, 4:2:2, and 4:4:4 with a variable indicating color difference subsampling, it can be processed regardless of the difference in color difference samples: inputTensor[0][cy][cx] = ComponentSample[0][cx*2 ][cy*2 ] inputTensor[1][cy][cx] = ComponentSample[0][cx*2+1][cy*2 ] inputTensor[2][cy][cx] = ComponentSample[0][cx*2 ][cy*2+1] inputTensor[3][cy][cx] = ComponentSample[0][cx*2+1][cy*2+1] ChromaOffset = 1<<(BitDepthC-1) inputTensor[4][cy][cx] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[1][cx*2 / SW][cy*2 / SH] inputTensor[5][cy][cx] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[2][cx*2 / SW][cy*2 / SH] Also, since cx = x / 2 and cy = y / 2, the NN filter section 611 may be derived as follows for the above-mentioned ranges of x and y: It may be derived as follows: inputTensor[0][y / 2][x / 2] = ComponentSample[0][x / 2*2 ][y / 2*2 ] inputTensor[1][y / 2][x / 2] = ComponentSample[0][x / 2*2+1][y / 2*2 ] inputTensor[2][y / 2][x / 2] = ComponentSample[0][x / 2*2 ][y / 2*2+1] inputTensor[3][y / 2][x / 2] = ComponentSample[0][x / 2*2+1][y / 2*2+1] ChromaOffset = 1<<(BitDepthC-1) inputTensor[4][y / 2][x / 2] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[1][x / SW][y / SH] inputTensor[5][y / 2][x / 2] = ChromaFormatIdc==0? ChromaOffset : ComponentSample[2][x / SW][y / SH] (Conversion from the NN output data of the post-filter) Based on the value of nnrpf_output_format_idc, the NN filter section 611 derives the output image outSamples from the NN output data outputTensor[][][] of the 3D array, which is the output data of the NN filter . Specifically, as shown in FIG. 20 and below, the NN filter section 611 derives an image as follows based on the value of nnrpf_output_format_idc and the color difference subsampling values outSW and outSH of the output image. Note that for outSW and outSH, values derived based on OutputChromaFormatIdc decoded from the encoded data are used. Hereinafter, x and y represent the coordinates of the luminance pixels of the output image. For example, in outputTensor, the ranges of x and y are x = 0..outLumaWidth-1 and y = outLumaHeight-1, respectively. cx and cy represent the coordinates of the color difference pixels of the output image, and the ranges of cx and cy are cx = 0..outChromaWidth-1 and cy = 0..outChromaHeight-1, respectively. The following NN filter section 611 processes this range. outSamplesL, outSamplesCb, and outSamplesCr represent the luminance channel, color difference (Cb) channel, and color difference (Cr) channel of the output image, respectively.
[0197] When nnrpf_output_format_idc is 0, the NN filter section 611 derives outSamplesL as follows.
[0198] outSamplesL[x][y] = outputTensor[0][y][x] When nnrpf_output_format_idc is 1, outSamplesCb and outSamplesCr are derived as follows.
[0199] outSamplesCb[cx][cy] = outputTensor[0][cy][cx] outSamplesCr[cx][cy] = outputTensor[1][cy][cx] When nnrpf_output_format_idc is 2, outSamplesL, outSamplesCb, and outSamplesC are derived as follows.
[0200] outSamplesL[x][y] = outputTensor[0][y][x] outSamplesCb[x / outSW][y / outSH] = outputTensor[1][y][x] outSamplesCr[x / outSW][y / outSH] = outputTensor[2][y][x] Alternatively, outSamplesCb and outSamplesCr may be derived as follows.
[0201] outSamplesCb[cx][cy] = outputTensor[1][cy*outSH][cx*outSW] outSamplesCr[cx][cy] = outputTensor[2][cy*outSH][cx*outSW] When nnrpf_output_format_idc is 3, outSamplesL is derived as follows: outSamplesL[x / 2*2 ][y / 2*2 ] = outputTensor[0][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2 ] = outputTensor[1][y / 2][x / 2] outSamplesL[x / 2*2 ][y / 2*2+1] = outputTensor[2][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2+1] = outputTensor[3][y / 2][x / 2] Alternatively, outSamplesL may be derived as follows: outSamplesL[cx*2 ][cy*2 ] = outputTensor[0][cy][cx] outSamplesL[cx*2+1][cy*2 ] = outputTensor[1][cy][cx] outSamplesL[cx*2 ][cy*2+1] = outputTensor[2][cy][cx] outSamplesL[cx*2+1][cy*2+1] = outputTensor[3][cy][cx] Furthermore, when nnrpf_output_format_idc is 3 and the output image is in the 4:2:0 format (ChromaFormatIdc of the output image is 1, SW = SH = 2), outSamplesCb and outSamplesCr are derived as follows: outSamplesCb[cx][cy] = outputTensor[4][cy][cx] outSamplesCr[cx][cy] = outputTensor[5][cy][cx] Or, when the output image is in the 4:2:2 format (ChromaFormatIdc of the output image is 2, SW = 2, SH = 1) then outSamplesCb and outSamplesCr are derived as follows: outSamplesCb[cx][cy / 2*2 ] = outputTensor[4][cy][cx] outSamplesCb[cx][cy / 2*2+1] = outputTensor[4][cy][cx] outSamplesCr[cx][cy / 2*2 ] = outputTensor[5][cy][cx] outSamplesCr[cx][cy / 2*2+1] = outputTensor[5][cy][cx] Or, when the output image is in the 4:4:4 format (ChromaFormatIdc of the output image is 3, SW = SH = 1), outSamplesCb and outSamplesCr are derived as follows: outSamplesCb[cx / 2*2 ][cy / 2*2 ] = outputTensor[4][cy][cx] outSamplesCb[cx / 2*2+1][cy / 2*2+1] = outputTensor[4][cy][cx] outSamplesCr[cx / 2*2 ][cy / 2*2 ] = outputTensor[5][cy][cx] outSamplesCr[cx / 2*2+1][cy / 2*2+1] = outputTensor[5][cy][cx] Alternatively, it may be derived as follows: for (j=0; j<outSH; j++) for (i=0; i<outSW; i++) outSamplesCb[cx / outSW*outSW+i][cy / outSH*outSH+j] = outputTensor[4][cy][cx] outSamplesCr[cx / outSW*outSW+i][cy / outSH*outSH+j] = outputTensor[5][cy][cx] Also, in the case of YUV4:2:0 format, since cx = x / 2 and cy = y / 2, the NN filter unit 611 may derive outSamplesL, outSamplesCb, and outSamplesCr as follows: outSamplesL[x / 2*2 ][y / 2*2 ] = outputTensor[0][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2 ] = outputTensor[1][y / 2][x / 2] outSamplesL[x / 2*2 ][y / 2*2+1] = outputTensor[2][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2+1] = outputTensor[3][y / 2][x / 2] outSamplesCb[x / outSW][y / outSH] = outputTensor[4][y / 2][x / 2] outSamplesCr[x / outSW][y / outSH] = outputTensor[5][y / 2][x / 2] Note that outSW and outSH may be set and processed to be the same as the color difference sampling of the input data.
[0202] outSW = SW outSW = SH FIG. 22 is a diagram showing the relationship between the number of channels of a tensor and conversion processing.
[0203] Furthermore, as shown below, the NN filter unit 611 may derive NN input data as follows according to the number of channels NumOutChannels(numTensors) of the output tensor.
[0204] When NumOutChannels is 1, outSamplesL is derived as follows.
[0205] outSamplesL[x][y] = outputTensor[0][y][x] When NumOutChannels is 2, outSamplesCb and outSamplesCr are derived as follows.
[0206] outSamplesCb[cx][cy] = outputTensor[0][cy][cx] outSamplesCr[cx][cy] = outputTensor[1][cy][cx] When NumOutChannels is 3, outSamplesL, outSamplesCb, and outSamplesCr are derived as follows.
[0207] outSamplesL[x][y] = outputTensor[0][y][x] outSamplesCb[x / outSW][y / outSH] = outputTensor[1][y][x] outSamplesCr[x / outSW][y / outSH] = outputTensor[2][y][x] Alternatively, outSamplesCb and outSamplesCr may be derived as follows.
[0208] outSamplesCb[cx][cy] = outputTensor[1][cy*outSH][cx*outSW] outSamplesCr[cx][cy] = outputTensor[2][cy*outSH][cx*outSW] When NumOutChannels is 6, outSamplesL, outSamplesCb, and outSamplesCr are derived as follows.
[0209] outSamplesL[x / 2*2 ][y / 2*2 ] = outputTensor[0][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2 ] = outputTensor[1][y / 2][x / 2] outSamplesL[x / 2*2 ][y / 2*2+1] = outputTensor[2][y / 2][x / 2] outSamplesL[x / 2*2+1][y / 2*2+1] = outputTensor[3][y / 2][x / 2] outSamplesCb[x / outSW][y / outSH] = outputTensor[4][y / 2][x / 2] outSamplesCr[x / outSW][y / outSH] = outputTensor[5][y / 2][x / 2] <Second Embodiment> FIG. 21 is a diagram showing another embodiment of the syntax table of the NN filter SEI. In this embodiment in addition to the syntax element nnrpf_io_idc in Non-Patent Document 3, the syntax element nnrpf_additional_input_idc is used. This has the effect of enabling flexible identification of the conversion method even when the types of input and output increase and the combinations increase. Note that the description of the same parts as in the first embodiment is omitted.
[0210] The header decoding unit 3020 decodes the input / output image information nnrpf_io_idc and the additional input information nnrpf_additional_info_idc when nnrpf_mode_idc is 1 or 2 as shown in FIG. 21 (indicating that the SEI is data of a new post-filter).
[0211] The header decoding unit 3020 derives the formats of the input tensor and output tensor of the model as follows based on the value of nnrpf_io_idc. When nnrpf_io_idc == 0: numInChannels = numOutChannels = 1 When nnrpf_io_idc == 1: numInChannels = numOutChannels = 2 When nnrpf_io_idc == 2: numInChannels = numOutChannels = 3 When nnrpf_io_idc == 3: numInChannels = numOutChannels = 4 When nnrpf_io_idc == 4: numInChannels = numOutChannels = 6 In this embodiment, it is derived such that the values of numInChannels and numOutChannels are equal. However, the association between nnrpf_io_idc and numInChannles and numOutChannels is not limited to this, and other combinations may be used. Also, instead of using nnrpf_io_idc, the number of channels may be directly decoded. Also, conversion processing may be performed using the same variable numInOutChannels, numTensors without distinguishing between numInChannels and numOutChannels.
[0212] Furthermore, the header decoding unit 3020, based on the value of nnrpf_additional_info_idc, for input and output Derive the values of the variables useSliceQPY, usebSY, and usebSC that indicate the use of input channels other than common image components. Each of them is a flag indicating whether SliceQPY, bSY, bSCb, and bSCr are used as additional input channels. In this embodiment, a single flag usebSC controls both bSCb and bSCr. Note that SliceQPY is the luminance quantization parameter in the slice to which a certain coordinate belongs. bSY, bSCb, and bSCR are arrays that store the values of the block strength of the deblocking filter at a certain coordinate in the luminance and chrominance difference (Cb, Cr) channels of the decoded image, respectively. The derivation of the above flags from nnrpf_additional_info_idc is, for example as follows. When nnrpf_additional_info_idc == 0: useSliceQPY = 0, usebSY = 0, usebSC = 0 When nnrpf_additional_info_idc == 1: useSliceQPY = 1, usebSY = 0, usebSC = 0 When nnrpf_additional_info_idc == 2: useSliceQPY = 0, usebSY = 1, usebSC = 0 When nnrpf_additional_info_idc == 3: useSliceQPY = 1, usebSY = 1, usebSC = 0 When nnrpf_additional_info_idc == 4: useSliceQPY = 0, usebSY = 0, usebSC = 1 When nnrpf_additional_info_idc == 5: useSliceQPY = 1, usebSY = 0, usebSC = 1 When nnrpf_additional_info_idc == 6: useSliceQPY = 0, usebSY = 1, usebSC = 1 When nnrpf_additional_info_idc == 7: useSliceQPY = 1, usebSY = 1, usebSC = 1 Note that the association between nnrpf_additional_info_idc and useSliceQPY, usebSY, and usebSC is not limited to this, and other combinations may also be used.
[0213] Alternatively, as follows, each bit of nnrpf_additional_indo_idc can be associated with each flag, and it may be derived by calculation. useSliceQPY = nnrpf_additional_info_idc & 1 usebSY = (nnrpf_additional_info_idc >> 1) & 1 usebSC = (nnrpf_additional_info_idc >> 2) & 1 Alternatively, the header decoding unit 3020 may not encode nnrpf_additional_indo_idc, but instead decode the flags useSliceQPY, usebSY, and usebSC from the encoded data.
[0214] Note that the types of flags are not limited to the above, and flags may be derived in the same way for other information. For example, the slice quantization parameter SliceQPC for color difference, the reference image, the predicted image, and the QP value for each encoding block unit may be used.
[0215] The header decoding unit 3020 further decodes nnrpf_patch_size_minus1, which indicates the size of the processing unit (patch) of the model specified in the SEI (the number of pixels in the horizontal and vertical directions) - 1. At this time, the variable patchSize representing the size of the patch is derived by the following formula. patchSize = nnrpf_patch_size_minus1 + 1 Furthermore, the header decoding unit 3020 decodes nnrpf_overlap. nnrpf_overlap indicates the width of the region adjacent to the patch and input to the model together with the patch. That is, the size of the height x width of the tensor input to the model is (nnrpf_overlap*2+patchSize) x (nnrpf_overlap*2+patchSize) pixels.
[0216] (Another example of the details of the NN filter unit 611) The NN filter unit 611 in this embodiment performs three steps: STEP1: conversion from the image ComponentSample to the input tensor inputTensor, STEP2: applying the neural network filter processing postProcessingFilter to the inputTensor to output the outputTensor, and STEP3: deriving an image from the derived outputTensor.
[0217] SliceQPY is the quantization parameter qP of the luminance dose in the slice to which a certain coordinate belongs. bSY, bSCb, and bsCR are arrays that store the values of the block strength (bS) of the deblocking filter at a certain coordinate in the luminance and chrominance difference (Cb, Cr) channels of the decoded image, respectively. bS may be a value derived from the prediction mode of the block, the motion vector, the difference value of the pixel value, etc. in the deblocking filter processing.
[0218] (Pseudo code) The NN filter unit 611 may perform the processing of STEP1, STEP2, and STEP3 in units of a predetermined patch size from the upper left coordinate of the image. The following pseudo code is an overview of the processing in units of patches. for(cTop=0; cTop<PicHeightInLumaSamples; cTop+=patchSize*rH) { for (cLeft = 0; cLeft < PicWidthInLumaSamples; cLeft += patchSize * rW) { <STEP1: Conversion to Input Tensor of Image and Additional Information> <STEP2: Application of Filter Processing> <STEP3: Conversion of Output Tensor to Image> } } Here, rH and rW are variables that determine the increment amounts of the loop variables related to chrominance subsampling. Also, rH and rW are the vertical and horizontal magnification factors for deriving the size of the tensor patch on the image.
[0219] The NN filter unit 611 may be derived as follows using chrominance subsampling. rH = SH rW = SW Alternatively, the NN filter unit 611 may be derived as follows using the number of input channels of the tensor as well. rH = numInChannles == 6? 2 : 1 rW = (numInChannles == 4 || numInChannels == 6)? 2 : 1 Alternatively, the NN filter unit 611 may use the chrominance subsampling inside the tensor and the chrominance subsampling of the input image to derive rW and rH, for example, by the following formula. rH = numInChannels == 2? SH : numInChannles == 6? 2 : 1 rW = numInChannels == 2? SW : (numInChannles == 4 || numInChannels == 6)? 2 : 1 rH and rW indicate that a tensor of patch size (height x width) patchSize x patchSize corresponds to an area of patchSize * rW x patchSize * rH on the input luminance image (width x height). Hereinafter, processing examples from STEP1 to STEP3 will be described.
[0220] (STEP1) The NN filter section 611 derives a patch including an overlap. If the overlap part is included, the size of the input data is patchSize+nnrpf_overlap*2 x patchSize+nnrpf_overlap*2 pixels in both vertical and horizontal directions. The NN filter section 611 derives each element of one patch including the overlap in the input tensor as follows in the following pseudo-code.
[0221] for (yP = -nnrpf_overlap; yP < patchSize+nnprf_overlap; yP++) { for (xP = -nnrpf_overlap; xP < patchSize+nnprf_overlap; xP++) { yP1 = yP+nnrpf_overlap xP1 = xP+nnrpf_overlap yT = Clip3(0, PicHeightInLumaSamples-1, cTop+yP*rH) yL = Clip3(0, PicWidthInLumaSamples-1, cLeft+xP*rW) yB = Clip3(0, PicHeightInLumaSamples-1, yT+rH-1) xR = Clip3(0, PicWidthInLumaSamples-1, yL+rW-1) cy = yT / SH cx = yL / SW <Derivation of input tensor: Conversion of input image> <Derivation of input tensor: Conversion of additional information> } } yP and xP are loop variables in the height direction (H) and width direction (W) of the tensor. yP1 and xP1 are indices offset by the overlap size to reference elements of the input tensor including the overlap region. yT, yL, yB, xR are the coordinate values on the luminance component corresponding to yP and xP in the patch to be processed. cy and cx are the coordinate values on the chrominance component of the input image corresponding to yP and xP in the patch to be processed. These variables may be derived as needed during each process described below. Next, the conversion process of the input image will be described. Next, the conversion process of the input image will be described.
[0222] (Conversion of Input Image) Within the loop of STEP1, the NN filter unit 611 converts the input image according to the value of numInChannels as shown in the following pseudo-code, and derives the values of each channel corresponding to the tensor coordinates yP, xP: Within the loop of STEP1, the NN filter unit 611 converts the input image according to the value of numInChannels as shown in the following pseudo-code, and derives the values of each channel corresponding to the tensor coordinates yP, xP: as follows: if (numInChannels == 1) { inputTensor[0][yP1][xP1] = ComponentSamples[0][yT][yL] ch_pos = 1 } else if (numInChannels == 2) { inputTensor[0][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[1][cx][cy] inputTensor[1][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[2][cx][cy] ch_pos = 2 } else if (numInChannels == 3) { inputTensor[0][yP1][xP1] = ComponentSamples[0][yT][yL] inputTensor[1][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[1][cx][cy] inputTensor[2][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[2][cx][cy] ch_pos = 3 } else if (numInChannels == 4) { inputTensor[0][yP1][xP1] = ComponentSamples[0][yL][yT] inputTensor[1][yP1][xP1] = ComponentSamples[0][yL][yB] inputTensor[2][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[1][cx][cy] inputTensor[3][yP1][xP1] = ChromaFormatIdc == 0? ChromaOffset : ComponentSamples[2][cx][cy] ch_pos = 4 } else if (numInChannels == 6) { inputTensor[0][yP1][xP1] = ComponentSamples[0][yL][yT] inputTensor[1][yP1][xP1] = ComponentSamples[0][xR][yT] inputTensor[2][yP1][xP1] = ComponentSamples[0][yL][yB] inputTensor[3][yP1][xP1] = ComponentSamples[0][xR][yB] inputTensor[4][yP1][xP1] = ChromaFormatIdc == 0 ? ChromaOffset : ComponentSamples[1][cx][cy] inputTensor[5][yP1][xP1] = ChromaFormatIdc == 0 ? ChromaOffset : ComponentSamples[2][cx][cy] ch_pos = 6 } In the above, ch_pos may be derived as ch_pos = numInChannels.
[0223] When numInChannels == 1, the image data of one channel of the input tensor is derived from the luminance component of the decoded image.
[0224] When numInChannels == 2, the image data of two channels of the input tensor is derived from the two chrominance components of the decoded image.
[0225] When numInChannels == 3, the 4:4:4 format image data represented by three channels is derived from the luminance component and chrominance components of the decoded image and set to the input tensor.
[0226] When numInChannels == 4, the 4:2:2 format image data represented by a two-channel luminance image and a two-channel chrominance image is derived from the luminance component and chrominance components of the decoded image and set to the input tensor.
[0227] When numInChannels == 6, 4:2:0 format image data represented by a 4-channel luminance image and a 2-channel chrominance image is derived from the luminance component and chrominance components of the decoded image and set to the input tensor. Derive and set it to the input tensor.
[0228] In any case, when the decoded image has no chrominance component (ChromaFormatIdc == 0), the NN filter unit 611 sets ChromaOffset to the chrominance channel. Also, ch_pos is an index indicating the next available channel in the input tensor. When adding data to the input tensor, it is advisable to use the channels after the value of ch_pos.
[0229] (Derivation of additional information) Next, the NN filter unit 611 adds additional information to the inputTensor based on the additional input information (flag). For example, if useSliceQPY is true, the value of SliceQPY is set to the next channel of the inputTensor (inputTensor[ch_pos]). Note that when setting a value to a channel, increase the value of ch_pos according to the number of channels used. Next, if usebSY is true, the value of bSY at the position corresponding to the patch region is set to the next channel of the inputTensor. Further, if usebSC is true and the decoded image has chrominance components, the values of bSCb and bSCr at the positions corresponding to the patch region are set to the next channel of the inputTensor. At this time, the NN filter unit 611 may set the values of SliceQPY, bSY, bSCb, and bSCr as they are, or may set the values after conversion. The following shows an example of deriving the channels of the additional information: if (useSliceQPY) { inputTensor[ch_pos][yP1][xP1] = 2^((SliceQPY - 42) / 6) ch_pos += 1 } if (usebSY) { inputTensor[ch_pos][yP1][xP1] = Clip3(0,(1<<BitDepthY)-1, bSY[yL][yT]< <(BitDepthY-2)) ch_pos += 1 } if (usebSC && InputChromaFormatIdc!=0) { inputTensor[ch_pos][yP1][xP1] = Clip3(0, (1<<BitDepthC)-1, bSCb[cx][cy] << (BitDepthC-2)) ch_pos += 1 inputTensor[ch_pos][yP1][xP1] = Clip3(0, (1<<BitDepthC)-1, bSCr[cx][cy] << (BitDepthC-2)) ch_pos += 1 } As additional information other than the above, an example is shown where the NN filter unit 611 adds a second image different from the input image to the input tensor. The following example is the process when adding the channels of the reference image refPicY using the flag useRefPicY : if (useRefPicY) { inputTensor[ch_pos][yP1][xP1] = refPicY[yL][yT] ch_pos += 1 } The flag useRefPicY may be derived from nnrpf_additional_info_idc or decoded from the encoded data, similar to the flags of other additional information such as useSliceQPY. Also, in addition to the reference image, the same applies when adding the channels of the predicted image predPicY for one frame : if (usePredPicY) { inputTensor[ch_pos][yP1][xP1] = predPicY[yL][yT] ch_pos += 1 } Note that the method of providing additional information to the post-filter is not limited to the above. It may be derived as another parameter separate from inputTensor and input to the post-filter processing.
[0230] (STEP2) In STEP2, the NN filter unit 611 performs post-filter processing using the derived inputTensor as input, and obtains an output tensor outputTensor that is the result of the filter processing: outputTensor = postProcessingFilter(inputTensor) In this embodiment, the size of the output tensor outputTensor in the height direction (H) × width direction (W) is patchSize x patchSize, and it is assumed that no overlap region is included. However, this is not limited thereto, and an output tensor including an overlap region may be cropped to a size of patchSize x patchSize and used.
[0231] (STEP3) In STEP3, the NN filter unit 611 derives image data from the output tensor outputTensor. The NN filter unit 611 derives the pixel values of each component outSamplesL, outSamplesCb, and outSamplesCr of the output image from the data of the output tensor outputTensor. At this time, as shown in the following pseudo-code, the derivation of the image data is performed based on the number of channels (data format) of the output tensor and the color difference subsampling of the output image. The following pseudo-code is an example of the processing in STEP3 : for(yP=0, ySrc=cTop; yP<patchSize; yP++, ySrc+=rH) for(xP = 0, xSrc = cLeft; xP < patchSize; xP++, xSrc += rW) { if (numOutChannels == 1) { / / rW = 1, rH = 1 if (pfp_component_idc != 1) { outSamplesL[xSrc][ySrc] = outputTensor[0][yP][xP] } else { outSamplesL[xSrc][ySrc] = ComponentSamples[0][xSrc][ySrc] } if (OutputChromaFormatIdc != 0 && pfp_component_idc != 0) { cyOut = ySrc / outSH cxOut = xSrc / outSW outSamplesCb[cxOut][cyOut] = ComponentSamples[1][ySrc / SH][xSrc / SW] outSamplesCr[cxOut][cyOut] = ComponentSamples[2][ySrc / SH][xSrc / SW] } } else if (numOutChannels == 2) { / / rW = SW,rH = SH for (dy = 0; dy < rH, dy++) { for (dx = 0; dx < rW, dx++) { outSamplesL[ySrc + dy][xSrc + dx] = ComponentSamples[0][ySrc + dy][xSrc + dx] if (OutputChromaFormatIdc != 0) { cyOut = (ySrc + dy) / outSH cxOut = (xSrc + dx) / outSW if (pfp_component_idc!=0) { outSamplesCb[cxOut][cyOut] = outputTensor[0][yP][xP] outSamplesCr[cxOut][cyOut] = outputTensor[1][yP][xP] } else { outSamplesCb[cxOut][cyOut] = ComponentSamples[1][(ySrc+dy) / SH] [ (xSrc+dx) / SW ] outSamplesCr[cxOut][cyOut] = ComponentSamples[2][(ySrc+dy) / SH] [ (xSrc+dx) / SW ] } } } } } else if (numOutChannels==3) { / / rW=1, rH=1 if (pfp_component_idc!=1) { outSamplesL[xSrc][ySrc] = outputTensor[0][yP][xP] } else { outSamplesL[xSrc][ySrc] = ComponentSamples[0][xSrc][ySrc] } if (OutputChromaFormatIdc!=0) { cyOut = ySrc / outSH cxOut = xSrc / outSW if (pfp_component_idc!=0) { outSamplesCb[cxOut][cyOut] = outputTensor[1][yP][xP] outSamplesCr[cxOut][cyOut] = outputTensor[2][yP][xP]} else { outSamplesCb[cxOut][cyOut] = ComponentSamples[1][xSrc / SW][ySrc / SH] outSamplesCr[cxOut][cyOut] = ComponentSamples[2][xSrc / SW][ySrc / SH] } } } else if (numOutChannels==4) { / / rW=2, rH=1 if (pfp_component_idc!=1) { outSamplesL[xSrc ][ ySrc ] = outputTensor[0][yP][xP] outSamplesL[xSrc+1][ ySrc ] = outputTensor[1][yP][xP] } else { outSamplesL[xSrc ][ ySrc ] = ComponentSamples[0][xSrc ][ ySrc ] outSamplesL[xSrc+1][ ySrc ] = ComponentSamples[0][xSrc+1][ ySrc ] } if (OutputChromaFormatIdc!=0) { for (dx=0; dx<rW; dx++) { cyOut = (ySrc+dy) / outSH cxOut = (xSrc+dx) / outSW if (pfp_component_idc!=0) { outSamplesCb[cxOut][cyOut] = outputTensor[2][yP][xP] outSamplesCr[cxOut][cyOut] = outputTensor[3][yP][xP] } else { outSamplesCb[cxOut][cyOut] = ComponentSamples[1][(xSrc+dx) / SW] (ySrc+dy) / SH] outSamplesCr[cxOut][cyOut] = ComponentSamples[2][(xSrc+dx) / SW] (ySrc+dy) / SH] } } } } else if (numOutChannels==6) { / / rW=2, rH=2 if (pfp_component_idc!=1) { outSamplesL[xSrc ][ySrc ] = outputTensor[0][yP][xP] outSamplesL[xSrc+1][ySrc ] = outputTensor[1][yP][xP] outSamplesL[xSrc ][ySrc+1] = outputTensor[2][yP][xP] outSamplesL[xSrc+1][ySrc+1] = outputTensor[3][yP][xP] } else { outSamplesL[xSrc ][ySrc ] = ComponentSamples[0][xSrc ][ySrc ] outSamplesL[xSrc+1][ySrc ] = ComponentSamples[0][xSrc+1][ySrc ] outSamplesL[xSrc ][ySrc+1] = ComponentSamples[0][xSrc ][ySrc+1] outSamplesL[xSrc+1][ySrc+1] = ComponentSamples[0][xSrc+1][ySrc+1] } if (OutputChromaFormatIdc!=0) { for (dy = 0; dy < rH; dy++) { for (dx = 0; dx < rW; dx++) { cyOut = (ySrc + dy) / outSH cxOut = (xSrc + dx) / outSW if (pfp_component_idc != 0) { outSamplesCb[cxOut][cyOut] = outputTensor[4][yP][xP] outSamplesCr[cxOut][cyOut] = outputTensor[5][yP][xP] } else { outSamplesCb[cxOut][cyOut] = ComponentSamples[1][(xSrc + dx) / SW] [(ySrc + dy) / SH] outSamplesCr[cxOut][cyOut] = ComponentSamples[2][(xSrc + dx) / SW] [(ySrc + dy) / SH] } } } } } } } When numOutChannels is 1, the NN filter unit 611 outputs using one channel of outputTensor to derive the luminance component of the image. When the output image includes color difference components, the NN filter unit 611 copies and uses the color difference components of the input image.
[0232] When numOutChannels is 2, the NN filter unit 611 outputs using two channels of outputTensor to derive two color difference components of the image. Also, the NN filter unit 611 uses the luminance component of the input image as the luminance component of the output image.
[0233] When numOutChannels is 3, the NN filter unit 611 derives the outputTensor into the luminance component and two chrominance components of the output image.
[0234] When numOutChannels is 4, the NN filter unit 611 uses two channels of the outputTensor to derive the luminance component of the output image, and uses the other two channels to derive two chrominance components.
[0235] When numOutChannels is 6, the NN filter unit 611 uses four channels of the outputTensor to derive the luminance component of the output image, and uses the other two channels to derive two chrominance components.
[0236] In any case of numOutChannels, if the output image does not have chrominance components, the chrominance channels of the output tensor are not processed. Also, when numOutChannels is 2 or more and includes chrominance channels, the values outSH and outSW derived from the chrominance subsampling of the output image are used to perform a conversion to conform to the chrominance component format of the output image. When numOutChannels is 2 the values SH(SubHightC) and SW(SubWidthC) derived from the chrominance subsampling of the input image are also referred to and a conversion is performed to conform to the chrominance component format of the output image. Furthermore, when numOutChannels is other than 2 and pfp_component_idc is 1, the luminance component of the output image is not updated by the output tensor and the luminance component of the input image is used. Similarly, when numOutChannels is other than 1 and pfp_component_idc is 0, the chrominance component of the output image is not updated by the output tensor and the chrominance component of the input image is used.
[0237] Moreover, when numOutChannels is other than 2 and pfp_component_idc is 1, the luminance component of the output image is not updated by the output tensor and the luminance component of the input image is used. Similarly, when numOutChannels is other than 1 and pfp_component_idc is 0, the chrominance component of the output image is not updated by the output tensor and the chrominance component of the input image is used.
[0238] In this way, the NN filter unit 611 determines the number of channels of the output tensor and the chrominance sub-sub-sub-sampling of the output image. Based on the sampling, the pixel value of each component of the output image is derived. 611 derives rW*rH pixels on the output image corresponding to one pixel of the output tensor represented by yP and xP. This can be done in a loop (dy, dx) as in the pseudocode above when numOutChannels is 2 or 6, or the loop can be omitted or expanded as in the other cases. may be set to always loop twice for both dy and dx.
[0239] In the above example, the NN filter unit 611 performs multiple inductions for the same color difference coordinates of the output image. However, if the coordinates overlap, the second and subsequent derivations may be omitted. do not have.
[0240] Furthermore, in the above example, cTop and cLeft are configured to increment by patchSize*rH and patchSize*rW as coordinate values on the image, respectively, but this is not limited to this. It may be configured as a loop that increments by patchSize*SH and patchSize*SW using subsampling, or it may be a loop that always increments by 1 or 2. Similarly, ySrc and xSrc are configured to increment by rH and rW, respectively, as coordinate values on the image, but this is not limited to this. It may be configured as a loop that increments by SH and SW using chrominance subsampling of the input image, or a loop that increments by outSH and outSW using chrominance subsampling of the output image, or it may be a loop that always increments by 1.
[0241] A moving image decoding apparatus includes a predicted image derivation unit that decodes a predicted image as described above, and a residual decoding unit that decodes a residual. The apparatus decodes input / output image information that specifies the number of channels of the input / output tensors of a neural network model, and additional input information. Using the input / output image information, a part of the input tensor is derived from a first image. Further, using the additional input information, another part of the input tensor is derived from a second image different from the first image, or from encoding information related to the derivation of the predicted image or the decoding of the residual. This enables the properties and features of the input image to be transmitted to the neural network model in more detail, enhancing the filtering effect.
[0242] Furthermore, as described above, the addition value of the loop variable is changed using the parameters of chroma subsampling, and loop processing is performed in a raster scan manner. Within the loop, an input tensor is derived from an image, and a deep learning filter is applied to the input tensor to derive an output tensor. This enables the input tensor to be derived with the same processing even in the case of different color samplings.
[0243] (Summary) This application may also be configured to decode encoded data including input tensor identification parameters that specify the correspondence between the channels of the input tensor of a neural network model and color components.
[0244] This application may also be configured such that a relational expression for deriving an input tensor from an input image is defined according to the input tensor identification parameters and the chroma subsampling of the input image.
[0245] The input tensor identification parameters of this application may also be configured to include specifying any one of 1 channel, 2 channels, 3 channels, and 6 channels.
[0246] This application may also be configured to include means for deriving an input tensor from an input image according to the input tensor identification parameters.
[0247] This application may be configured to decode encoded data including output tensor identification parameters that specify the correspondence between channels and color components of the output tensor of a neural network model. It may also be.
[0248] This application may also be configured such that a relational expression for deriving an output image from an output tensor is defined according to the output tensor identification parameter and the color difference subsampling of the output image.
[0249] The output tensor identification parameter of this application may be configured to include specifying any one of 1 channel, 2 channels, 3 channels, and 6 channels.
[0250] This application may also be configured to include means for deriving an output image from an output tensor according to the output tensor identification parameter.
[0251] Thus, this SEI includes information on the input data format to the NN filter and the output data format of the NN filter. As a result, it is possible to easily select a method for appropriately converting a decoded image into input data of the NN filter and a method for appropriately converting output data of the NN filter into an output image without loading and analyzing the model.
[0252] (Configuration for deriving input tensor identification parameter and output tensor parameter) Note that the NN filter unit 611 does not decode the input tensor identification parameter and the output tensor identification parameter from the additional data, but may decode them from the encoded data or derive them from the topology of the NN model identified by a URI or the like.
[0253] The NN filter unit 611 derives nnrpf_input_format_idc as follows according to the number of channels NumInChannels of the input data inputTensor of the NN model. In the case of 1 channel, nnrpf_input_format_idc = 0 In the case of 2 channels, nnrpf_input_format_idc = 1 In the case of 3 channels, nnrpf_input_format_idc = 2 In the case of 6 channels, nnrpf_input_format_idc = 3 The NN filter unit 611 derives nnrpf_output_format_idc as follows according to the number of channels NumOutChannels of the output data outputTensor of the NN model as follows In the case of 1 channel, nnrpf_output_format_idc = 0 In the case of 2 channels, nnrpf_output_format_idc = 1 In the case of 3 channels, nnrpf_output_format_idc = 2 In the case of 6 channels, nnrpf_output_format_idc = 3 According to the above configuration, the NN filter unit 611 analyzes the number of dimensions of the input data and output data of the encoded data transmitted or the specified NN model, and performs conversion from the input image to the input tensor and conversion from the output tensor to the output image according to the analysis results (input tensor identification parameters, output tensor identification parameters). As a result, it is possible to specify the relationship between the color components and channels not specified in the NN model itself, prepare the NN input data, and obtain the output image from the NN output data. Note that this SEI may include information on the complexity of the neural network work model
[0254] (Configuration of the image encoding device) Next, the configuration of the image encoding device 11 according to the present embodiment will be described. FIG. 7 is a block diagram showing the configuration of the image encoding device 11 according to the present embodiment The image encoding device 11 includes a predicted image generation unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, and a loop - Prediction filter 107, prediction parameter memory (prediction parameter storage unit, frame memory) 108, reference picture memory (reference image storage unit, frame memory) 109, encoding parameter determination Unit 110, parameter encoding unit 111, prediction parameter derivation unit 120, and entropy encoding unit 104.
[0255] The prediction image generation unit 101 generates a prediction image for each CU. The prediction image generation unit 101 includes the inter-prediction image generation unit 309 and the intra-prediction image generation unit 310 that have been described above, and the description thereof is omitted.
[0256] The subtraction unit 102 subtracts the pixel value of the prediction image of the block input from the prediction image generation unit 101 from the pixel value of the image T to generate a prediction error. The subtraction unit 102 outputs the prediction error to the conversion / quantization unit 103.
[0257] The conversion / quantization unit 103 calculates conversion coefficients for the prediction error input from the subtraction unit 102 by frequency conversion, and derives quantized conversion coefficients by quantization. The conversion / quantization unit 103 Outputs the quantized conversion coefficients to the parameter encoding unit 111 and the inverse quantization / inverse conversion unit 105.
[0258] The inverse quantization / inverse conversion unit 105 is the same as the inverse quantization / inverse conversion unit 311 (Fig. 5) in the image decoder 31 and the description thereof is omitted. The calculated prediction error is output to the addition unit 106.
[0259] The parameter encoding unit 111 includes a header encoding unit 1110, a CT information encoding unit 1111, and a CU encoding unit 1112 (prediction mode encoding unit). The CU encoding unit 1112 further includes a TU encoding unit 1114 and the following describes the general operations of each module.
[0260] The header encoding unit 1110 performs encoding processing on parameters such as header information, segmentation information, prediction information, and quantized conversion coefficients.
[0261] The CT information encoding unit 1111 encodes QT, MT (BT, TT) split information, etc.
[0262] The CU encoding unit 1112 encodes CU information, prediction information, split information, etc.
[0263] When the TU contains prediction errors, the TU encoding unit 1114 encodes QP update information and quantized prediction errors.
[0264] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter-prediction parameters (predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX), intra-prediction parameters (intra_luma_mpm_flag, intra_luma_mpm_idx, intra_luma_mpm_reminder, intra_chroma_pred_mode), and quantized transform coefficients to the parameter encoding unit 111. supply.
[0265] The entropy encoding unit 104 receives the quantized transform coefficients and encoding parameters (split information, prediction parameters) from the parameter encoding unit 111. The entropy encoding unit 104 entropy-encodes these to generate and output the encoded data Te. encode them to generate and output the encoded data Te.
[0266] The prediction parameter derivation unit 120 is a means including an inter-prediction parameter encoding unit 112 and an intra-prediction parameter encoding unit 113, and derives intra-prediction parameters and intra-prediction parameters from the parameters input from the encoding parameter determination unit 110. The derived intra-prediction parameters and intra-prediction parameters are output to the parameter encoding unit 111. are output.
[0267] (Configuration of the inter-prediction parameter encoding unit) As shown in FIG. 8, the inter-prediction parameter encoding unit 112 includes a parameter encoding control unit 1121 and an inter-prediction parameter derivation unit 303. Inter-prediction parameter derivation The unit 303 has a configuration common to the image decoding apparatus. The parameter encoding control unit 1121 includes a merge index derivation unit 11211 and a vector candidate index derivation unit 11212.
[0268] The merge index derivation unit 11211 derives merge candidates and the like, and outputs them to the inter-prediction parameter derivation unit 303. The vector candidate index derivation unit 11212 derives prediction vector candidates and the like, and outputs them to the inter-prediction parameter derivation unit 303 and the parameter encoding unit 111.
[0269] (Configuration of the intra-prediction parameter encoding unit 113) The intra-prediction parameter encoding unit 113 includes a parameter encoding control unit 1131 and an intra-prediction parameter derivation unit 304. The intra-prediction parameter derivation unit 304 has a configuration common to the image decoding apparatus.
[0270] The parameter encoding control unit 1131 derives IntraPredModeY and IntraPredModeC. Further, it determines intra_luma_mpm_flag by referring to mpmCandList[]. These prediction parameters are output to the intra-prediction parameter derivation unit 304 and the parameter encoding unit 111.
[0271] However, different from the image decoding apparatus, the inputs to the inter-prediction parameter derivation unit 303 and the intra-prediction parameter derivation unit 304 are in the encoding parameter determination unit 110 and the prediction parameter memory 108, and are output to the parameter encoding unit 111.
[0272] The addition unit 106 adds the pixel values of the predicted block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization / inverse transformation unit 105 for each pixel to generate a decoded image. The addition unit 106 stores the generated decoded image in the reference picture memory 109.
[0273] The loop filter 107 applies a deblocking filter, SAO, and ALF to the decoded image generated by the addition unit 106. Note that the loop filter 107 does not necessarily include the above three types of filters. For example, it may have a configuration including only the deblocking filter.
[0274] The prediction parameter memory 108 stores the prediction parameters generated by the encoding parameter determination unit 110 at predetermined positions for each target picture and CU.
[0275] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at predetermined positions for each target picture and CU.
[0276] The encoding parameter determination unit 110 selects one set from among a plurality of sets of encoding parameters. The encoding parameters are the QT, BT, or TT splitting information, prediction parameters, or parameters to be encoded generated in relation to these, as described above. The predicted image generation unit 101 generates a predicted image using these encoding parameters.
[0277] The encoding parameter determination unit 110 calculates the amount of information and the RD cost value indicating the encoding error for each of the plurality of sets. The RD cost value is, for example, the sum of the amount of code and the value obtained by multiplying the mean squared error by a coefficient λ. The amount of code is the amount of information of the encoded data Te obtained by entropy encoding the quantization error and the encoding parameters. The mean squared error is calculated in the subtraction unit 102. is the sum of the squares of the prediction errors. The coefficient λ is a real number greater than a preset zero. The encoding parameter determination unit 110 selects a set of encoding parameters that minimizes the calculated cost value. The encoding parameter determination unit 110 outputs the determined encoding parameters to the parameter encoding unit 111 and the prediction parameter derivation unit 120.
[0278] Note that a part of the image encoding device 11 and the image decoding device 31 in the above-described embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generation unit 308, the inverse quantization / inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the predicted image generation unit 101, the subtraction unit 102, the transform / quantization unit 103, the entropy encoding unit 104, the inverse quantization / inverse transform unit 105, the loop filter 107, the encoding parameter determination unit 110, the parameter encoding unit 111, the prediction parameter derivation unit 120 may be realized by a computer. In that case, a program for realizing this control function is recorded on a computer-readable recording medium, and the program recorded on this recording medium is read into a computer system and executed to realize it. Here, the "computer system" refers to a computer system built in either the image encoding device 11 or the image decoding device 31, and includes hardware such as an OS and peripheral devices. Also, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., a computer system This refers to a storage device such as a hard disk incorporated in the stem. Further, the "computer-readable recording medium" may include those that hold a program dynamically for a short time, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and those that hold a program for a certain period of time, like the volatile memory inside a computer system serving as a server or client in that case. Also, the above program may be for realizing a part of the aforementioned functions, and may further be realizable in combination with a program already recorded in the computer system for realizing the aforementioned functions.
[0279] Also, part or all of the image encoding device 11 and the image decoding device 31 in the above-described embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be made into a processor individually, or part or all of them may be integrated and made into a processor. Also, the method of integrating into a circuit is not limited to LSI and may be realized by a dedicated circuit or a general-purpose processor. Also, when a technology for integrating into a circuit that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit using that technology may be used.
[0280] As described above, one embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.
[0281] (NNR) Neural Network Coding and Representation (NNR) is an international standard for efficiently compressing a neural network (NN). By compressing a trained NN, it becomes possible to improve the efficiency when storing and transmitting the NN. The outline of the encoding / decoding process of NNR will be described below.
[0282] The outline of the encoding / decoding process of NNR will be described below.
[0283] FIG. 15 is a diagram showing an NNR encoding device and decoding device.
[0284] The NN encoding device 801 includes a preprocessing unit 8011, a quantization unit 8012, and an entropy encoding unit 8013. The NN encoding device 801 inputs the NN model O before compression, performs quantization of the NN model O in the quantization unit 8012, and obtains a quantized model Q. Before quantization, the NN encoding device 801 may repeatedly apply parameter reduction methods such as pruning and sparsification in the preprocessing unit 8011. Then, the entropy encoding unit 8013 applies entropy encoding to the quantized model Q to obtain a bitstream S for storing and transmitting the NN model.
[0285] The NN decoding device 802 includes an entropy decoding unit 8021, a parameter restoration unit 8022, and a postprocessing unit 8023. The NN decoding device 802 first inputs the transmitted bitstream S, performs entropy decoding of S in the entropy decoding unit 8021, and obtains an intermediate model RQ. If the operation environment of the NN model supports inference using the quantization representation used by RQ, RQ may be output and used for inference. Otherwise, the parameter restoration unit 8022 restores the parameters of RQ to the original representation to obtain an intermediate model RP. If the sparse tensor representation to be used can be processed in the operation environment of the NN model, RP may be output and used for inference. Otherwise, a reconstructed NN model R that does not include a tensor or structural representation different from the NN model O is obtained and output.
[0286] The NNR standard has decoding methods for numerical representations of specific NN parameters such as integers and floating-point numbers.
[0287] The decoding method NNR_PT_INT decodes a model consisting of integer-valued parameters. The decoding method NNR_PT_FLOAT extends NNR_PT_INT by adding a quantization step size delta. Multiply this delta by the above integer value to generate a scaled integer. delta is derived from the integer quantization parameter qp and the granularity parameter qp_density of delta as follows. It is derived from the meter qp and the granularity parameter qp_density of delta as follows.
[0288] mul = 2^(qp_density) + (qp & (2^(qp_density)-1)) delta = mul * 2^((qp >> qp_density)-qp_density) (Format of the trained NN) The representation of the trained NN consists of two elements: a topology representation such as the size of the layers and the connections between the layers, and a parameter representation such as weights and biases.
[0289] The topology representation is covered in native formats such as TensorFlow and PyTorch. However, for improved interoperability, there are exchange formats such as the Open Neural Network Exchange Format (ONNX) and the Neural Network Exchange Format (NNEF). Although it is covered in native formats such as TensorFlow and PyTorch, for improved interoperability, there are exchange formats such as the Open Neural Network Exchange Format (ONNX) and the Neural Network Exchange Format (NNEF).
[0290] Also, in the NNR standard, the topology information nnr_topology_unit_payload is transmitted as part of the NNR bitstream containing the compressed parameter tensor. This enables interoperability not only with exchange formats but also with topology information represented in native formats. This enables interoperability not only with exchange formats but also with topology information represented in native formats.
[0291] (SEI for post-filtering purposes) The SEI for post-filtering purposes indicates the purpose of the post-filtering process and is for the post-filtering process. Describe the input / output information according to the purpose. FIG. 13 shows an example of the syntax of this SEI. It is.
[0292] First, as an input of this SEI, define the InputChromaFormatIdc (ChromaFormatIdc) of the input image. The value of sps_chroma_format_idc of the encoded data is assigned to this value. The value of pfp_id indicates the identification number of the post-filtering process specified by other mechanisms. In this embodiment, it is associated with the NN filter SEI. This SEI message is applied in output order to the current decoded image and all subsequent decoded images until a new CLVS (Coded Layer Video Sequence) starts or the bitstream ends in the current layer.
[0293] The value of pfp_id includes an identification number used to identify the post-filtering process. The identification number is assumed to take values from 0 to 2^20 - 1, and values from 2^20 to 2^21 - 1 are reserved for future use.
[0294] The value of pfp_purpose indicates the purpose of the post-filtering process identified by pfp_id. The value of pfp_purpose ranges from 0 to 2^32 - 2. Other values of pfp_purpose are reserved for future specification. Note that the decoder for additional information ignores the post_filter_purpose SEI message containing reserved values of pfp_purpose. It is.
[0295] A value of 0 for pfp_purpose indicates an improvement in visual quality. That is, it means that a post-filtering process that performs image restoration processing without image resolution conversion is applied.
[0296]
[0297] The value 1 of pfp_purpose specifies the width or height of the trimmed output image. That is, it means that post-filtering with conversion of the image resolution is applied.
[0298] When the value of pfp_purpose is 1, there are syntax elements pfp_pic_width_in_luma_samples and pfp_pic_height_in_luma_samples.
[0299] pfp_pic_width_in_luma_samples specifies the width of the luminance pixel array of the image obtained by applying the post-processing filter identified by pfp_id to the trimmed output image.
[0300] pfp_pic_height_in_luma_samples specifies the height of the luminance pixel array of the image obtained by applying the post-processing filter identified by pfp_id to the trimmed output image.
[0301] In the examples of Non-Patent Document 1 and Non-Patent Document 2, information on resolution conversion and inverse conversion associated with color difference format conversion could not be described well. In the present embodiment, the above problems are solved by clarifying input / output information.
[0302] When the value of pfp_purpose is 2, there are syntax elements indicating the color component to which post-filtering is applied and the information on the color difference format of the output image. That is, it means that post-filtering related to color difference format conversion is applied.
[0303] When the value of pfp_purpose is 2, there are syntax elements pfp_component_idc and pfp_output_diff_chroma_format_idc.
[0304] The pfp_component_idc specifies the color component to which the post-filtering process is applied.
[0305] A value of 0 for pfp_component_idc indicates that the post-filtering process is applied only to the luminance component.
[0306] A value of 1 for pfp_component_idc indicates that the post-filtering process is applied to the two chrominance components. to be applied.
[0307] A value of 2 for pfp_component_idc indicates that the post-filtering process is applied to all three color components. to be applied.
[0308] If the syntax element pfp_component_idc does not exist, the value of pfp_component_idc is assumed to be 2. to be estimated.
[0309] The pfp_output_diff_chroma_format_idc is filter update information, indicating the difference value between the identification value of the chrominance format output by the post-filtering process and the identification value of the input chrominance format. Note that the value of pfp_output_diff_chroma_format_idc must be in the range from 0 to 2. And the variable OutputChromaFormatIdc, which is the identification value of the chrominance format output by the post-filtering process, is derived as follows. not. And the variable OutputChromaFormatIdc, which is the identification value of the chrominance format output by the post-filtering process, is derived as follows. OutputChromaFormatIdc = InputChromaFormatIdc + pfp_output_diff_chroma_format_idc
[0310] OutputChromaFormatIdc = InputChromaFormatIdc + pfp_output_diff_chroma_format_idc Here, InputChromaFormatIdc is described in the SPS of the encoded data. The value of sps_chroma_format_idc is the identification value of the chroma format of the decoded image. A value of 0 indicates monochrome (4:0:0), a value of 1 indicates 4:2:0, a value of 2 indicates 4:2:2, and a value of 3 indicates 4:4:4. Post-filter The variable OutputChromaFormatIdc, which is the identification value of the chroma format output by the post-filtering process, is Input Similar to ChromaFormatIdc, a value of 0 indicates monochrome (4:0:0), a value of 1 indicates 4:2:0, a value of 2 indicates 4:2:2, and a value of 3 indicates 4:4:4.
[0311] Note that in this embodiment, in order to derive OutputChromaFormatIdc, the difference value between the identification value of the chroma format to be output and the identification value of the input chroma format is used, and the value of OutputChromaFormatIdc is set to be the same as or greater than the value of InputChromaFormatIdc. However, without using the difference value, OutputChromaFormatIdc may be directly used as a syntax element. In this case, the value of OutputChromaFormatIdc can be defined independently of the value of InputChromaFormatIdc which has the advantage. In this case, if the syntax element OutputChromaFormatIdc does not exist it is assumed to be the same as the value of InputChromaFormatIdc.
[0312] The NN filter section 611 assumes php_output_diff_chroma_format_idc to be 0 when the update information php_output_diff_chroma_format_idc does not exist in the encoded data. Thus, when the php_output_diff_chroma_format_idc does not exist in the encoded data, the NN filter section 611 may be set as follows. OutputChromaFormatIdc = ChromaFormatIdc A moving image decoding apparatus includes a predicted image derivation unit that decodes a predicted image as described above and a residual decoding unit that decodes a residual. When there is no filter update information, the header decoding unit 3020 estimates the color difference format of the output related to color difference subsampling in the input color difference format, thereby achieving the effect that even when the filter update information is not specified, accurate filter processing can be used for update.
[0313] The NN filter unit 611 derives a variable indicating the color difference subsampling of the output image as follows as follows.
[0314] outSubWidthC = outSW = 1, outSubHeightC = outSH = 1 (OutputChromaFormatIdc == 0) outSubWidthC = outSW = 2, outSubHeightC = outSH = 2 (OutputChromaFormatIdc == 1) outSubWidthC = outSW = 2, outSubHeightC = outSH = 1 (OutputChromaFormatIdc == 2) outSubWidthC = outSW = 1, outSubHeightC = outSH = 1 (OutputChromaFormatIdc == 3) In this way, by defining the input components and output format of the post-filter processing for color difference format conversion, the input and output data of the post-filter processing for color difference format conversion can be clarified.
[0315] Note that the pfp_component_idc that specifies the color component to which the above post-filter processing is applied distinguishes between luminance and color difference, but it may simply indicate the number of components. Specifically it may have the following semantics.
[0316] A value of 0 for pfp_component_idc indicates applying post-filtering to one component. This is shown as follows.
[0317] A value of 1 for pfp_component_idc indicates applying post-filtering to two components. This is shown as follows.
[0318] A value of 2 for pfp_component_idc indicates applying post-filtering to all three components. This is shown as follows.
[0319] The NN filter unit 611 may switch the NN model according to pfp_component_idc.
[0320] When pfp_component_idc == 0: The NN filter unit 611 selects an NN model that derives a one-channel three-dimensional tensor from a one-channel three-dimensional tensor and performs filtering. This is shown as follows.
[0321] When pfp_component_idc == 1: The NN filter unit 611 selects an NN model that derives a two-channel three-dimensional tensor from a two-channel three-dimensional tensor and performs filtering. This is shown as follows.
[0322] When pfp_component_idc == 2: The NN filter unit 611 selects an NN model that derives a three-channel three-dimensional tensor from a three-channel three-dimensional tensor and performs filtering. This is shown as follows.
[0323] As described above, there is an effect of reducing the processing amount by selecting an appropriate NN model according to the color component to be applied.
[0324] The NN filter unit 611 has one channel for one component and two channels for two components according to pfp_component_idc. The nnrpf_input_format_idc may be derived as follows so that it becomes an NN model with 2 channels for 2 components and 3 channels for 3 components. nnrpf_input_format_idc = 0 (pfp_component_idc==0) nnrpf_input_format_idc = 1 (pfp_component_idc==1) nnrpf_input_format_idc = 2 (pfp_component_idc==2) That is, nnrpf_input_format_idc = pfp_component_idc As another example, the NN filter unit 611 may derive the nnrpf_input_format_idc as follows so that it becomes an NN model with 1 channel for 1 component, 2 channels for 2 components, and 6 channels for 3 components according to pfp_component_idc. nnrpf_input_format_idc = 0 (pfp_component_idc==0) nnrpf_input_format_idc = 1 (pfp_component_idc==1) nnrpf_input_format_idc = 3 (pfp_component_idc==2) That is, nnrpf_input_format_idc = pfp_component_idc < 2? pfp_component_idc : 3 As described above, there is an effect that an appropriate tensor format of the NN model can be selected according to the color component to be applied and the processing can be performed.
[0325] The NN filter unit 611 directly derives the inputTensor according to pfp_component_idc as described above. It may be output. Alternatively, the NN filter unit 611 may switch the value of the input image ComponentSample and the value of the NN output data outTensor according to pfp_component_idc, and derive the output image in the following process. When pfp_component_idc == 0 outSamplesL[x][y] = outTensor[0][y][x] outSamplesCb[x*2 / outSW][y*2 / outSH] = ComponentSample[1][x*2 / SW][y*2 / SH] outSamplesCr[x*2 / outSW][y*2 / outSH] = ComponentSample[2][x*2 / SW][y*2 / SH] When pfp_component_idc == 1 outSamplesL[x][y] = ComponentSample[0][x][y] outSamplesCb[x*2 / outSW][y*2 / outSH] = outTensor [0][x*2 / SW][y*2 / SH] outSamplesCr[x*2 / outSW][y*2 / outSH] = outTensor [1][x*2 / SW][y*2 / SH] When pfp_component_idc == 2 outSamplesL[x][y] = outTensor[0][x][y] outSamplesCb[x*2 / outSW][y*2 / outSH] = outTensor[1][x*2 / SW][y*2 / SH] outSamplesCr[x*2 / outSW][y*2 / outSH] = outTensor[2][x*2 / SW][y*2 / SH] Also, in the above example, the identification value of the color difference format output by the post-filter processing is shown as the difference value from the identification value of the input color difference format, but it may also be directly described by the syntax element.
[0326] In this embodiment, a post-filter target SEI is defined independently of the NN post-filter SEI, and the inputs and outputs of the post-filter processing are defined. However, a syntax similar to that of the NN post-filter SEI may be defined, and the problems can be solved similarly.
[0327] Describing this embodiment with reference to FIG. 1, an image decoding device that decodes encoded data obtained by encoding an image is provided, and a resolution inverse conversion device that converts the resolution of the image decoded by the image decoding device. The moving image decoding device is characterized by having an inverse conversion information decoding device that decodes color component information input to the resolution inverse conversion device and color difference format information output therefrom.
[0328] Further, a moving image encoding device includes an image encoding device that encodes an image, and an inverse conversion information encoding device that encodes color component information input to a resolution inverse conversion device that converts the resolution of the encoded image and color difference format information output therefrom.
[0329] A moving image decoding device according to an aspect of the present invention is a moving image decoding device including a predicted image derivation unit that decodes a predicted image and a residual decoding unit that decodes a residual, characterized in that an input tensor is derived from parameters specifying the number of channels of the input tensor and the output tensor of a neural network model, or an image is derived from the output tensor.
[0330] A moving image decoding device including a predicted image derivation unit that decodes a predicted image and a residual decoding unit that decodes a residual, decodes input / output image information specifying the number of channels of the input / output tensors of a neural network model and additional input information, derives a part of the input tensor from a first image using the input / output image information, and further uses the additional input information to obtain a second image different from the first image or encodes information related to the derivation of the predicted image or the decoding of the residual to derive another part of the input tensor.
[0331] Using the parameters of color difference subsampling to change the addition value of the loop variable, performing loop processing on the raster scan, and within the loop, deriving an input tensor from an image, and applying a deep learning filter to the input tensor to derive an output tensor.
[0332] A moving image decoding device including a predicted image derivation unit for decoding a predicted image and a residual decoding unit for decoding a residual, wherein the header unit estimates the color difference format of the output related to color difference subsampling in the input color difference format when there is no filter update information.
[0333] A moving image encoding device according to an aspect of the present invention is a moving image encoding device including a predicted image derivation unit for decoding a predicted image and a residual encoding unit for encoding a residual, Deriving an input tensor from parameters specifying the number of channels of the input tensor and the output tensor of a neural network model, or deriving an image from the output tensor.
[0334] Embodiments of the present invention are not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. That is, embodiments obtained by appropriately combining technical means modified within the scope shown in the claims are also included in the technical scope of the present invention.
Industrial Applicability
[0335] Embodiments of the present invention can be suitably applied to a moving image decoding device that decodes encoded data in which image data is encoded, and a moving image encoding device that generates encoded data in which image data is encoded. Further, it can be suitably applied to the data structure of the encoded data generated by the moving image encoding device and referred to by the moving image decoding device.
Explanation of Signs
[0336] 1 Moving image transmission system 30 Moving image decoding device 31 Image decoding device 301 Entropy decoding section 302 Parameter decoding section 303 Inter-prediction parameter derivation section 304 Intra-prediction parameter derivation section 305, 107 Loop filter 306, 109 Reference picture memory 307, 108 Prediction parameter memory 308, 101 Predicted image generation section 309 Inter-predicted image generation section 310 Intra-predicted image generation section 311, 105 Inverse quantization and inverse transformation section 312, 106 Addition section 320 Prediction parameter derivation section 10 Moving image encoding device 11 Image encoding device 102 Subtraction section 103 Transformation and quantization section 104 Entropy encoding section 110 Encoding parameter determination section 111 Parameter encoding section 112 Inter-prediction parameter encoding section 113 Intra-prediction parameter encoding section 120 Prediction parameter derivation section 71 Inverse transformation information creation device 81 Inverse transformation information encoding device 91 Inverse transformation information decoding device 611 NN filter section
Claims
1. An image decoding device that decodes an image from encoded data, comprising: a header decoding unit for deriving a filter strength and a bit depth for the filtering process; a neural network filter unit that derives input image information that specifies a format of an input tensor for a neural network model; the neural network filter unit derives a part of the input tensor using an image based on the input image information, sets additional information in another part of the input tensor, derives an output tensor by performing a neural network filtering process using the input tensor, and derives a filtered image from the output tensor; An image decoding device, characterized in that the additional information is derived using the filter strength and the bit depth.
2. An image encoding device that encodes an image, comprising: a header encoder for deriving a filter strength and a bit depth for the filtering process; a neural network filter unit that derives input image information that specifies a format of an input tensor for a neural network model; the neural network filter unit derives a part of the input tensor using an image based on the input image information, sets additional information in another part of the input tensor, derives an output tensor by performing a neural network filtering process using the input tensor, and derives a filtered image from the output tensor; An image encoding device characterized in that the additional information is derived using the filter strength and the bit depth.
3. 1. An image decoding method for decoding an image from encoded data, comprising the steps of: Derive a filter strength and a bit depth for the filtering process; Derive input image information that specifies a format of an input tensor for the neural network model; deriving a portion of the input tensor using an image based on the input image information; Set additional information in another part of the input tensor; Deriving an output tensor by performing a neural network filtering process using the input tensor; An image decoding method comprising: deriving a filtered image from the output tensor.
4. 1. An image encoding method for encoding an image, comprising the steps of: Derive a filter strength and a bit depth for the filtering process; Derive input image information that specifies a format of an input tensor for the neural network model; deriving a portion of the input tensor using an image based on the input image information; Set additional information in another part of the input tensor; Deriving an output tensor by performing a neural network filtering process using the input tensor; Deriving a filtered image from the output tensor; An image coding method, characterized in that the additional information is derived using the filter strength and the bit depth.
Citation Information
Patent Citations
Supplemental enhancement information messages for neural network based video post processing
US20200304836A1
Neural network representation formats
WO2021064013A2