Video decoding device, video encoding device, and transmission method

JP2024092440A5Pending Publication Date: 2025-12-10SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022208358
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Existing video encoding and decoding methods, such as those described in Non-Patent Documents 1 and 2, face issues with image quality deterioration at low transmission rates due to encoding distortion and inefficiencies in generating auxiliary extension information when the display order of pictures differs from the encoding/decoding order, particularly in prediction structures like Hierarchical Bi-prediction.

Method used

A video decoding device and encoding device that utilize a post-filter processing system with neural networks, where auxiliary extension information is generated and encoded to indicate whether to perform post-filter processing on a picture-by-picture basis, considering the decoding or display output order, and specify the neural network to be applied, ensuring efficient encoding and decoding regardless of prediction structure.

Benefits of technology

This configuration improves image quality and enables efficient encoding and decoding of auxiliary extension information even at low transmission rates, addressing the issues of encoding distortion and order discrepancies in existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To solve a problem in which, in video coding and decoding methods, it is possible to improve the image quality, which has been deteriorated by coding distortion due to low transmission rate, by post-filtering using a neural network, however, depending on the prediction structure, auxiliary enhancement information cannot be generated efficiently.SOLUTION: A video decoding device according to an aspect of the present invention includes an image decoding device that decodes encoded data to generate a decoded image, a post-filter processing device that performs post-filter processing on the decoded image, and an auxiliary enhancement information decoding device that decodes auxiliary enhancement information indicating whether to perform post-filter processing on a picture-by-picture basis in the post-filter processing device, and in the auxiliary enhancement information, when it is determined whether to continue the post-filter processing on the picture-by-picture basis, information indicating whether to continue in decoding order or in display output order is decoded.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] An embodiment of the present invention relates to a video encoding device and a video decoding device. [Background technology]

[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding moving images, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.

[0003] Specific examples of video encoding methods include H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding).

[0004] In such a video coding method, images (pictures) constituting a video are managed in a hierarchical structure consisting of slices obtained by dividing images, coding tree units (CTUs) obtained by dividing slices, coding units (sometimes called coding units: CUs) obtained by dividing coding tree units, and transform units (TUs) obtained by dividing coding units, and are coded / decoded for each CU.

[0005] In such a video coding method, a predicted image is usually generated based on a locally decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods for generating a predicted image include inter-prediction and intra-prediction.

[0006] Moreover, Non-Patent Document 1 can be cited as a recent example of a video encoding and decoding technique.

[0007] Non-Patent Document 1 discloses a video encoding and decoding method with extremely high encoding efficiency.

[0008] Non-Patent Document 2 specifies a supplemental enhancement information (SEI) message for transmitting image properties, display methods, timing, and the like simultaneously with encoded data, and discloses an SEI that transmits the topology and parameters of a neural network filter used as a post-filter in units that allow random access. It also discloses an SEI that transmits whether post-filter processing is to be performed in picture units. [Prior art documents] [Non-patent literature]

[0009] [Non-Patent Document 1] ITU-T Recommendation H.266 [Non-Patent Document 2] Text of ISO / IEC 23002-7:202x (2nd Ed.) DAM1 Additional SEI messages, Nov. 2022. Summary of the Invention [Problem to be solved by the invention]

[0010] However, although the method disclosed in Non-Patent Document 1 is a video encoding and decoding method with extremely high coding efficiency, there is a problem in that when the transmission rate is low, image quality deteriorates due to coding distortion.

[0011] In addition, in the method disclosed in Non-Patent Document 2, it is possible to improve image quality when the transmission rate is low by post-filtering using a neural network. However, when efficiently compressing, encoding, transmitting, and decoding moving images, depending on the prediction structure, the display order of pictures and the encoding and decoding order may differ. In such cases, there is a problem that auxiliary enhancement information cannot be generated efficiently. [Means for solving the problem]

[0012] A video decoding device according to one embodiment of the present invention is characterized in that it includes an image decoding device that decodes encoded data to generate a decoded image, a postfilter processing device that performs postfilter processing on the decoded image, an auxiliary enhancement information decoding device that decodes auxiliary enhancement information indicating whether or not to perform postfilter processing on a picture-by-picture basis in the postfilter processing device, and in that the auxiliary enhancement information decodes information indicating whether to continue postfilter processing on a picture-by-picture basis, whether to continue in decoding order or in display output order.

[0013] The auxiliary enhancement information decoding device is also characterized in that it decodes information for specifying a neural network to be applied to postfilter processing.

[0014] A moving image coding device according to one embodiment of the present invention is characterized in that it comprises an image coding device which codes an input image, an auxiliary extension information generating device which generates auxiliary extension information indicating whether or not to perform post-filter processing on a picture-by-picture basis, and an auxiliary extension information coding device which codes, in the auxiliary extension information, information indicating whether to continue post-filter processing on a picture-by-picture basis, to continue in decoding order or in display output order.

[0015] The auxiliary extension information encoding device is characterized in that it encodes information for specifying a neural network to which postfiltering is to be applied as auxiliary extension information. Effect of the Invention

[0016] This configuration can solve the problem that, depending on the prediction structure, when the picture display order differs from the encoding and decoding order, supplementary enhancement information cannot be generated efficiently. [Brief description of the drawings]

[0017] [Figure 1] 1 is a schematic diagram showing a configuration of a video transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of encoded data. [Diagram 3] FIG. 2 is a diagram showing a hierarchical structure of encoded data in units of sequences. [Figure 4] FIG. 1 is a schematic diagram showing a configuration of an image decoding device. [Diagram 5] 11 is a flowchart illustrating a schematic operation of an image decoding device. [Figure 6] FIG. 1 is a block diagram showing a configuration of an image encoding device. [Figure 7] FIG. 1 is a diagram showing an overview of the syntax of the neural network post-filter characteristic (NNPFC) SEI. [Figure 8] 11 is a diagram showing an example of the configuration of a syntax table of a neural network postfilter characteristic (NNPFC) SEI that defines auxiliary extension information in this embodiment. FIG. [Figure 9] A diagram showing the syntax of Neural Network Post Filter Activation (NNPFA) SEI. [Figure 10] This is an example of a prediction structure in which the display output order and the encoding and decoding order are different. [Figure 11] A figure showing the syntax of the Neural Network Post Filter Activation (NNPFA) SEI that specifies the auxiliary extension information of this embodiment. [Figure 12] A diagram showing the syntax of an SEI payload, which is a container for an SEI message. [Figure 13] FIG. 13 is a flowchart showing the processing of the post-filter processor 61. [Figure 14]FIG. 1 is a diagram showing an encoding device and a decoding device of an NNC. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0019] FIG. 1 is a schematic diagram showing the configuration of another moving image transmission system according to the present embodiment.

[0020] The video transmission system 1 is a system that transmits coded data obtained by coding an image, and decodes and displays the coded data. The video transmission system 1 is made up of a video coding device 10, a network 21, a video decoding device 30, and an image display device 41.

[0021] The video encoding device 10 is composed of an image encoding device (image encoding unit) 11, an auxiliary extension information creation device (auxiliary extension information creation unit) 71, an auxiliary extension information encoding device (auxiliary extension information encoding unit) 81, and a pre-filter processing device (pre-filter processing unit) 1001.

[0022] The video encoding device 10 creates a prefiltered image T2 from an input video T1 using a prefilter processing device 51, compresses and encodes the image using an image encoding device 11, and analyzes the input video T1 and a locally decoded image T3 from the image encoding device 11 to generate auxiliary extension information to be input to a postfilter processing device 61 using an auxiliary extension information creation device 71, encodes the information using an auxiliary extension information encoding device 81, generates encoded data Te, and transmits the encoded data Te to a network 21.

[0023] The video decoding device 30 includes an image decoding device (image decoding unit) 31, an auxiliary extension information decoding device (auxiliary extension information decoding unit) 91, and a post-filter processing device (host filter processing unit) 61.

[0024] The video decoding device 30 decodes the encoded data Te received from the network 21 using the image decoding device 31 and the auxiliary extension information decoding device 91, and performs post-filter processing on the decoded image Td1 using the auxiliary extension information in the post-filter processing device 61, and outputs the post-filter decoded image Td2 to the image display device 41.

[0025] The postfilter processor 61 may output the decoded image Td1 as is without performing postfiltering on the auxiliary enhancement information.

[0026] The image display device 41 displays all or a part of the post-filter image Td2 output from the post-filter processing device 1002. The image display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Examples of the display form include a stationary display, a mobile display, and an HMD. When the image decoding device 31 has high processing capability, it displays a high-quality image, and when it has only low processing capability, it displays an image that does not require high processing capability or display capability.

[0027] The network 21 transmits the encoded auxiliary extension information and the encoded data Te to the image decoding device 31. A part or all of the encoded auxiliary extension information may be included in the encoded data Te as auxiliary extension information SEI. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. The network 21 may also be replaced by a storage medium on which the encoded data Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).

[0028] As an example of a specific embodiment, in pre-filter processing, the input image may be reduced, and as auxiliary extension information, in post-filter processing, auxiliary extension information for neural network processing for enlarging the decoded image by super-resolution processing based on a neural network may be encoded and decoded.

[0029] As another example of a specific embodiment, in pre-filter processing, no particular processing is performed, and as auxiliary extension information, in post-filter processing, auxiliary extension information for neural network processing for restoring the decoded image to the input moving image by image restoration processing based on a neural network may be encoded and decoded.

[0030] In such a configuration, a framework is provided that enables efficient encoding and decoding of auxiliary extension information.

[0031] <Operator> The operators used in this specification are described below.

[0032] [[ID=18>> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is an OR assignment operator, and || represents a logical OR.

[0033] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0).

[0034] Clip3(a, b, c) is a function that clips c to a value between a and b (inclusive). It returns a if c < a, b if c > b, and c otherwise (where a <= b).

[0035] abs(a) is a function that returns the absolute value of a.

[0036] Int(a) is a function that returns the integer value of a.

[0037] floor(a) is a function that returns the largest integer less than or equal to a.

[0038] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0039] a / d represents the division of a by d (truncated to an integer).

[0040] (Structure of encoded data Te) Before going into a detailed description of the image encoding device 11 and the image decoding device 31 according to this embodiment, the data structure of the encoded data Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be described with reference to Figures 2 and 3.

[0041] The coded data Te is a bitstream consisting of multiple CVS (Coded Video Sequence) and EoB (End of Bitstream) NAL units as shown in FIG. 2. The CVS consists of multiple AU (Access Unit) and EoS (End of Sequence) NAL units. The AU at the beginning of the CVS is called the CVSS (Coded Video Sequence Start) AU. The unit obtained by dividing the CVS into layers is called the CLVS (Coded Layer Video Sequence). The AU consists of one or more layered PUs (Picture Units) with the same output time. If the multilayer coding method is not adopted, the AU consists of one PU. The PU is a unit of coded data for one decoded picture consisting of multiple NAL units. The CLVS consists of PUs of the same layer, and the PU at the beginning of the CLVS is called the CLVSS (Coded Layer Video Sequence Start) PU. The CLVSS PU is limited to PUs that are randomly accessible IRAP (Intra Random Access Pictures) or GDR (Gradual Decoder Refresh Picture). A NAL unit consists of a NAL unit header and RBSP (Raw Byte Sequence Payload) data. The NAL unit header consists of 2 bits of 0 data, followed by a 6-bit nuh_layer_id that indicates the layer value, a 5-bit nuh_unit_type that indicates the NAL unit type, and a 3-bit nuh_temporal_id_plus1 that is the Temporal ID value plus 1.

[0042] Fig. 3 is a diagram showing a hierarchical structure of data in the coded data Te in units of PU. The coded data Te illustratively includes a sequence and a plurality of pictures constituting the sequence. Fig. 3 shows a diagram showing a coded video sequence that defines the sequence SEQ, a coded picture that defines the picture PICT, a coded slice that defines the slice S, coded slice data that defines the slice data, a coding tree unit included in the coded slice data, and a coding unit included in the coding tree unit.

[0043] In the coded video sequence, a set of data to be referred to by the image decoding device 31 in order to decode the sequence SEQ to be processed is defined. As shown in Fig. 3, the sequence SEQ includes a video parameter set VPS (Video Parameter Set), a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information).

[0044] The video parameter set VPS specifies a set of coding parameters common to multiple videos composed of multiple layers, as well as a set of coding parameters related to multiple layers and each individual layer included in the video.

[0045] The sequence parameter set SPS specifies a set of coding parameters that the image decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs is selected from the PPS.

[0046] Here, the sequence parameter set SPS includes the following syntax elements: ref_pic_resampling_enabled_flag: A flag that specifies whether or not to use a function that changes the resolution (resampling) when decoding each image included in a single sequence that references the target SPS. In other words, this flag indicates that the size of the reference picture referenced in generating a predicted image changes between each image indicated by a single sequence. If the value of this flag is 1, the resampling is applied, and if the value is 0, it is not applied. pic_width_max_in_luma_samples: A syntax element that specifies the width of the image with the largest width among the images in a single sequence, in units of luminance blocks. The value of this syntax element must not be 0 and must be an integer multiple of Max(8, MinCbSizeY), where MinCbSizeY is a value determined by the minimum size of a luminance block. pic_height_max_in_luma_samples: This syntax element specifies the height of the image with the maximum height among the images in a single sequence, in units of luminance blocks. The value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY).

[0047] The picture parameter set PPS defines a set of coding parameters that the image decoding device 31 refers to in order to decode each picture in the target sequence. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected for each picture in the target sequence.

[0048] Here, the picture parameter set PPS includes the following syntax elements: pps_pic_width_in_luma_samples: A syntax element that specifies the width of the target picture. The value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and equal to or less than sps_pic_width_max_in_luma_samples. pps_pic_height_in_luma_samples: A syntax element that specifies the height of the target picture. The value of this syntax element is required to be a non-zero integer multiple of Max(8, MinCbSizeY) and equal to or less than sps_pic_height_max_in_luma_samples. pps_conformance_window_flag: a flag indicating whether conformance (cropping) window offset parameters are to be signaled subsequently, and where the conformance window is to be displayed. If this flag is 1, the parameters are to be signaled, if it is 0, the conformance window offset parameters are not present. pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, pps_conf_win_bottom_offset: offset values ​​for specifying the left, right, top, and bottom positions of a picture output by decoding, with respect to a rectangular area specified by the output picture coordinates. If the value of pps_conformance_window_flag is 0, the values ​​of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are estimated to be 0.

[0049] Here, the variable ChromaFormatIdc of the chrominance format is the value of sps_chroma_format_id, and the variables SubWidthC and SubHightC are values ​​determined by this ChromaFormatIdc. In the case of a monochrome format, SubWidthC and SubHightC are both 1, in the case of a 4:2:0 format, SubWidthC and SubHightC are both 2, in the case of a 4:2:2 format, SubWidthC is 2 and SubHightC is 1, and in the case of a 4:4:4 format, SubWidthC and SubHightC are both 1. · pps_init_qp_minus26 is information for deriving the quantization parameter SliceQpY of the slice referenced by the PPS.

[0050] (Subpicture) A picture may be further divided into rectangular sub-pictures. The size of a sub-picture may be a multiple of the CTU. A sub-picture is defined as a set of tiles that are an integer number of consecutive tiles vertically and horizontally. In other words, a picture is divided into rectangular tiles, and a sub-picture is defined as a set of rectangular tiles. A sub-picture may be defined using the IDs of the top left tile and the bottom right tile of the sub-picture.

[0051] Fig. 6 is a conceptual diagram of an image to be processed in the video transmission system 1, showing changes in the resolution of the image over time. However, Fig. 6 does not distinguish whether the image is encoded or not. Fig. 6 shows an example in which an image is transmitted to the image decoding device 31 while adaptively changing the resolution using a picture parameter set PPS in the process of the video transmission system 1.

[0052] (Encoded Picture) A coded picture defines a set of data to be referenced by the image decoding device 31 in order to decode a picture PICT to be processed. As shown in Fig. 3, the picture PICT includes a picture header PH and slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).

[0053] Hereinafter, when there is no need to distinguish between slices 0 to NS-1, the subscripts of the symbols may be omitted. The same applies to other data that are included in the coded data Te described below and have subscripts.

[0054] The picture header contains the following syntax elements:

[0055] pic_temporal_mvp_enabled_flag is a flag that specifies whether or not temporal motion vector prediction is used for inter prediction of a slice associated with the picture header. If the value of the flag is 0, the syntax elements of the slice associated with the picture header are restricted so that temporal motion vector prediction is not used in decoding the slice. If the value of the flag is 1, it indicates that temporal motion vector prediction is used in decoding the slice associated with the picture header. If the flag is not specified, the value is presumed to be 0.

[0056] (Coded Slice) A coded slice defines a set of data to be referenced by the image decoding device 31 in order to decode a current slice S. As shown in Fig. 3, a slice includes a slice header and slice data.

[0057] The slice header includes a group of coding parameters to be referred to by the image decoding device 31 in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates a slice type is an example of a coding parameter included in the slice header.

[0058] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction when encoding, (2) a P slice that uses uni-prediction (L0 prediction) or intra prediction when encoding, and (3) a B slice that uses uni-prediction (L0 prediction or L1 prediction), bi-prediction, or intra prediction when encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P or B slice, it refers to a slice including a block that can use inter prediction.

[0059] In addition, the slice header may include a reference to a picture parameter set PPS (pic_parameter_set_id).

[0060] (Encoded slice data) The coded slice data specifies a set of data to be referenced by the image decoding device 31 in order to decode the slice data to be processed. The slice data includes a CTU, as shown in the coded slice header in Fig. 3. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice, and is also called a Largest Coding Unit (LCU).

[0061] (coding tree unit) 3 specifies a set of data that the image decoding device 31 refers to in order to decode a CTU to be processed. The CTU is divided into coding units CU, which are basic units of encoding processing, by recursive quad tree division (QT (Quad Tree) division), binary tree division (BT (Binary Tree) division), or ternary tree division (TT (Ternary Tree) division). BT division and TT division are collectively called multi tree division (MT (Multi Tree) division). A node of a tree structure obtained by recursive quad tree division is called a coding node. Intermediate nodes of the quad tree, binary tree, and ternary tree are coding nodes, and the CTU itself is specified as the top coding node.

[0062] CT includes, as CT information, a CU split flag (split_cu_flag) indicating whether CT splitting is performed, a QT split flag (qt_split_cu_flag) indicating whether QT splitting is performed, an MT split direction (mtt_split_cu_vertical_flag) indicating the split direction of MT splitting, and an MT split type (mtt_split_cu_binary_flag) indicating the split type of MT splitting. split_cu_flag, qt_split_cu_flag, mtt_split_cu_vertical_flag, and mtt_split_cu_binary_flag are transmitted for each encoding node.

[0063] Different trees may be used for luminance and chrominance. The type of tree is indicated by treeType. For example, when using a common tree for luminance (Y, cIdx=0) and chrominance (Cb / Cr, cIdx=1,2), the common single tree is indicated by treeType=SINGLE_TREE. When using two different trees (DUAL trees) for luminance and chrominance, the luminance tree is indicated by treeType=DUAL_TREE_LUMA and the chrominance tree is indicated by treeType=DUAL_TREE_CHROMA.

[0064] (Encoding Unit) 3 specifies a set of data to be referenced by the image decoding device 31 in order to decode a coding unit to be processed. Specifically, a CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header specifies a prediction mode, etc.

[0065] The prediction process may be performed on a CU basis, or on a sub-CU basis by further dividing a CU. If the size of a CU and a sub-CU are the same, there is one sub-CU in the CU. If the size of a CU is larger than that of a sub-CU, the CU is divided into sub-CUs. For example, if the CU is 8x8 and the sub-CU is 4x4, the CU is divided into 2 parts horizontally and 2 parts vertically, into 4 sub-CUs.

[0066] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is a prediction within the same picture, while inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).

[0067] The transform and quantization processes are performed in units of CUs, but the quantized transform coefficients may be entropy coded in units of sub-blocks such as 4x4.

[0068] (Prediction parameters) The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.

[0069] Hereinafter, prediction parameters of inter prediction will be described. Inter prediction parameters are composed of prediction list use flags predFlagL0 and predFlagL1, reference picture indexes refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether or not a reference picture list (L0 list, L1 list) is used, and when the value is 1, the corresponding reference picture list is used. Note that in this specification, when "a flag indicating whether or not XX" is written, a flag other than 0 (for example, 1) is XX, 0 is not XX, and 1 is treated as true and 0 is treated as false in logical negation, logical product, etc. (similarly below). However, in an actual device or method, other values ​​can be used as true and false values.

[0070] Syntax elements for deriving inter-prediction parameters include, for example, an affine flag affine_flag used in merge mode, a merge flag merge_flag, a merge index merge_idx, an MMVD flag mmvd_flag, an inter-prediction identifier inter_pred_idc for selecting a reference picture to be used in AMVP mode, a reference picture index refIdxLX, a prediction vector index mvp_LX_idx for deriving a motion vector, a difference vector mvdLX, and a motion vector precision mode amvr_mode.

[0071] (Configuration of an image decoding device) The configuration of an image decoding device 31 (FIG. 4) according to this embodiment will be described.

[0072] The image decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, a prediction image generating unit (prediction image generating device) 308, an inverse quantization and inverse transform unit 311, an adder 312, and a prediction parameter derivation unit 320. Note that, in accordance with the image encoding device 11 described below, the image decoding device 31 may also be configured not to include the loop filter 305.

[0073] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022 (prediction mode decoding unit), and the CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice header (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data. The TU decoding unit 3024 decodes QP update information (quantization correction value) and quantization prediction error (residual_coding) from the encoded data when a prediction error is included in the TU.

[0074] In cases other than the skip mode (skip_mode==0), the TU decoding unit 3024 decodes the QP update information and the quantized prediction error from the encoded data. More specifically, in cases of skip_mode==0, the TU decoding unit 3024 decodes a flag cu_cbp indicating whether or not the current block includes a quantized prediction error, and decodes the quantized prediction error when cu_cbp is 1. When cu_cbp does not exist in the encoded data, it derives 0.

[0075] The TU decoding unit 3024 decodes an index mts_idx indicating a transformation base from the coded data. The TU decoding unit 3024 also decodes an index stIdx indicating the use of a secondary transformation and a transformation base from the coded data. When stIdx is 0, it indicates no application of a secondary transformation, when it is 1, it indicates one transformation of a set (pair) of secondary transformation bases, and when it is 2, it indicates the other transformation of the pair.

[0076] Furthermore, the TU decoding unit 3024 may decode a sub-block transform flag cu_sbt_flag. When cu_sbt_flag is 1, the CU is divided into a plurality of sub-blocks, and the residual of only one specific sub-block is decoded. Furthermore, the TU decoding unit 3024 may decode a flag cu_sbt_quad_flag indicating whether the number of sub-blocks is 4 or 2, cu_sbt_horizontal_flag indicating the division direction, and cu_sbt_pos_flag indicating a sub-block including a non-zero transform coefficient.

[0077] The predicted image generating unit 308 includes an inter predicted image generating unit 309 and an intra predicted image generating unit 310 .

[0078] In the following, an example will be described in which CTU and CU are used as processing units, but the present invention is not limited to this example and processing may be performed in sub-CU units. Alternatively, CTU and CU may be read as blocks and sub-CU as sub-blocks, and processing may be performed in block or sub-block units.

[0079] The entropy decoding unit 301 performs entropy decoding on the coded data Te input from the outside, and decodes each code (syntax element). There are two types of entropy coding: a method of variable-length coding the syntax element using a context (probability model) adaptively selected according to the type of syntax element and surrounding circumstances, and a method of variable-length coding the syntax element using a predetermined table or formula. The former CABAC (Context Adaptive Binary Arithmetic Coding) stores the CABAC state of the context (probability state index pStateIdx that specifies the type (0 or 1) and probability of the most probable symbol) in memory. The entropy decoding unit 301 initializes all CABAC states at the beginning of a segment (tile, CTU row, slice). The entropy decoding unit 301 converts the syntax element into a binary string (Bin String) and decodes each bit of the Bin String. When a context is used, a context index ctxInc is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the used context is updated. Bits that do not use a context are decoded with equal probability (EP, bypass), and the derivation of ctxInc and the CABAC state are omitted. The decoded syntax elements include prediction information for generating a predicted image and a prediction error for generating a difference image.

[0080] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The decoded code is, for example, a prediction mode predMode, merge_flag, merge_idx, inter_pred_idc, refIdxLX, mvp_LX_idx, mvdLX, amvr_mode, etc. Control of which code to decode is performed based on an instruction from the parameter decoding unit 302.

[0081] (Basic flow) FIG. 5 is a flowchart illustrating a schematic operation of the image decoding device 31.

[0082] (S1100: Decode Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as the VPS, SPS, and PPS from the encoded data.

[0083] (S1200: Decode slice information) The header decoding unit 3020 decodes the slice header (slice information) from the encoded data.

[0084] Thereafter, the image decoding device 31 repeats the processes from S1300 to S5000 for each CTU included in the target picture to derive a decoded image of each CTU.

[0085] (S1300: Decode CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.

[0086] (S1400: Decode CT Information) The CT information decoding unit 3021 decodes the CT from the encoded data.

[0087] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data.

[0088] (S1510: Decoding CU information) The CU decoding unit 3022 decodes the CU information, prediction information, the TU split flag split_transform_flag, the CU residual flags cbf_cb, cbf_cr, cbf_luma, and the like from the encoded data.

[0089] (S1520: TU information decoding) When a prediction error is included in a TU, the TU decoding unit 3024 decodes the QP update information, the quantization prediction error, and the transform index mts_idx from the encoded data. Note that the QP update information is a difference value from the quantization parameter predicted value qPpred, which is a predicted value of the quantization parameter QP.

[0090] (S2000: Generation of predicted image) The predicted image generation unit 308 generates a predicted image for each block included in the current CU based on prediction information.

[0091] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transform unit 311 executes inverse quantization and inverse transform processing for each TU included in the target CU.

[0092] (S4000: Generate decoded image) The adder 312 adds the predicted image supplied from the predicted image generation unit 308 and the prediction error supplied from the inverse quantization and inverse transform unit 311 to generate a decoded image of the current CU.

[0093] (S5000: Loop Filter) The loop filter 305 applies a loop filter such as a deblocking filter, SAO, or ALF to the decoded image to generate a decoded image.

[0094] Non-Patent Document 1 describes a video encoding and decoding method with extremely high encoding efficiency, but when the transmission rate is low, there is a problem that image quality deteriorates due to encoding distortion.

[0095] In addition, in Non-Patent Document 2, it is possible to improve image quality when the transmission rate is low by post-filtering using a neural network. However, when efficiently compressing, encoding, transmitting, and decoding moving images, depending on the prediction structure, the display order of pictures and the encoding and decoding order may differ. In such cases, there is a problem that auxiliary enhancement information cannot be generated efficiently.

[0096] In this embodiment, even when the transmission rate is low, it is possible to improve the image quality and efficiently code and decode the supplemental extension information regardless of the prediction structure.

[0097] (Neural network post-filter characteristics (NNPFC) SEI Fig. 7 shows an outline of the syntax of the Neural Network Post-Filter Characteristics (NNPFC) SEI message of Non-Patent Document 2. The NNPFC SEI message specifies the neural network to be applied as a post-filter process. The application of a specific post-filter process to a specific picture is indicated by the Neural Network Post-Filter Activation SEI message (Fig. 8).

[0098] To apply this SEI message, the following variables must be defined: The width and height in luminance pixels of the picture decoded by the image decoding device 31 A picture with idx in the range 0 to numInputPics-1 that is used as input for postfiltering, with luminance pixel array CroppedYPic[idx] and chrominance pixel array CroppedCbPic[idx] and CroppedCrPic[idx]. - Pixel bit depth of the luminance pixel array BitDepthY ·Pixel bit length of chrominance pixel array BitDepthC Variables SubWidthC and SubHeightC indicating the chrominance format indicated by the chrominance format ChromaFormatIdc of the picture decoded by the image decoding device 31. When the chrominance format is 4:2:0, the variables SubWidthC and SubHeightC are both 2, when the chrominance format is 4:2:2, the variables SubWidthC are 2 and the variable SubHeightC are 1, and when the chrominance format is 4:4:4, the variables SubWidthC and SubHeightC are both 1. nnpfc_auxiliary_inp_idc indicates the presence of auxiliary augmentation information in the input tensor of the neural network postfilter, and if its value is equal to 1, it inputs the deblocking filtering strength control value StrengthControlVal, which is a real number ranging from 0 to 1, as the auxiliary augmentation information.

[0099] The syntax element nnpfc_id indicates an identification number that can be used to identify the post-filter process.

[0100] If the NNPFC SEI message is the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value within the current CLVS, the following applies: Indicates that this SEI message is a basic post-filter process. · This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer until the end of the current CLVS. If an NNPFC SEI message is a repetition of a previous NNPFC SEI message in the current CLVS in decoding order, the successor semantics apply as if this SEI message was the only NNPFC SEI message with the same content in the current CLVS.

[0101] If the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies: This SEI message indicates that it is an update relative to the previous base post-filter in decoding order using the same nnpfc_id value. This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer until the end of the current CLVS or until the next NNPFC SEI message with the given nnpfc_id value in the current layer. nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bit stream that specifies post-filtering or is an update relative to an underlying post-processing filter with the same nnpfc_id value.

[0102] nnpfc_mode_idc equal to 1 indicates that the post-filter associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, of the form identified by the tag URI nnpfc_tag_uri.

[0103] nnpfc_reserved_zero_bit_a indicates 0.

[0104] nnpfc_tag_uri contains a tag URI with syntax and semantics specified in IETF RFC4151. It is used to store format and related information about the neural network used as the base postfilter, or for updates related to postfilters with the same nnpfc_id value. Note that nnpfc_tag_uri allows the format of the neural network data specified by nnrpf_uri to be uniquely identified without the need for a registration authority. If nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it indicates that the neural network data identified by nnpfc_uri is ISO / IEC 15938-17 compliant and encoded with NNC (Neural Network Coding).

[0105] The nnpfc_uri contains a URI with syntax and semantics specified in IETF Internet Standard 66 to identify the neural network to be used as a postfilter or update associated with a postfilter with the same nnpfc_id value.

[0106] nnpfc_formatting_and_purpose_flag equal to 1 indicates the presence of syntax elements related to the purpose, input format, output format, and complexity of the filter. nnpfc_formatting_and_purpose_flag equal to 0 indicates the absence of syntax elements related to the purpose, input format, output format, and complexity of the filter.

[0107] If this SEI message is the first NNPFC SEI message in decoding order with the particular nnpfc_id value in the current CLVS, then nnpfc_formatting_and_purpose_flag shall be equal to 1. If this SEI message is not the first NNPFC SEI message in decoding order with the particular nnpfc_id value in the current CLVS, then the value of nnpfc_formatting_and_purpose_flag shall be equal to 0.

[0108] nnpfc_purpose indicates the purpose of the postfilter processing. The value of nnpfc_purpose is 1 for image quality improvement, 2 for chrominance upsampling from 4:2:0 chrominance format to 4:2:2 or 4:4:4, or chrominance upsampling from 4:2:2 chrominance format to 4:4:4, 3 for increasing the width or height of the cropped decoded output image without changing the chrominance format, 4 for increasing the width or height of the decoded output image and upsampling the chrominance format, and 5 for picture rate upsampling.

[0109] The problem with non-patent document 2 is that when the value of nnpfc_purpose is 4, the value of the syntax element nnpfc_out_sub_c_flag can be read, but the values ​​of the necessary syntax elements nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples cannot be read.

[0110] Therefore, in this embodiment, the syntax is changed so that, as shown in Figure 8, "else if( nnpfc_purpose == 3 || nnpfc_purpose == 4 ) {" is changed to "if( nnpfc_purpose == 3 || nnpfc_purpose == 4 ) {" so that when the value of nnpfc_purpose is 4, not only the syntax element nnpfc_out_sub_c_flag but also the values ​​of the syntax elements nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples can be read.

[0111] Another solution is to not use "else if" for nnpfc_purpose, and in addition to the above modifications, you can change "else if( nnpfc_purpose == 5 ) {" to "if( nnpfc_purpose == 5 ) {".

[0112] nnpfc_inp_format_idc indicates how to convert pixel values ​​input to the neural network for post-filtering. When nnpfc_inp_format_idc is 0, it indicates that the input values ​​are real numbers, and when nnpfc_inp_format_idc is 1, it indicates that the input values ​​are unsigned integers.

[0113] If the value of nnpfc_inp_sample_idc is 0, the functions InpY and InpC are derived as follows:

[0114] InpY(x) = x÷((1< <BitDepthY)-1) InpC(x) = x ÷ ((1< <BitDepthC)-1) If the value of nnpfc_inp_sample_idc is 1, the functions InpY and InpC are derived as follows:

[0115] shift=BitDepthY-inpTensorBitDepth if(inpTensorBitDepth>=BitDepthY) InpY(x)=x<<(inpTensorBitDepth-BitDepthY) else InpY(x)=Clip3(0,(1< <inpTensorBitDepth)-1,(x+(1<<(shift-1)))> >shift) shift=BitDepthC-inpTensorBitDepth if(inpTensorBitDepth>=BitDepthC) InpC(x)=x<<(inpTensorBitDepth-BitDepthC) else InpC(x) = Clip3(0,(1< <inpTensorBitDepth)-1,(x+(1<<(shift-1)))> >shift) If nnpfc_inp_format_idc is 1, the syntax element nnpfc_inp_tensor_bitdepth_minus8 is encoded and decoded. The value obtained by adding 8 to nnpfc_inp_tensor_bitdepth_minus8 is the bit length inpTensorBitDepth when the pixel values ​​input to the neural network for input post-filtering are unsigned integers, and is derived using the following formula:

[0116] inpTensorBitDepth = nnpfc_inp_tensor_bitdepth_minus8 + 8 The problem with Non-Patent Document 2 is that when nnpfc_inp_format_idc is 1, if a deblocking filtering strength control value StrengthControlVal that is a real value between 0 and 1 exists, a method for converting the value of StrengthControlVal into an unsigned integer is not defined.

[0117] Therefore, the variable StrengthControlVal is redefined as follows:

[0118] StrengthControlVal=floor(StrengthControlVal*((1 << inpTensorBitDepth)-1)) By performing such a derivation, even when nnpfc_inp_format_idc is 1, post-filter processing using a neural network with the deblocking filtering strength control value being an unsigned integer value can be realized.

[0119] As another solution, the strength value StrengthControlValInDecoder obtained from the image decoding device 31 may be obtained to perform branching as follows.

[0120] if (nnpfc_inp_format_idc == 0) StrengthControlVal= StrengthControlValInDecoder else if (nnpfc_inp_format_idc == 1) StrengthControlVal =floor(StrengthControlValInDecoder*((1 << inpTensorBitDepth)-1) Note that, in the image decoding device 31, StrengthControlVal (StrengthControlValInDecoder) may be set by normalizing the value of the quantization parameter obtained from the encoded data to a decimal number between 0 and 1, as in the following value.

[0121] StrengthControlVal = SliceQpY of the first slice of the target picture ÷ NormQP Furthermore, you may explicitly clip the value to 0..1.

[0122] StrengthControlVal = Clip3(0.0, 1.0, SliceQpY of the first slice of the target picture ÷ NormQP) Note that SliceQpY is the value of the quantization parameter at the beginning of a slice decoded from encoded data, and if it is a negative value, it is set to 0. NormQP is a value for normalization, and is the maximum value of the quantization parameter or the maximum value plus 1. For example, in the case of Non-Patent Document 1, the maximum value of the quantization parameter is 63, so the value of NormQP is set to 63 or 64. In the cases of H.264 / AVC and H.265 / HEVC, the maximum value of the quantization parameter is 51, so the value of NormQP may be set to 51 or 52.

[0123] In another embodiment, the maximum value is assumed to be the pixel bit length of the luminance, StrengthControlValInDecoder = SliceQpY, and then the function InpY is used to StrengthControlVal = InpY(StrengthControlValInDecoder) It is also possible to use the following.

[0124] nnpfc_reserved_zero_bit_b shall be 0, and nnpfc_payload_byte[i] shall be the i-th byte of the bitstream encoded in NNC in accordance with ISO / IEC 15938-17.

[0125] (Neural Network Post Filter Activation (NNPFA) SEI) 8 shows the syntax of the Neural Network Post-Filter Activation (NNPFA) SEI message of Non-Patent Document 2. The Neural Network Post-Filter Activation NNPFA SEI message activates or deactivates the application of a target neural network post-filtering identified by nnpfa_target_id for post-filtering of a sequence of pictures.

[0126] nnpfa_target_id indicates the target picture neural network postfiltering. It identifies one or more NNPFC SEI messages with nnpfc_id equal to nnfpa_target_id for the current picture.

[0127] An NNPFA SEI message with a particular value of nnpfa_target_id should not be present on the current PU unless one or both of the following conditions are true: ·There is an NNPFC SEI message in the current CLVS with nnpfc_id equal to a specific value of nnpfa_target_id present in a PU preceding the current PU in decoding order. · There is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id for the current PU.

[0128] If a PU contains both an NNPFC SEI message with a particular value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a particular value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order.

[0129] nnpfa_cancel_flag, when set to 1, indicates that the continuity of the target neural network postfiltering set by the previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is cancelled, i.e. the target neural network postfiltering is not performed.

[0130] If it has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag is 0, it will not be used unless activated by another NNPFA SEI message. If nnpfa_cancel_flag is 0, it indicates that nnpfa_persistence_flag will persist.

[0131] nnnpfa_persistence_flag indicates that the target neural network postfilter processing of the current layer should persist in display output order.

[0132] If nnpfa_persistence_flag is 0, it indicates that the target neural network postfiltering applies to postfiltering of the current picture only.

[0133] nnpfa_persistence_flag, when set to 1, indicates that the target neural network postfiltering should be applied to the current image and all subsequent pictures of the current layer until one or more of the following conditions become true: -New CLVS for the current layer is started Bitstream ends · Images in the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 are output after the current image in the display output order.

[0134] Note that neural network post-filtering is not applied to subsequent pictures of the current layer associated with an NNPFA SEI message if the nnpfa_target_id and nnpfa_cancel_flag are the same as those of the current SEI message and are 1.

[0135] In Non-Patent Document 2, by setting nnpfa_persistence_flag to 1, it is possible to continue postfiltering for subsequent pictures in display output order without sending NNPFA SEI. However, when the encoding and decoding order is different from the display output order, for example, when a prediction structure such as a so-called hierarchical B structure (Hierarchical Bi-prediction Structure) is adopted, it is difficult for the auxiliary enhancement information generating device 71 to generate auxiliary enhancement information of the NNPFA SEI.

[0136] When generating the auxiliary extension information of the NNPFA SEI, it is usually considered that whether to perform postfilter processing and which neural network model to use is determined in the coding order (= decoding order), but it becomes necessary to generate the NNPFA SEI so that it is possible to determine whether the postfilter processing by the same neural network model is continuous in the display output order. In this case, if it is not known in advance whether to perform postfilter processing on a picture and which model to use, it is not possible to determine at the timing of encoding on a picture-by-picture basis whether the postfilter processing by the same neural network model is continuous in the display output order.

[0137] In such a case, it is necessary to set nnpfa_persistence_flag to 0 and send NNPFA SEI for all pictures for which postfilter processing is activated (always transmit), or to delay until the display output order is determined, and then create auxiliary enhancement information, encode the auxiliary enhancement information, and output the encoded data.

[0138] Specifically, FIG. 10 shows an example of a hierarchical B structure as an example of a prediction structure in which the display output order and the encoding / decoding order are different. Picture P at No. 8 in the display output order is No. 1 in the encoding / decoding order. Meanwhile, picture B at No. 7 in the display output order is No. 8 in the encoding / decoding order. In the case of postfiltering to improve image quality, whether or not to apply postfiltering is determined on a picture-by-picture basis immediately after encoding. However, when postfiltering is activated for picture P at No. 8 in the display output order, picture B immediately before it at No. 7 in the display output order is No. 8 in the encoding / decoding order, so that encoding / decoding processing has not been executed and it is not known whether postfiltering is continuing or not.

[0139] Therefore, in this embodiment, in the syntax of the NNPA SEI, as shown in Fig. 11, when the value of nnpfa_persistence_flag is 1, the syntax element decoding_order_flag is coded and decoded. When decoding_order_flag is 1, it indicates that postfilter processing of pictures continues in the decoding order (coding order). When decoding_order_flag is 0, it indicates that postfilter processing of pictures continues in the display output order.

[0140] With this configuration, if the encoding / decoding order and the display output order differ, it is possible to create auxiliary enhancement information by setting decoding_order_flag to 1, which indicates that post-filter processing will continue in the encoding / decoding order.

[0141] Another solution is to use output_order_flag instead of decoding_order_flag, where output_order_flag is set to 1 to indicate that postfiltering of pictures continues in display output order, and output_order_flag is set to 0 to indicate that postfiltering of pictures continues in decoding order (coding order).

[0142] Alternatively, a syntax element called nnpfa_presistence_idc can be defined instead of nnpfa_presistence_flag, with the descriptor being ue(v), and a value of 0 indicates that the target neural network postfiltering is applied to the current picture only. A value of 1 for nnpfa_presistence_idc indicates that postfiltering of pictures continues in display output order, and a value of 2 for nnpfa_presistence_idc indicates that postfiltering of pictures continues in decoding order (encoding order).

[0143] Another solution is to describe the decoding_order_flag and output_order_flag in the NNPFC SEI. In this case, the NNPFA SEI is not changed and the syntax in Figure 9 is used.

[0144] According to this embodiment, even if the encoding / decoding order and the display output order are different, it is possible to efficiently realize activation of post-filtering on a picture-by-picture basis.

[0145] Furthermore, the auxiliary extension information creating device 71, the auxiliary extension information encoding device 81, and the auxiliary extension information decoding device 91 may commonly hold generic network parameters. The auxiliary extension information creating device 71 uses a framework such as a neural network postfilter characteristic SEI to create network parameters for partially updating the commonly held generic network as auxiliary extension information. The auxiliary extension information may then be encoded by the auxiliary extension information encoding device 81 and decoded by the auxiliary extension information decoding device 91. With this configuration, the amount of code for the auxiliary extension information can be reduced, and auxiliary extension information corresponding to the input image T can be created, encoded, and decoded.

[0146] In addition, in order to support multiple formats as the transmission format of the network parameters, a parameter (identifier) ​​indicating the format may be sent. Furthermore, the actual supplementary extension information following the identifier may be transmitted as a byte string.

[0147] The auxiliary extension information of the network parameters decoded by the auxiliary extension information decoder 91 is input to the postfilter processor 61 .

[0148] The postfilter processor 61 performs post-image processing using a neural network by using the decoded auxiliary enhancement information (neural network postfilter characteristics SEI, neural network postfilter activation SEI) to restore the decoded video Td.

[0149] The auxiliary extension information encoding device 81 encodes the auxiliary extension information based on the syntax tables of Figures 7, 8, 9, and 11. The auxiliary extension information is encoded as auxiliary extension information SEI, multiplexed into the encoded data Te output by the image encoding device 11, and output to the network 21.

[0150] The auxiliary extension information decoding device 91 decodes the auxiliary extension information from the encoded data Te based on the syntax tables of Figures 7, 8, 9, and 11, and sends the decoded result to the post-filter processing device 61 and the image recognition device 51. The auxiliary extension information decoding device 91 decodes the auxiliary extension information encoded as auxiliary extension information SEI.

[0151] The post-filtering device 61 performs post-image processing on the decoded video image Td using the decoded video image Td and the auxiliary extension information to generate a post-image processing image To.

[0152] Furthermore, general-purpose network parameters may be held in common by the auxiliary extension information creating device 71, the auxiliary extension information encoding device 81, and the auxiliary extension information decoding device 91. In the auxiliary extension information creating device 71, network parameters for partially updating the commonly held general-purpose network may be created as auxiliary extension information, which may then be encoded by the auxiliary extension information encoding device 81 and decoded by the auxiliary extension information decoding device 91. With this configuration, the amount of code for the auxiliary extension information can be reduced, and auxiliary extension information corresponding to the input image T can be created, encoded, and decoded.

[0153] In addition, in order to support multiple formats as the transmission format of the network parameters, a parameter (identifier) ​​indicating the format may be sent. Furthermore, the actual supplementary extension information following the identifier may be transmitted as a byte string.

[0154] The auxiliary extension information of the network parameters decoded by the auxiliary extension information decoder 91 is input to the postfilter processor 61 .

[0155] In addition, in the example of this embodiment, the syntax in SEI is shown, but it is not limited to SEI, and syntax such as SPS, PPS, APS, slice header, etc. may be used.

[0156] In this embodiment, in such a configuration, a method is provided that can improve image quality even when the transmission rate is low, and can efficiently code and decode supplemental extension information regardless of the prediction structure.

[0157] In addition, a part of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 308, the inverse quantization and inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the predicted image generating unit 101, the subtraction unit 102, the transform and quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, the parameter coding unit 111, and the prediction parameter derivation unit 120 may be realized by a computer. In this case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into and executed by a computer system. In addition, the "computer system" referred to here is a computer system built into either the image encoding device 11 or the image decoding device 31, and includes hardware such as an OS and peripheral devices. In addition, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording medium" may also include devices that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and devices that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such cases. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0158] In addition, a part or the whole of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be individually made into a processor, or a part or the whole may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0159] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.

[0160] [Application example] The above-mentioned video encoding device 10 and video decoding device 30 can be mounted and used in various devices that transmit, receive, record, and play back videos. Note that the video may be a natural video captured by a camera or the like, or an artificial video (including CG and GUI) generated by a computer or the like.

[0161] (SEI payload) FIG. 12 is a diagram showing the syntax of the SEI payload, which is a container of the SEI message.

[0162] This is called when nal_unit_type is PREFIX_SEI_NUT. PREFIX_SEI_NUT indicates that the SEI is located before the slice data.

[0163] When payloadType is 210, the neural network postfilter feature SEI is called.

[0164] When payloadType is 211, the neural network post-filter activation SEI is called.

[0165] (SEI decoding and post-filtering) The header decoding unit 3020 reads the SEI payload, which is a container of the SEI message, and decodes the neural network post-filter characteristics SEI message. For example, the header decoding unit 3020 decodes nnpfc_id, nnpfc_mode_idc, nnpfc_formatting_and_purpose_flag, nnpfc_purpose, nnpfc_reserved_zero_bit_a, nnpfc_uri_tag[i], nnpfc_uri[i], nnpfc_reserved_zero_bit_b, and nnpfc_payload_byte[i].

[0166] 13 is a flowchart showing the processing of the post-filter processor 61. The post-filter processor 61 performs the following processing in accordance with the parameters of the SEI message.

[0167] S6001: Read processing volume and accuracy from SEI.

[0168] S6002: If the amount of post processing exceeds the processable complexity, the process ends. If not, the process proceeds to S6003.

[0169] S6003: End if the accuracy exceeds the processable accuracy of the post-filter processor 61. If not, proceed to S6004.

[0170] S6004: A network model is identified from the SEI, and the topology of the post-filter processing device 61 is set.

[0171] S6005: Derive network model parameters from the updated information of SEI.

[0172] S6006: The derived network model parameters are loaded into the post-filter processor 61.

[0173] S6007: The post-filter processor 61 executes filtering and outputs the result to the outside.

[0174] However, the SEI is not necessarily required for constructing luma samples and chroma samples in the decoding process. (Details of the post-filter processing unit 61) The NN filter unit performs filtering using a neural network model, using the input image inputTensor and input parameters (e.g., QP, bS, etc.). The input image may be an image for each component, or an image with multiple components as channels. In addition, the input parameters may be assigned to a channel different from the image.

[0175] The NN filter unit may repeatedly apply the following process.

[0176] The NN filter part derives the output image outputTensor by convolving the inputTensor with the kernel k[m][i][j] (conv, convolution) and adding bias. Here, nn=0..n-1, xx=0..width-1, yy=0..height-1, and Σ represents the summation for mm, i, and j, respectively.

[0177] outputTensor[nn][xx][yy]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][xx+i-of][yy+j-of]+bias[nn]) In the case of 1x1 Conv, Σ represents the sum of mm=0..m-1, i=0, j=0. In this case, of=0 is set. In the case of 3x3 Conv, Σ represents the sum of mm=0..m-1, i=0..2, j=0..2. In this case, of=1 is set. n is the number of channels of outSamples, m is the number of channels of inputTensor, width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor. of is the size of the padding area placed around inputTensor to make the sizes of inputTensor and outputTensor the same. In the following, when the output of the NN filter part is a value (correction value) rather than an image, the output is represented as corrNN instead of outputTensor.

[0178] Note that if you write inputTensor and outputTensor in CHW format instead of CWH format, it is equivalent to the following process.

[0179] outputTensor[nn][yy][xx]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][yy+j-of][xx+i-of]+bias[nn]) In addition, a process called Depth-wise Conv, shown in the following formula, may be performed. Here, nn=0..n-1, xx=0..width-1, yy=0..height-1, and Σ represents the summation for i and j, respectively. n is the number of channels of outputTensor and inputTensor, width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor.

[0180] outputTensor[nn][xx][yy]=ΣΣ(k[nn][i][j]*inputTensor[nn][xx+i-of][yy+j-of]+bias[nn]) Also, a nonlinear process called Activate, for example, ReLU, may be used. ReLU(x) = x >= 0 ? x : 0 Alternatively, leakyReLU shown in the following formula may be used.

[0181] leakyReLU(x) = x >= 0 ? x : a * x Here, a is a predetermined value, for example, 0.1 or 0.125. In order to perform integer arithmetic, all the values ​​of k, bias, and a above may be integers, and a right shift may be performed after conv.

[0182] With ReLU, 0 is always output for values ​​less than 0, and the input value is output as is for values ​​greater than or equal to 0. On the other hand, with leakyReLU, linear processing is performed for values ​​less than 0 with the gradient set by a. With ReLU, the gradient for values ​​less than 0 disappears, which can make it difficult for learning to progress. With leakyReLU, the gradient for values ​​less than 0 remains, making the above problem less likely to occur. Of the above leakyReLU(x), PReLU, which uses a parameterized value of a, may also be used.

[0183] (NNC) Neural Network Coding (NNC) is an international standard ISO / IEC15938-17 for efficiently compressing neural networks (NNs). Compressing trained NNs makes it possible to store and transmit NNs more efficiently.

[0184] The following provides an overview of the encoding and decoding processes of the NNC.

[0185] FIG. 14 is a diagram showing an encoding device and a decoding device of the NNC.

[0186] The NN coding device 801 has a pre-processing unit 8011, a quantization unit 8012, and an entropy coding unit 8013. The NN coding device 801 receives an uncompressed NN model O, and the quantization unit 8012 quantizes the NN model O to obtain a quantized model Q. The NN coding device 801 may repeatedly apply a parameter reduction method such as pruning or sparsification before quantization in the pre-processing unit 8011. Thereafter, the entropy coding unit 8013 applies entropy coding to the quantized model Q to obtain a bit stream S for storing and transmitting the NN model.

[0187] The NN decoding device 802 includes an entropy decoding unit 8021, a parameter restoration unit 8022, and a post-processing unit 8023. The NN decoding device 802 first inputs the transmitted bit stream S, and the entropy decoding unit 8021 performs entropy decoding of S to obtain an intermediate model RQ. If the operating environment of the NN model supports inference using the quantized representation used in RQ, the RQ may be output and used for inference. If not, the parameter restoration unit 8022 restores the parameters of RQ to their original representation to obtain an intermediate model RP. If the sparse tensor representation used can be processed in the operating environment of the NN model, the RP may be output and used for inference. If not, a reconstructed NN model R that does not include a tensor or structural representation different from the NN model O is obtained and output.

[0188] The NNC standard has decoding methods for certain numeric representations of NN parameters, including integers and floating point.

[0189] The decoding method NNR_PT_INT decodes a model with integer-valued parameters. The decoding method NNR_PT_FLOAT extends NNR_PT_INT by adding a quantization step size delta, which is multiplied by the integer value to produce a scaled integer. Delta is derived from the integer quantization parameter qp and the granularity parameter qp_density of delta as follows:

[0190] mul = 2^(qp_density) + (qp & (2^(qp_density)-1)) delta = mul * 2^((qp >> qp_density)-qp_density) (Format of trained NN) The representation of a trained NN consists of two elements: a topological representation, such as the size of layers and the connections between layers, and a parameter representation, such as weights and biases.

[0191] Topology representation is covered by native formats such as Tensorflow and PyTorch, but to improve interoperability, exchange formats such as Open Neural Network Exchange Format (ONNX) and Neural Network Exchange Format (NNEF) exist.

[0192] In addition, the NNC standard transmits topology information nnr_topology_unit_payload as part of the NNC bitstream containing compressed parameter tensors. This allows interoperability with topology information expressed in native format as well as exchange formats. (Image Encoding Device Configuration) Next, the configuration of the image encoding device 11 according to this embodiment will be described. Fig. 6 is a block diagram showing the configuration of the image encoding device 11 according to this embodiment. The image encoding device 11 includes a prediction image generating unit 101, a subtraction unit 102, a transformation and quantization unit 103, an inverse quantization and inverse transformation unit 105, an addition unit 106, a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determining unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.

[0193] The predicted image generating unit 101 generates a predicted image for each CU.

[0194] The subtraction unit 102 generates a prediction error by subtracting the pixel values ​​of the predicted image of the block input from the predicted image generation unit 101 from the pixel values ​​of the image T. The subtraction unit 102 outputs the prediction error to the transformation and quantization unit 103.

[0195] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction errors input from the subtraction unit 102, and derives quantized transform coefficients by quantizing the prediction errors. The transform / quantization unit 103 outputs the quantized transform coefficients to the parameter coding unit 111 and the inverse quantization / inverse transform unit 105.

[0196] The inverse quantization and inverse transform unit 105 is the same as the inverse quantization and inverse transform unit 311 (FIG. 6) in the image decoding device 31, and a description thereof will be omitted. The calculated prediction error is output to the addition unit .

[0197] The parameter coding unit 111 includes a header coding unit 1110, a CT information coding unit 1111, and a CU coding unit 1112 (prediction mode coding unit). The CU coding unit 1112 further includes a TU coding unit 1114. The following describes an outline of the operation of each module.

[0198] The header encoding unit 1110 performs encoding processing of parameters such as header information, division information, prediction information, and quantized transform coefficients.

[0199] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) division information and the like.

[0200] The CU encoding unit 1112 encodes the CU information, prediction information, division information, and so on.

[0201] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized prediction error.

[0202] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters and quantized transform coefficients to the parameter encoding unit 111.

[0203] The entropy coding unit 104 receives the quantized transform coefficients and the coding parameters from the parameter coding unit 111. The entropy coding unit 104 entropy codes these to generate and output coded data Te.

[0204] The prediction parameter derivation unit 120 derives inter prediction parameters and intra prediction parameters from the parameters input from the encoding parameter determination unit 110. The derived inter prediction parameters and intra prediction parameters are output to the parameter encoding unit 111.

[0205] The adder 106 generates a decoded image by adding, for each pixel, the pixel value of the predicted block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in a reference picture memory 109.

[0206] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily have to include the above three types of filters, and may be configured, for example, as only a deblocking filter.

[0207] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.

[0208] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined position for each current picture and CU.

[0209] The coding parameter determination unit 110 selects one set from among a plurality of sets of coding parameters. The coding parameters are the above-mentioned QT, BT or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generation unit 101 generates a predicted image using these coding parameters.

[0210] In addition, a part of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 308, the inverse quantization and inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the predicted image generating unit 101, the subtraction unit 102, the transform and quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, the parameter coding unit 111, and the prediction parameter derivation unit 120 may be realized by a computer. In this case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into and executed by a computer system. In addition, the "computer system" referred to here is a computer system built into either the image encoding device 11 or the image decoding device 31, and includes hardware such as an OS and peripheral devices. In addition, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording medium" may also include devices that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and devices that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such cases. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0211] In addition, a part or the whole of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be individually made into a processor, or a part or the whole may be integrated into a processor. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0212] Although an embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes and the like are possible within the scope of the gist of the present invention. Explaining this embodiment based on FIG. 1, a video encoding device 10 is characterized by having an image encoding device 11 that encodes an input image, an auxiliary enhancement information generating device 71 that generates auxiliary enhancement information indicating whether or not to perform postfilter processing on a picture-by-picture basis, and an auxiliary enhancement information encoding device 81 that encodes information indicating whether to continue postfilter processing in the decoding order or in the display output order when determining whether to continue postfilter processing on a picture-by-picture basis in the auxiliary enhancement information. In addition, the auxiliary enhancement information encoding device 81 is characterized by encoding information that specifies a neural network to be applied to postfilter processing.

[0213] The video decoding device 30 is characterized by including an image decoding device 31 that decodes images from encoded data, a postfilter processing device 61 that performs postfilter processing on the decoded images decoded by the image decoding device 31, and an auxiliary extension information decoding device 91 that decodes auxiliary extension information indicating whether or not to perform postfilter processing on a picture-by-picture basis in the postfilter processing device 61, and decoding information indicating whether to continue postfilter processing in the decoding order or in the display output order when determining whether to continue postfilter processing on a picture-by-picture basis in the auxiliary extension information. Also, the auxiliary extension information decoding device 91 is characterized by decoding information that specifies a neural network to be applied to the postfilter processing.

[0214] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]

[0215] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data in which image data is coded, and a video coding device that generates coded data in which image data is coded, and can also be suitably applied to the data structure of coded data that is generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]

[0216] 1. Video transmission system 30 Video Decoding Device 31 Image Decoding Device 301 Entropy Decoding Unit 302 Parameter Decoding Unit 305, 107 Loop Filter 306, 109 Reference Picture Memory 307, 108 Prediction parameter memory 308, 101 Prediction image generation unit 311, 105 Inverse quantization and inverse transformation unit 312, 106 Addition section 320 Prediction Parameter Derivation Unit 10 Video Encoding Device 11 Image encoding device 102 Subtraction section 103 Transformation and Quantization Section 104 Entropy coding unit 110 Encoding parameter determination unit 111 Parameter Encoding Unit 120 Prediction parameter derivation part 41 Image display device 51 Pre-filter processing device 61 Post-filter processing device 71 Auxiliary extension information creation device 81 Supplementary extension information coding device 91 Auxiliary extension information decoding device

Claims

1. A neural network post-filter characteristic SEI message includes an auxiliary extension information decoding unit that decodes a first syntax element and a second syntax element from the message, and derives a first variable using an intensity control value according to the value of the first syntax element and the value of the second syntax element; a first variable being set to an integer value obtained by multiplying the intensity control value by a value obtained by shifting 1 to the left by the value of a chrominance bit depth and then subtracting 1 from the value, when the value of the first syntax element is 1 and the value of the second syntax element is 1.

2. A video decoding device as described in Claim 1, characterized in that when the value of the first syntax element is 1 and the value of the second syntax element is 1, this indicates that the input tensor contains two color difference matrices and one auxiliary input matrix.

3. The first syntax element indicates whether auxiliary input data is present in an input tensor of a neural network postfilter; 2. The video decoding device according to claim 1, wherein the second syntax element is in the form of an input tensor and indicates a method for ordering pixels in a decoded image.

4. A neural network post-filter characteristic SEI message is encoded into a first syntax element and a second syntax element, and an auxiliary extension information encoding unit is configured to derive a first variable using an intensity control value according to a value of the first syntax element and a value of the second syntax element; a first syntax element that is a chrominance bit depth value, a second syntax element that is a chrominance bit depth value, ...

5. A method for transmitting a bitstream encoded by a video encoding device, comprising: transmitting the bitstream; The bitstream comprises: a neural network post-filter characteristics SEI message including a first syntax element and a second syntax element; A transmission method characterized in that, when the value of the first syntax element is 1 and the value of the second syntax element is 1, the first variable is set to an integer value obtained by multiplying the intensity control value by a value obtained by shifting 1 to the left by the value of the chrominance bit depth and then subtracting 1 from the value.