Video encoding device and video decoding device

JP2024002451A5Inactive Publication Date: 2025-06-30SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022101630
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-06-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Benefits of technology

【0015】 このような構成にすることで、ニューラルネットワークの処理を効率よくかつ正確に実行することが可能となる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To solve a problem in which, in the conventional neural network post-filter characteristic SEI specification, a patch size larger than the picture size of a decoded image can be defined.SOLUTION: A video decoding device according to an embodiment of the present invention includes an image decoding device that decodes encoded data to generate a decoded image, and a resolution inverse conversion device using a neural network that converts the decoded image to a specified resolution using inverse conversion information. In the resolution inverse conversion device, information indicating the number of horizontal pixels and vertical pixels of the patch size, which is the processing unit of the neural network is decoded, the maximum values of the number of horizontal pixels and the number of vertical pixels of the patch size are set as the number of horizontal pixels and the number of vertical pixels of the decoded image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] An embodiment of the present invention relates to a video encoding device and a video decoding device. [Background technology]

[0002] In order to efficiently transmit or record moving images, a moving image encoding device is used that generates encoded data by encoding moving images, and a moving image decoding device is used that generates a decoded image by decoding the encoded data.

[0003] Specific examples of video encoding methods include H.264 / AVC and H.265 / HEVC (High-Efficiency Video Coding).

[0004] In such a video coding method, images (pictures) constituting a video are divided into slices obtained by dividing the images, coding tree units (CTUs) obtained by dividing the slices, coding tree units (CTUs) obtained by dividing the coding tree units, and so on. The coding unit (sometimes called a coding unit (CU)) that is to be encoded, and The coding unit is divided into transform units (TUs), which are managed in a hierarchical structure, and the coding unit is encoded / decoded for each CU.

[0005] In such a video coding method, a predicted image is usually generated based on a locally decoded image obtained by encoding / decoding an input image, and a prediction error (sometimes called a "difference image" or "residual image") obtained by subtracting the predicted image from the input image (original image) is coded. Methods for generating a predicted image include inter-prediction and intra-prediction.

[0006] H.274 includes a number of features that allow the transmission of image characteristics, display methods, and timing information simultaneously with the encoded data. A Supplemental Enhancement Information (SEI) message is defined for this purpose.

[0007] Non-Patent Document 1 discloses a method of explicitly defining an SEI that transmits the topology and parameters of a neural network filter used as a postfilter, and a method of indirectly defining the SEI as reference information. [Prior art documents] [Non-patent literature]

[0008] [Non-Patent Document 1] S. McCarthy, T. Chujoh, MM Hannuksela, GJ Sullivan and Y.-K. Wang, "Additional SEI messages for VSEI (Draft 1)," JVET-Z2006, June 21, 2022. Summary of the Invention [Problem to be solved by the invention]

[0009] However, the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1 is In this case, there was a problem that a patch size larger than the picture size of the decoded image could be defined.

[0010] In addition, in the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, There was an issue that the variable indicating the maximum value of the network parameter could overflow with a 32-bit integer or 64-bit integer value. [Means for solving the problem]

[0011] A video decoding device according to one aspect of the present invention includes: an image decoding device that decodes encoded data to generate a decoded image; a resolution inverse conversion device using a neural network for converting the decoded image into a specified resolution using inverse conversion information; In the resolution inverse conversion device, information indicating the number of horizontal pixels and the number of vertical pixels of a patch size, which is a unit of processing of a neural network, is decoded; The maximum values ​​of the number of horizontal pixels and the number of vertical pixels of the patch size are set to the number of horizontal pixels and the number of vertical pixels of the decoded image.

[0012] Moreover, a video decoding device according to one aspect of the present invention includes: an image decoding device that decodes encoded data to generate a decoded image; a resolution inverse conversion device using a neural network for converting the decoded image into a specified resolution using inverse conversion information; Decoding information indicating a maximum value of a network parameter of the neural network in the resolution inverse conversion device; When deriving the maximum value of the network parameter, the maximum value is set to a value that does not overflow in 32-bit integer arithmetic.

[0013] Furthermore, a video encoding device according to an aspect of the present invention comprises: an image encoding device that encodes an image to generate encoded data; an inverse conversion information generating device for generating inverse conversion information for inversely converting a resolution of a decoded image when the encoded data is decoded; An inverse transformation information encoding device that encodes the inverse transformation information as supplemental extension information, In the resolution inverse conversion information encoding device, information indicating the number of horizontal pixels and the number of vertical pixels of a patch size, which is a unit of processing of a neural network, is encoded; The maximum values ​​of the number of horizontal pixels and the number of vertical pixels of the patch size are set to the number of horizontal pixels and the number of vertical pixels of the encoded image.

[0014] Furthermore, a video encoding device according to an aspect of the present invention comprises: an image encoding device that encodes an image to generate encoded data; an inverse conversion information generating device for generating inverse conversion information for inversely converting a resolution of a decoded image when the encoded data is decoded; An inverse transformation information encoding device that encodes the inverse transformation information as supplemental extension information, encoding information indicating a maximum value of a network parameter of a neural network in the resolution inverse conversion information encoding device; When deriving the maximum value of the network parameter, the maximum value is set to a value that does not overflow in 64-bit integer arithmetic. Effect of the Invention

[0015] With this configuration, it becomes possible to execute neural network processing efficiently and accurately. [Brief description of the drawings]

[0016] [Figure 1] 1 is a schematic diagram showing a configuration of a video transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of encoded data. [Diagram 3] 1 is a conceptual diagram of an image to be processed in the video transmission system according to the present embodiment. [Figure 4] FIG. 2 is a conceptual diagram showing an example of a reference picture and a reference picture list. [Diagram 5] FIG. 1 is a schematic diagram showing a configuration of an image decoding device. [Figure 6] 11 is a flowchart illustrating a schematic operation of an image decoding device. [Figure 7] FIG. 1 is a block diagram showing a configuration of an image encoding device. [Figure 8] 13 is a schematic diagram showing a configuration of an inter-prediction parameter encoding unit. [Figure 9] FIG. 13 is a diagram showing the neural network postfilter characteristic SEI syntax. [Figure 10] FIG. 1 illustrates the syntax of a neural network complexity element. [Figure 11] FIG. 1 is a diagram showing the formula 1 of the function Reflect(y,z). [Figure 12] FIG. 13 illustrates neural network post-filter activation SEI syntax. [Figure 13] A diagram showing the syntax of an SEI payload, which is a container for an SEI message. [Figure 14] FIG. 13 is a flowchart showing the process of the NN filter unit 611. [Figure 15] FIG. 13 is a diagram showing the configuration of a neural network of an NN filter unit 611. [Figure 16] FIG. 1 is a diagram showing an NNR encoding device and a decoding device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0017] (First embodiment) Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0018] FIG. 1 is a schematic diagram showing the configuration of a video transmission system according to this embodiment.

[0019] The video transmission system 1 encodes images having different resolutions, the resolutions of which have been converted. The video transmission system 1 comprises a video encoding device 10, a network 21, and a video It comprises an image decoding device 30 and an image display device 41 .

[0020] The video encoding device 10 is made up of a resolution conversion device (resolution conversion unit) 51, an image encoding device (image encoding unit) 11, an inverse conversion information creation device (inverse conversion information creation unit) 71, and an inverse conversion information encoding device (inverse conversion information encoding unit) 81.

[0021] The video decoding device 30 includes an image decoding device (image decoding unit) 31, a resolution inverse conversion device (resolution inverse conversion unit) 61, and an inverse conversion information decoding device (inverse conversion information decoding unit) 91.

[0022] The resolution conversion device 51 converts the resolution of an image T included in a video to generate an image of a different resolution. to the image coding device 11. The resolution conversion device 51 also supplies inverse conversion information indicating whether or not the resolution of the image is converted to the image coding device 11. If the information indicates resolution conversion, the image coding device 11 sets resolution conversion information ref_pic_resampling_enabled_flag, which will be described later, to 1, and includes it in a sequence parameter set SPS (Sequence Parameter Set) of the coded data Te for coding.

[0023] The inverse conversion information generating device 71 generates inverse conversion information based on an image T1 included in a video. The inverse conversion information is derived or selected from the relationship between an input image T1 before resolution conversion and an image Td1 after resolution conversion, encoding, and decoding.

[0024] Inverse transformation information is input to the inverse transformation information encoding device 81. The inverse transformation information encoding device 81 encodes the inverse transformation information to generate encoded inverse transformation information, and transmits the encoded inverse transformation information to the network 21.

[0025] A variable resolution image T2 is input to the image encoding device 11. The image encoding device 11 encodes image size information of the input image in PPS units using the framework of RPR (Reference Picture Resampling), and transmits the encoded image size information to the image decoding device 31.

[0026] In FIG. 1, the inverse transformation information coding device 81 is not connected to the image coding device 11, but the inverse transformation information coding device 81 and the image coding device 11 may communicate necessary information as appropriate.

[0027] The network 21 transmits the coded inverse transformation information and the coded data Te to the image decoding device 31. A part or all of the coded inverse transformation information is coded as supplemental extension information SEI. The encoded data Te may be included in the data Te. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a two-way communication network, and may be a one-way communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. The network 21 may also be replaced by a storage medium on which the encoded data Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blue-ray Disc: registered trademark).

[0028] The image decoding device 31 decodes the coded data Te transmitted by the network 21 to generate a variable resolution decoded image Td 1 and supplies it to the resolution inverse conversion device 61 .

[0029] The inverse conversion information decoding device 91 decodes the coded inverse conversion information transmitted by the network 21 to generate inverse conversion information, and supplies the generated inverse conversion information to the resolution inverse conversion device 61 .

[0030] 1, the inverse transformation information decoding device 91 is illustrated separately from the image decoding device 31, but the inverse transformation information decoding device 91 may be included in the image decoding device 31. For example, the inverse transformation information decoding device 91 may be included in the image decoding device 31 separately from each functional unit of the image decoding device 31. Also, although not connected to the image decoding device 31 in FIG. 1, the inverse transformation information decoding device 91 and the image decoding device 31 may communicate necessary information as appropriate.

[0031] When the resolution conversion information indicates resolution conversion, the resolution inverse conversion device 61 inversely converts the decoded image of the image decoding device 31 into a resolution-converted image based on the image size information included in the encoded data and the inverse conversion information. Methods for inversely converting a resolution-converted image include post-filter processing such as super-resolution processing using a neural network.

[0032] Furthermore, when the resolution conversion information indicates a normal resolution, the resolution inverse conversion device 61 may perform post-filtering using a neural network, execute a resolution inverse conversion process to restore the input image T1, and generate a decoded image Td2.

[0033] The image display device 41 receives one or more decoded images Td2 from the resolution inverse conversion device 61. The image display device 41 displays all or part of the image. The image display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. The display may be in the form of a stationary display, a mobile display, an HMD, or the like. When the image decoding device 31 has high processing power, it displays high quality images, and when it has only low processing power, it displays images that do not require high processing power or display power.

[0034] FIG. 3 is a conceptual diagram of an image to be processed in the video transmission system shown in FIG. FIG. 3 is a diagram showing the change in resolution of the image over time. The system does not distinguish whether the image is coded or not. In the process, the resolution of the image is reduced and the image is transmitted to the image decoding device 31. As shown in FIG. 3, the resolution conversion device 51 usually reduces the amount of information to be transmitted. To do this, a conversion is performed to make the image resolution the same as or lower than the resolution of the input image.

[0035] <operator> The operators used in this specification are listed below.

[0036] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, and | is a bitwise OR , |= is an OR assignment operator, and || indicates a logical OR.

[0037] x? y : z is a ternary operator that takes y when x is true (non-zero) and z when x is false (0). is.

[0038] Clip3(a, b, c) is a function that clips c to a value between a and b, inclusive. It returns a if c < a, b if c > b, and c otherwise (assuming a <= b). That is, it returns a if c < a, b if c > b, and c otherwise (where a <= b).

[0039] abs(a) is a function that returns the absolute value of a.

[0040] Int(a) is a function that returns the integer value of a.

[0041] floor(a) is a function that returns the largest integer less than or equal to a.

[0042] ceil(a) is a function that returns the smallest integer greater than or equal to a.

[0043] a / d represents the division of a by d, rounded down to the nearest integer.

[0044] a ^ b represents power(a, b). When a = 2, it is equal to 1 << b.

[0045] <Structure of the encoded data Te>[[]]END]] Prior to the detailed description of the image encoding device 11 and the image decoding device 31 according to this embodiment, the data structure of the encoded data Te generated by the image encoding device 11 and decoded by the image decoding device 31 will be described.

[0046] FIG. 2 is a diagram showing the hierarchical structure of the data in the encoded data Te. The encoded data Te is , exemplarily, includes a sequence and a number of pictures that make up the sequence. a coded video sequence defining a sequence SEQ; a coded picture defining a picture PICT; A slice definition for slice S, an encoding slice definition for slice data, A diagram showing slice data, coding tree units contained in the coded slice data, and coding units contained in the coding tree units is shown.

[0047] (Coded Video Sequence) In a coded video sequence, an image decoding device is used to decode the sequence SEQ to be processed. The standard specifies a set of data referenced by the video parameter set (VPS) 31. As shown in Fig. 2, the sequence SEQ includes a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture PICT, and supplemental enhancement information (SEI).

[0048] In the video parameter set VPS, for a video image that consists of multiple layers, A set of coding parameters common to a plurality of video streams and a set of coding parameters associated with a plurality of layers and each individual layer contained in the video stream are defined.

[0049] The sequence parameter set SPS specifies a set of coding parameters that the image decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are Note that multiple SPSs may exist. In that case, one of the multiple SPSs is selected from the PPS. Select.

[0050] Here, the sequence parameter set SPS includes the following syntax elements: ref_pic_resampling_enabled_flag: A flag that specifies whether or not to use a function for varying resolution (resampling) when decoding each image included in a single sequence that references the target SPS. In other words, this flag indicates that the size of the reference picture referenced in generating a predicted image changes between each image indicated by a single sequence. When the value of this flag is 1, the resampling function is enabled. is applied, if it is 0 it is not applied. pic_width_max_in_luma_samples: The maximum width of any image in a sequence. This syntax element specifies the width of the image in luminance block units. The value of this syntax element must not be 0 and must be an integer multiple of Max(8, MinCbSizeY). Here, MinCbSizeY is a value determined by the minimum size of a luminance block. pic_height_max_in_luma_samples: This syntax element specifies the height of the image with the maximum height among the images in a single sequence, in units of luminance blocks. The value of this syntax element must not be 0 and must be an integer multiple of Max(8, MinCbSizeY). is required.

[0051] The picture parameter set PPS specifies the number of pictures to be decoded for each picture in the target sequence. A set of coding parameters to be referred to by the image decoding device 31 is specified. If present, one of multiple PPSs may be selected for each picture in the target sequence. Select.

[0052] Here, the picture parameter set PPS includes the following syntax elements: pps_pic_width_in_luma_samples: Syntax element that specifies the width of the target picture. The value of this syntax element is not 0, is an integer multiple of Max(8, MinCbSizeY), and The value must be less than or equal to sps_pic_width_max_in_luma_samples. pps_pic_height_in_luma_samples: A syntax element that specifies the height of the target picture. The value of this syntax element is not 0, but an integer multiple of Max(8, MinCbSizeY). It is also required that the value be less than or equal to sps_pic_height_max_in_luma_samples. pps_conformance_window_flag: Conformance (cropping) window offset This flag indicates whether the conformance parameter is to be notified next and indicates where to display the conformance window. If this flag is 1, the parameter is signaled and is 0, the conformance window offset parameter is not present. Indicates that there is no pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, pps_conf_win_bottom_offset: offset values ​​for specifying the left, right, top, and bottom positions of a picture output by decoding, with respect to a rectangular area specified by the output picture coordinates. If the value of pps_conformance_window_flag is 0, the values ​​of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset are estimated to be 0.

[0053] Here, the chroma format variable ChromaFormatIdc is the value of sps_chroma_format_id. The variables SubWidthC and SubHightC are determined by this ChromaFormatIdc. For monochrome formats, SubWidthC and SubHightC are both 1, for 4:2:0 formats, SubWidthC and SubHightC are both 2, for 4:2:2 formats, SubWidthC is 2 and SubHightC is 1, and for 4:4:4 formats, SubWidthC and SubHightC are both 1. · pps_init_qp_minus26 is information for deriving the quantization parameter SliceQpY of the slice referenced by the PPS.

[0054] (Subpicture) A picture may be further divided into rectangular sub-pictures. The size of a sub-picture may be a multiple of the CTU. A sub-picture is defined as a set of an integer number of consecutive tiles in both the vertical and horizontal directions. That is, a picture is divided into rectangular tiles, and a subpicture is defined as a set of rectangular tiles. A subpicture may be defined using the IDs of its top left tile and bottom right tile.

[0055] (Encoded Picture) The coded picture defines a set of data to be referenced by the image decoding device 31 in order to decode the picture PICT to be processed. As shown in FIG. 2, the picture PICT is PH, including slices 0 to NS-1 (NS is the total number of slices contained in the picture PICT) .

[0056] It also includes information (ph_qp_delta) for deriving the quantization parameter SliceQpY that is updated at the picture level.

[0057] SliceQpY = 26 + pps_init_qp_minus26 + ph_qp_delta In the following, when there is no need to distinguish between slices 0 to NS-1, the subscripts of the signs are The same applies to other data to which subscripts are added that are included in the encoded data Te described below.

[0058] (Coded Slice) In the case of the coded slice, the image decoding device 31 refers to the coded slice S to decode the slice S to be processed. As shown in Figure 2, a slice consists of a slice header and and contains slice data.

[0059] The slice header includes a group of coding parameters to be referred to by the image decoding device 31 in order to determine a decoding method for the current slice. Slice type designation information (slice_type) that designates a slice type is an example of a coding parameter included in the slice header.

[0060] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction when encoding, (2) a P slice that uses uni-prediction (L0 prediction) or intra prediction when encoding, and (3) a B slice that uses uni-prediction (L0 prediction or L1 prediction), bi-prediction, or intra prediction when encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P or B slice, it refers to a slice including a block that can use inter prediction.

[0061] Also included is information (sh_qp_delta) for deriving the quantization parameter SliceQpY that is updated at the slice level.

[0062] SliceQpY = 26 + pps_init_qp_minus26 + sh_qp_delta In addition, the slice header may include a reference to a picture parameter set PPS (pic_parameter_set_id).

[0063] (Encoded slice data) The coded slice data defines a set of data to be referenced by the image decoding device 31 in order to decode the slice data to be processed. The slice data is the coded slice data shown in FIG. As shown in the header, it contains a CTU, which is a fixed size (e.g. A block of 64x64 pixels is sometimes called a Largest Coding Unit (LCU).

[0064] (coding tree unit) 2 specifies a set of data that the image decoding device 31 refers to in order to decode a CTU to be processed. The CTU is divided into coding units CU, which are basic units of encoding processing, by recursive quad tree division (QT (Quad Tree) division), binary tree division (BT (Binary Tree) division), or ternary tree division (TT (Ternary Tree) division). BT division and TT division are collectively called multi tree division (MT (Multi Tree) division). A node of a tree structure obtained by recursive quad tree division is called a coding node. Intermediate nodes of a quad tree, binary tree, and ternary tree are coding nodes, and the CTU itself is specified as the top coding node.

[0065] (Encoding Unit) FIG. 2 shows data to be referenced by the image decoding device 31 in order to decode the coding unit to be processed. Specifically, a CU consists of a CU header CUH, prediction parameters, and transformation parameters. The CU header includes a quantization transform coefficient, a quantization metric, etc. The CU header specifies a prediction mode, etc.

[0066] The prediction process may be performed in units of CUs, or in units of sub-CUs obtained by further dividing a CU.

[0067] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is a prediction within the same picture, while inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).

[0068] Transformation and quantization are performed in units of CUs, but quantization coefficients are stored in units of subblocks such as 4x4. It may be entropy coded.

[0069] (Prediction parameters) The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.

[0070] Hereinafter, prediction parameters of inter prediction will be described. Inter prediction parameters are composed of prediction list use flags predFlagL0 and predFlagL1, reference picture indexes refIdxL0 and refIdxL1, and motion vectors mvL0 and mvL1. predFlagL0 and predFlagL1 are flags indicating whether or not a reference picture list (L0 list, L1 list) is used, and when the value is 1, the corresponding reference picture list is used. Note that in this specification, when "a flag indicating whether or not XX" is written, a flag other than 0 (for example, 1) is XX, 0 is not XX, and 1 is treated as true and 0 is treated as false in logical negation, logical product, etc. (similarly below). However, in an actual device or method, other values ​​can be used as true and false values.

[0071] (Reference picture list) The reference picture list is a list of reference pictures stored in the reference picture memory 306. FIG. 4 is a conceptual diagram showing an example of a reference picture and a reference picture list. In the conceptual diagram of an example of a reference picture in FIG. 4, the rectangles represent pictures and the arrows represent the positions of the pictures. Reference relationship, the horizontal axis is time, and I, P, and B in the rectangle are intra picture, uni-prediction picture, and bi-prediction picture, respectively. The numbers in the rectangles indicate the decoding order of predicted pictures. Tid is the value of TemporalID, which indicates the depth of the hierarchy. and is sent in the NAL unit header. As shown in the figure, the decoding order of pictures is I0, P1 , B2, B3, B4, and the display order is I0, B3, B2, B4, P1. The reference picture list is an example of a reference picture list for a frame (a frame of a picture that is a reference picture). A reference picture list is a list indicating candidates for a picture (slice), and one picture (slice) may have one or more reference picture lists. In the example shown in the figure, the target picture B3 has two reference picture lists, an L0 list RefPicList0 and an L1 list RefPicList1. In each CU, refIdxLX is used to specify which picture in the reference picture list RefPicListX (X=0 or 1) is actually referenced. The figure shows an example where refIdxL0=2 and refIdxL1=0. Note that LX is a description method used when there is no distinction between L0 prediction and L1 prediction, and hereinafter, parameters for the L0 list and parameters for the L1 list are distinguished by replacing LX with L0 and L1.

[0072] (Configuration of an image decoding device) The configuration of an image decoding device 31 (FIG. 5) according to this embodiment will be described.

[0073] The image decoding device 31 includes an entropy decoding unit 301, a parameter decoding unit (prediction image decoding device) 302, a loop filter 305, a reference picture memory 306, a prediction parameter memory 307, and a prediction image A generation unit (prediction image generation device) 308, an inverse quantization and inverse transformation unit 311, an addition unit 312, a prediction parameter The image decoding device 320 is configured to include a data derivation unit 320. In some configurations, the device 31 does not include a loop filter 305 .

[0074] The parameter decoding unit 302 further includes a header decoding unit 3020, a CT information decoding unit 3021, and a CU decoding unit 3022. The CU decoding unit 3022 further includes a TU decoding unit 3024. These may be collectively referred to as a decoding module. The header decoding unit 3020 decodes parameter set information such as VPS, SPS, PPS, and APS, and slice headers (slice information) from the encoded data. The CT information decoding unit 3021 decodes the CT from the encoded data. The CU decoding unit 3022 decodes the CU from the encoded data.

[0075] In a mode other than the skip mode (skip_mode==0), the TU decoding unit 3024 decodes the QP update information and the quantized prediction error from the encoded data.

[0076] The predicted image generating unit 308 includes an inter predicted image generating unit 309 and an intra predicted image generating unit 310. It is composed of:

[0077] The prediction parameter derivation unit 320 is configured to include an inter prediction parameter derivation unit 303 and an intra prediction parameter derivation unit 304 .

[0078] The entropy decoding unit 301 performs entropy decoding on the encoded data Te input from the outside. The entropy decoder 301 performs decoding to decode each code (syntax element). There are two types of entropy coding: one is to perform variable-length coding of syntax elements using a context (probability model) adaptively selected according to the type of syntax element and surrounding circumstances, and the other is to perform variable-length coding of syntax elements using a predetermined table or formula. The former CABAC (Context Adaptive Binary Arithmetic Coding) stores the CABAC state of the context (probability state index pStateIdx that specifies the type (0 or 1) and probability of the most probable symbol) in memory. The entropy decoder 301 initializes all CABAC states at the beginning of a segment (tile, CTU row, slice). The entropy decoder 301 converts the syntax element into a binary string (Bin String) and decodes each bit of the Bin String. When a context is used, a context index ctxInc is derived for each bit of the syntax element, the bit is decoded using the context, and the CABAC state of the used context is updated. Bits without context are decoded with equal probability (EP, bypass), and ctxInc derivation and CABAC state are omitted. The decoded syntax elements include prediction information for generating a predicted image and prediction error for generating a difference image.

[0079] The entropy decoding unit 301 outputs the decoded code to the parameter decoding unit 302. The control as to whether to decode the parameter is performed based on an instruction from the parameter decoding unit 302.

[0080] (Basic flow) FIG. 6 is a flowchart illustrating a schematic operation of the image decoding device 31.

[0081] (S1100: Decode Parameter Set Information) The header decoding unit 3020 decodes parameter set information such as the VPS, SPS, and PPS from the encoded data.

[0082] (S1200: Decode slice information) The header decoding unit 3020 decodes slice header information from the encoded data. Decode (slice information).

[0083] Hereinafter, the image decoding device 31 performs the steps S1300 to S5000 for each CTU included in the target picture. The process is repeated to derive a decoded image for each CTU.

[0084] (S1300: Decode CTU Information) The CT information decoding unit 3021 decodes the CTU from the encoded data.

[0085] (S1400: Decode CT Information) The CT information decoding unit 3021 decodes the CT from the encoded data.

[0086] (S1500: CU Decoding) The CU decoding unit 3022 performs S1510 and S1520 to decode the CU from the encoded data. Issued.

[0087] (S1510: Decode CU information) The CU decoding unit 3022 decodes CU information, prediction information, etc. from the encoded data. do.

[0088] (S1520: Decode TU information) When a TU includes a prediction error, the TU decoding unit 3024 decodes the encoded TU information. The QP update information and the quantized prediction error are decoded from the data. Note that the QP update information is a difference value from the quantization parameter predicted value qPpred, which is a predicted value of the quantization parameter QP.

[0089] (S2000: Generation of predicted image) The predicted image generation unit 308 generates a predicted image for each block included in the current CU based on the prediction information.

[0090] (S3000: Inverse Quantization and Inverse Transformation) The inverse quantization and inverse transform unit 311 executes inverse quantization and inverse transform processing for each TU included in the target CU.

[0091] (S4000: Decoded image generation) The adder 312 receives a predicted image from the predicted image generator 308 and The prediction error supplied from the inverse quantization and inverse transform unit 311 is added to the A decoded image is generated.

[0092] (S5000: Loop Filter) The loop filter 305 applies a loop filter such as a deblocking filter, SAO, or ALF to the decoded image to generate a decoded image.

[0093] (Configuration of inter-prediction parameter derivation unit) The loop filter 305 is a filter provided in the coding loop, and is used to remove block distortion and ringing. The loop filter 305 is a filter that removes distortion and improves image quality. The loop filter 305 applies a deblocking filter, a sample adaptive offset (SAO), and an adaptive filter to the decoded image of the CU generated by the adder 312. Apply a filter such as an automatic loop filter (ALF).

[0094] The DF unit 601 is a bS deriving unit that derives the strength bS of the deblocking filter in units of pixels, boundaries, and lines. The output unit 602 performs deblocking filter processing to reduce block noise. It is composed of a data section 602.

[0095] The DF unit 601 derives an edge degree edgeIdc indicating whether there is a partition boundary, a prediction block boundary, or a transform block boundary in the input image resPicture before NN (Neural Network) processing (processing by the NN filter unit 601), and a maximum filter length maxFilterLength of the deblocking filter. Furthermore, the DF unit 601 derives a strength bS of the deblocking filter from edgeIdc, the transform block boundary, and the coding parameters.

[0096] The reference picture memory 306 stores the decoded image of the CU in a predetermined Store it in position.

[0097] The prediction parameter memory 307 stores prediction parameters at a predetermined location for each CTU or CU.

[0098] The predicted image generating unit 308 receives the parameters derived by the prediction parameter derivation unit 320. The predicted image generating unit 308 also reads a reference picture from the reference picture memory 306. The predicted image generating unit 308 generates a block image using the parameters and the reference picture (reference picture block). A predicted image of the block is generated.

[0099] The inverse quantization and inverse transform unit 311 (residual decoding unit) inverse quantizes and inverse transforms the quantized transform coefficients input from the parameter decoding unit 302 to obtain transform coefficients.

[0100] (Neural network post-filter characteristics SEI) 9 shows the syntax of the neural network post-filter characteristics SEI message nn_post_filter_characteristics(payloadSize) of Non-Patent Document 1. The argument payloadSize indicates the number of bytes of this SEI message.

[0101] This SEI message is applied to each CVS (Coded Video Sequence). Note that CVS refers to IRAP (Intra Random Access Pictures) and GDR (Gradual Decoder Refresh Pictures). A randomly accessible access unit such as An access unit is a collection of pictures that are displayed at the same time. An IRAP may be Instantaneous Decoder Refresh (IDR), Clean Random Access (CRA), or Broken Link Access (BLA).

[0102] The following variables are defined in this SEI message:

[0103] The width and height of the decoded image, in units of luma pixels, are denoted here by InpPicWidthInLumaSamples and InpPicHeightInLumaSamples, respectively.

[0104] InpPicWidthInLumaSamples is set equal to pps_pic_width_in_luma_samples-SubWidthC*(pps_conf_win_left_offset+pps_conf_win_right_offset).

[0105] InpPicHeightInLumaSamples is set equal to pps_pic_height_in_luma_samples-SubHeightC*(pps_conf_win_top_offset+pps_conf_win_bottom_offset).

[0106] The decoded image is a two-dimensional array of luminance pixels, CroppedYPic[y][x], with vertical coordinate y and horizontal coordinate x, and two-dimensional arrays of chrominance pixels, CroppedCbPic[y][x] and CroppedCrPic[y][x]. Here, the y coordinate of the upper left corner of the pixel array is 0, and the x coordinate is 0.

[0107] The pixel bit length of the luminance of the decoded image is BitDepthY. The pixel bit length of the chrominance of the decoded image is BitDepthC. Note that both BitDepthY and BitDepthC are set equal to BitDepth.

[0108] The variable InpSubWidthC is the chrominance subsampling ratio for the horizontal luminance of the decoded image, and the variable InpSubHeightC is the chrominance subsampling ratio for the horizontal luminance of the decoded face image. Note that InpSubWidthC is set equal to the variable SubWidthC of the encoded data. InpSubHeightC is set equal to the variable SubHeightC of the encoded data.

[0109] The variable SliceQPY is set equal to the quantization parameter SliceQpY that is updated at the slice level of the encoded data.

[0110] nnpfc_id contains an identification number that can be used to identify the postfilter process. The value of nnpfc_id must be between 0 and 2^32-2 inclusive. nnpfc_id values ​​between 256 and 511 inclusive, or between 2^31 and 2^32-2 inclusive, are reserved for future use. Therefore, decoders shall ignore nnpfc_id values ​​between 256 and 511 inclusive, or between 2^31 and 2^32-2 inclusive.

[0111] If the value of nnpfc_mode_idc is 0, the associated postfiltering is disabled in this specification. Specifies that the value is determined by external means that are not specified.

[0112] If the value of nnpfc_mode_idc or nnpfc_id is 1, the associated postfilter processing is Indicates that the neural network contained in this SEI message is represented by the ISO / IEC 15938-17 bit stream.

[0113] The value of nnpfc_mode_idc MUST be a value between 0 and 255, inclusive. Values ​​of nnpfc_mode_idc greater than 1 are reserved for future specification and therefore MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification SHALL ignore SEI messages containing reserved values ​​of nnpfc_mode_idc. If the current Coded Layer Video Sequence (CLVS) contains a preceding Neural Network Post-Filter Characteristics SEI message, in decoding order, that has the same value of nnpfc_id as the value of this SEI message's nnpfc_id for which at least one of the following two conditions apply:

[0114] This SEI message has the same content as the previous Neural Network Post Filter Characteristics SEI message, but provides an update to the neural network. The values ​​of nnpfc_mode_idc and nnpfc_payload_byte[i] are different to achieve this.

[0115] This SEI message has the same content as the neural network post-filter characteristics SEI message described above.

[0116] This SEI message contains the current decoded picture and all subsequent ones up to the end of the current CLVS. It performs post-filtering on all decoded pictures.

[0117] nnpfc_purpos indicates the purpose of post-filter processing. The value of nnpfc_purpose must be between 0 and 2^32-2. Values ​​of nnpfc_purpose greater than 4 are reserved for future specifications. This is reserved for future use only and MUST NOT be present in bitstreams conforming to this version of the specification. Decoders conforming to this version of the specification MUST NOT SEI messages containing reserved values ​​SHALL be ignored.

[0118] A value of 0 for nnpfc_purpose indicates unknown or undefined.

[0119] When the value of nnpfc_purpose is 1, the purpose is to improve image quality.

[0120] If the value of nnpfc_purpose is 2, upsampling is performed to 4:2:2 chrominance format or 4:4:4 chrominance format, or upsampling from 4:2:2 chrominance format to 4:4:4 chrominance format. A value of 3 for nnpfc_purpose increases the width or height of the decoded output image without changing the chroma format.

[0121] A value of 4 for nnpfc_purpose will increase the width or height of the decoded output image and upsample the chrominance format.

[0122] The nnpfc_out_sub_width_c_flag and nnpfc_out_sub_height_c_flag are used to derive the variables outSubWidthC and outSubHeightC, respectively, which specify the chrominance to luma subsampling ratio of the resulting postfiltered image. If not present, nnpfc_out_sub_width_c_flag and nnpfc_out_sub_height_c_flag are inferred to be both equal to 0. If nnpfc_out_sub_width_c_flag and nnpfc_out_sub_height_c_flag are present, the sum of nnpfc_out_sub_width_c_flag and nnpfc_out_sub_height_c_flag must be greater than 0.

[0123] outSubWidthC=InpSubWidthC-nnpfc_out_sub_width_c_flag outSubHeightC=InpSubHeightC-nnpfc_out_sub_height_c_flag A bitstream conformance requirement is that both outSubWidthC and outSubHeightC be greater than 0.

[0124] In the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, there is a problem that a non-existent chrominance format in which outSubWidthC is 1 and outSubHeightC is 2 can be defined.

[0125] Therefore, the requirement for bitstream compliance is that both outSubWidthC and outSubHeightC are greater than 0, and the value of outSubWidthC is greater than the value of outSubHeightC. do.

[0126] By processing a bitstream defined in this way, the problem of a non-existent chrominance format being defined (derived in the process) can be solved.

[0127] As another solution, the bitstream compliance requirement is that both outSubWidthC and outSubHeightC are greater than 0, and the value of outSubWidthC is not 1, and the value of outSubHeightC is not 2.

[0128] As another solution, a requirement for bitstream compliance is that both outSubWidthC and outSubHeightC are greater than 0, and the value of outSubWidthC is equal to or greater than the value of outSubHeightC.

[0129] The syntax elements nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the image luma pixel array obtained by applying the postfilter identified by nnpfc_id to the decoded image. If nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to InpPicWidthInLumaSamples and InpPicHeightInLumaSamples, respectively.

[0130] A value of 0 for the syntax element nnpfc_component_last_flag specifies that the second dimension of the input tensor inputTensor is used for postfiltering, and the output tensor outputTensor resulting from the postfiltering is used for the channels.

[0131] A value of 1 for nnpfc_component_last_flag specifies that the last dimension of the input tensor is used for postfiltering, and the output tensor outputTensor resulting from the postfiltering is used for the channels.

[0132] The syntax element nnpfc_inp_sample_idc indicates how to convert the pixel values ​​of the decoded image into input values ​​for post-filter processing. When the value of nnpfc_inp_sample_idc is 0, 1, 2, or 3, the input values ​​for post-filter processing are in the floating-point value format of binary16, binary32, binary64, or binary128 specified in IEEE754-2019, respectively, and the functions InpY and InpC are specified as follows:

[0133] InpY(x) = x÷((1< <BitDepthY)-1) InpC(x) = x ÷ ((1< <BitDepthC)-1) If the value of nnpfc_inp_sample_idc is 4, the input to the postfilter is an unsigned integer. In numerical terms, the functions InpY and InpC are specified as follows:

[0134] shift=BitDepthY-inpTensorBitDepth if(inpTensorBitDepth>=BitDepthY) InpY(x)=x<<(inpTensorBitDepth-BitDepthY) else InpY(x)=Clip3(0,(1< <inpTensorBitDepth)-1,(x+(1<<(shift-1)))> >shift) shift=BitDepthC-inpTensorBitDepth if(inpTensorBitDepth>=BitDepthC) InpC(x)=x<<(inpTensorBitDepth-BitDepthC) else InpC(x) = Clip3(0,(1< <inpTensorBitDepth)-1,(x+(1<<(shift-1)))> >shift) The variable inpTensorBitDepth is obtained from the syntax element nnpfc_inp_tensor_bitdepth_minus8, which will be described later.

[0135] The value of nnpfc_inp_sample_idc must be between 0 and 255. Values ​​of nnpfc_inp_sample_idc greater than 4 are reserved for future use and should not be used in this specification. A decoder conforming to this version of the specification shall not be present in any SEI message containing a reserved value of nnpfc_inp_sample_idc. The message shall be ignored.

[0136] The value of the syntax element nnpfc_inp_tensor_bitdepth_minus8 plus 8 is the input integer The pixel bit depth of the tensor's luminance pixel values. The value of the variable inpTensorBitDepth is as follows: is derived.

[0137] inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8 As a bitstream compliance requirement, the value of nnpfc_inp_tensor_bitdepth_minus8 must be in the range 0 to 24.

[0138] The syntax element nnpfc_inp_order_idc specifies the pixel order of the decoded image for post-filter processing. The semaphores of nnpfc_inp_order_idc in the range 0 to 3 are used to indicate how to order the inputs of The metric specifies the process of deriving the input tensor inputTensor for each value of nnpfc_inp_order_idc. It also specifies the vertical pixel coordinate cTop and horizontal pixel coordinate cLeft from which the input tensor is derived. Specifies the top left pixel position of the patch of pixels to be decoded. If the chrominance format of the decoded image is not 4:2:0 The value of nnpfc_inp_order_idc cannot be 3. The value of nnpfc_inp_order_idc must be between 0 and 255. Values ​​of nnpfc_inp_order_idc greater than 3 are reserved for future specification. Reserved and therefore MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification SHALL ignore SEI messages containing a reserved value of nnpfc_inp_order_idc.

[0139] A patch is a rectangular array of pixels from a component of an image (e.g., luma and chroma components).

[0140] The syntax element nnpfc_constant_patch_size_flag, when set to 0, indicates that the postfilter process accepts as input patch sizes that are positive integer multiples of the patch sizes specified by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. When nnpfc_patch_width_minus1 is set to 1, the postfilter process accepts as input patch sizes specified by nnpfc_patch_height_minus1.

[0141] The value of the syntax element nnpfc_patch_width_minus1 plus 1 indicates the number of horizontal pixels of the patch size required for input to the postfilter process if the value of nnpfc_constant_patch_size_flag is 1. If the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_width_minus1+1) can be used as the number of horizontal pixels of the patch size used for input to the postfilter process. The value of nnpfc_patch_width_minus1 must be between 0 and 32766 inclusive.

[0142] The value of the syntax element nnpfc_patch_height_minus1 plus 1 indicates the number of vertical pixels of the patch size required for input to the postfilter process if the value of nnpfc_constant_patch_size_flag is 1. If the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_height_minus1+1) can be used as the number of vertical pixels of the patch size used for input to the postfilter process. The value of nnpfc_patch_height_minus1 must be between 0 and 32766 inclusive.

[0143] In the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, There was a problem that it was possible to define a patch size larger than the chat size.

[0144] Therefore, the value of nnpfc_patch_width_minus1 must be between 0 and InpPicWidthInLumaSamples-1. The variable InpPicWidthInLumaSamples is the pixel value of the luminance width of the input image of this SEI. It indicates a prime number. In addition, the value of nnpfc_patch_height_minus1 is greater than or equal to 0 and less than or equal to InpPicHeightInLumaSamples-1. The variable InpPicHeightInLumaSamples indicates the number of pixels of the luminance height of the image to which this SEI is input.

[0145] In this way, by defining a value range of the bitstream syntax and making it the target of processing, it is possible to solve the problem of being able to define a patch size larger than the picture size (outputting a patch size larger than the picture size).

[0146] The syntax element nnpfc_overlap specifies the horizontal overlap of adjacent input tensors. Specifies the number of pixels and the number of vertical pixels. The value of nnpfc_overlap must be between 0 and 16383. It won't happen.

[0147] Variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling , verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are as follows: Derive.

[0148] inpPatchWidth=nnpfc_patch_width_minus1+1 inpPatchHeight=nnpfc_patch_height_minus1+1 outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / InpPicWidthInLumaSamples outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / InpPicHeightInLumaSamples horCScaling=InpSubWidthC / outSubWidthC verCScaling=InpSubHeightC / outSubHeightC outPatchCWidth=outPatchWidth*horCScaling outPatchCHeight=outPatchHeight*verCScaling overlapSize=nnpfc_overlap Note that outPatchWidth*InpPicWidthInLumaSamples must be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*InpPicHeightInLumaSamples must be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.

[0149] The syntax element nnpfc_padding_type specifies the padding process when referring to pixel locations outside the boundaries of the decoded image.

[0150] The value of nnpfc_padding_type must be between 0 and 15 inclusive.

[0151] If the value of nnpfc_padding_type is 0, the value of pixel positions outside the boundary of the decoded image is set to 0.

[0152] If the value of nnpfc_padding_type is 1, the pixel position outside the boundary of the decoded image is set as the boundary value. do.

[0153] If the value of nnpfc_padding_type is 2, the pixel values ​​outside the boundary of the decoded image are set to the boundary value. Let the value be the specular reflection value. To define the value more precisely, we define the function InpSampleVal. The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes the vertical pixel position y, the horizontal pixel position x, the image height picHeight, the image width picWidth, and the pixel array croppedPic, and returns the value of sampleVal derived as follows:

[0154] if(nnpfc_padding_type==0) if(y<0||x<0||y>=picHeight||x>=picWidth) sampleVal=0 else sampleVal=croppedPic[y][x] else if(nnpfc_padding_type==1) sampleVal=croppedPic[Clip3(0,picHeight-1,y)][Clip3(0,picWidth-1,x)] else / *nnpfc_padding_type==2* / sampleVal=croppedPic[Reflect(picHeight-1,y)][Reflect(picWidth-1,x)] The function Reflect(y,z) can be expressed by Equation 1 in FIG.

[0155] nnpfc_padding_type values ​​greater than 3 are reserved for future use. and therefore MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification SHALL ignore SEI messages containing reserved values ​​of nnpfc_padding_type.

[0156] A patch is a rectangular array of pixels from a component of an image (e.g., luma and chroma components).

[0157] When the syntax element nnpfc_constant_patch_size_flag is 0, post-filter processing is performed using nnpfc nnpfc_patch_width_minus1 indicates that the input patch size should be a positive integer multiple of the patch size indicated by nnpfc_patch_height_minus1. In this case, the postfilter process uses the patch size indicated by nnpfc_patch_height_minus1. Indicates that it is accepted as input.

[0158] The value of the syntax element nnpfc_patch_width_minus1 plus 1 indicates the number of horizontal pixels of the patch size required for input to the postfilter process if the value of nnpfc_constant_patch_size_flag is 1. If the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_width_minus1+1) can be used as the number of horizontal pixels of the patch size used for input to the postfilter process. The value of nnpfc_patch_width_minus1 must be between 0 and 32766 inclusive.

[0159] The value of the syntax element nnpfc_patch_height_minus1 plus 1 indicates the number of vertical pixels of the patch size required for input to the postfilter process if the value of nnpfc_constant_patch_size_flag is 1. If the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_height_minus1+1) can be used as the number of vertical pixels of the patch size used for input to the postfilter process. The value of nnpfc_patch_height_minus1 must be between 0 and 32766 inclusive.

[0160] In the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, There was a problem in that it was possible to define a patch size larger than the picture size of the image.

[0161] Therefore, the value of nnpfc_patch_width_minus1 must be between 0 and InpPicWidthInLumaSamples-1. The variable InpPicWidthInLumaSamples is the pixel value of the luminance width of the input image of this SEI. It indicates a prime number. In addition, the value of nnpfc_patch_height_minus1 is greater than or equal to 0 and less than or equal to InpPicHeightInLumaSamples-1. The variable InpPicHeightInLumaSamples indicates the number of pixels of the luminance height of the image to which this SEI is input.

[0162] Defining the range of the syntax in this way solves the problem that a patch size larger than the picture size of the decoded image can be defined.

[0163] The syntax element nnpfc_overlap specifies the horizontal overlap of adjacent input tensors. Specifies the number of pixels and the number of vertical pixels. The value of nnpfc_overlap must be between 0 and 16383. It won't happen.

[0164] The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling , verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows :

[0165] inpPatchWidth = nnpfc_patch_width_minus1 + 1 inpPatchHeight = nnpfc_patch_height_minus1 + 1 outPatchWidth = (nnpfc_pic_width_in_luma_samples * inpPatchWidth) / InpPicWidthInLumaSamples outPatchHeight = (nnpfc_pic_height_in_luma_samples * inpPatchHeight) / InpPicHeightInLumaSamples horCScaling = InpSubWidthC / outSubWidthC verCScaling = InpSubHeightC / outSubHeightC outPatchCWidth = outPatchWidth * horCScaling outPatchCHeight = outPatchHeight * verCScaling overlapSize = nnpfc_overlap Note that outPatchWidth*InpPicWidthInLumaSamples must be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*InpPicHeightInLumaSamples must be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.

[0166] The syntax element nnpfc_padding_type specifies the padding process when referring to pixel locations outside the boundaries of the decoded image.

[0167] The value of nnpfc_padding_type must be between 0 and 15 inclusive.

[0168] If the value of nnpfc_padding_type is 0, the value of pixel positions outside the boundary of the decoded image is set to 0.

[0169] If the value of nnpfc_padding_type is 1, the pixel position outside the boundary of the decoded image is set as the boundary value. do.

[0170] If the value of nnpfc_padding_type is 2, the pixel values ​​outside the boundary of the decoded image are set to the boundary value. The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) is used to calculate the vertical pixel position of the input. Given the pixel position y, horizontal pixel position x, image height picHeight, image width picWidth, and pixel array croppedPic, returns the value of sampleVal derived as follows:

[0171] if(nnpfc_padding_type==0) if(y<0||x<0||y>=picHeight||x>=picWidth) sampleVal=0 else sampleVal=croppedPic[y][x] else if(nnpfc_padding_type==1) sampleVal=croppedPic[Clip3(0,picHeight-1,y)][Clip3(0,picWidth-1,x)] else / *nnpfc_padding_type==2* / sampleVal=croppedPic[Reflect(picHeight-1,y)][Reflect(picWidth-1,x)] The function Reflect(y,z) can be expressed by Equation 1 in FIG.

[0172] nnpfc_padding_type values ​​greater than 3 are reserved for future use. and therefore MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification SHALL ignore SEI messages containing reserved values ​​of nnpfc_padding_type.

[0173] If the value of nnpfc_inp_order_idc is 0, the input tensor only has a brightness matrix, so the number of channels is 1. The process for deriving the input tensors, DeriveInputTensors(), is as follows:

[0174] for(yP=-overlapSize;yP <inpPatchHeight+overlapSize;yP++) for(xP=-overlapSize;xP <inpPatchWidth+overlapSize;xP++){ inpVal=InpY(InpSampleVal(cTop+yP,cLeft+xP,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) if(nnpfc_component_last_flag==0) inputTensor[0][0][yP+overlapSize][xP+overlapSize]=inpVal else inputTensor[0][yP+overlapSize][xP+overlapSize][0]=inpVal } If the value of nnpfc_inp_order_idc is 1, then only the chrominance matrix is ​​present in the input tensor, so the number of channels is 2. The process for deriving the input tensors, DeriveInputTensors(), is as follows:

[0175] for(yP=-overlapSize;yP <inpPatchHeight+overlapSize;yP++) for(xP=-overlapSize;xP <inpPatchWidth+overlapSize;xP++){ inpCbVal=InpC(InpSampleVal(cTop+yP,cLeft+xP, InpPicHeightInLumaSamples / InpSubHeightC, InpPicWidthtInLumaSamples / InpSubWidthC,CroppedCbPic)) inpCrVal=InpC(InpSampleVal(cTop+yP,cLeft+xP, InpPicHeightInLumaSamples / InpSubHeightC, InpPicWidthtInLumaSamples / InpSubWidthC,CroppedCrPic)) if(nnpfc_component_last_flag==0){ inputTensor[0][0][yP+overlapSize][xP+overlapSize]=inpCbVal inputTensor[0][1][yP+overlapSize][xP+overlapSize]=inpCrVal }else{ inputTensor[0][yP+overlapSize][xP+overlapSize][0]=inpCbVal inputTensor[0][yP+overlapSize][xP+overlapSize][1]=inpCrVal } } If the value of nnpfc_inp_order_idc is 2, the luma and chroma matrices are present in the input tensor, so the number of channels is 3. The process for deriving the input tensors, DeriveInputTensors(), is as follows:

[0176] for(yP=-overlapSize;yP <inpPatchHeight+overlapSize;yP++) for(xP=-overlapSize;xP <inpPatchWidth+overlapSize;xP++){ yY=cTop+yP xY=cLeft+xP yC=yY / InpSubHeightC xC=xY / InpSubWidthC inpYVal=InpY(InpSampleVal(yY,xY,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) inpCbVal=InpC(InpSampleVal(yC,xC,InpPicHeightInLumaSamples / InpSubHeightC, InpPicWidthtInLumaSamples / InpSubWidthC,CroppedCbPic)) inpCrVal=InpC(InpSampleVal(yC,xC,InpPicHeightInLumaSamples / InpSubHeightC, InpPicWidthtInLumaSamples / InpSubWidthC,CroppedCrPic)) if(nnpfc_component_last_flag==0){ inputTensor[0][0][yP+overlapSize][xP+overlapSize]=inpYVal inputTensor[0][1][yP+overlapSize][xP+overlapSize]=inpCbVal inputTensor[0][2][yP+overlapSize][xP+overlapSize]=inpCrVal } else { inputTensor[0][yP+overlapSize][xP+overlapSize][0]=inpYVal inputTensor[0][yP+overlapSize][xP+overlapSize][1]=inpCbVal inputTensor[0][yP+overlapSize][xP+overlapSize][2]=inpCrVal } } If the value of nnpfc_inp_order_idc is 3, the input tensor has 4 luma matrices, 2 chroma matrices, and a quantization parameter matrix, so the number of channels is 7. The channels are derived in an interleaved manner. This nnpfc_inp_order_idc can only be used if the chroma format is 4:2:0. If nnpfc_inp_order_idc is equal to 3, the variable SliceQPY is set equal to the per-slice SliceQpY. The process for deriving the input tensors, DeriveInputTensors(), is as follows:

[0177] for(yP=-overlapSize;yP <inpPatchHeight+overlapSize;yP++) for(xP=-overlapSize;xP <inpPatchWidth+overlapSize;xP++){ yTL=cTop+yP*2 xTL=cLeft+xP*2 yBR=yTL+1 xBR=xTL+1 yC=cTop / 2+yP xC=cLeft / 2+xP inpTLVal=InpY(InpSampleVal(yTL,xTL,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) inpTRVal=InpY(InpSampleVal(yTL,xBR,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) inpBLVal=InpY(InpSampleVal(yBR,xTL,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) inpBRVal=InpY(InpSampleVal(yBR,xBR,InpPicHeightInLumaSamples, InpPicWidthtInLumaSamples,CroppedYPic)) inpCbVal=InpC(InpSampleVal(yC,xC,InpPicHeightInLumaSamples / 2, InpPicWidthtInLumaSamples / 2,CroppedCbPic)) inpCrVal=InpC(InpSampleVal(yC,xC,InpPicHeightInLumaSamples / 2, InpPicWidthtInLumaSamples / 2,CroppedCrPic)) if(nnpfc_component_last_flag==0){ inputTensor[0][0][yP+overlapSize][xP+overlapSize]=inpTLVal inputTensor[0][1][yP+overlapSize][xP+overlapSize]=inpTRVal inputTensor[0][2][yP+overlapSize][xP+overlapSize]=inpBLVal inputTensor[0][3][yP+overlapSize][xP+overlapSize]=inpBRVal inputTensor[0][4][yP+overlapSize][xP+overlapSize]=inpCbVal inputTensor[0][5][yP+overlapSize][xP+overlapSize]=inpCrVal inputTensor[0][6][yP+overlapSize][xP+overlapSize]=2*(SliceQPY-42) / 6 }else{ inputTensor[0][yP+overlapSize][xP+overlapSize][0]=inpTLVal inputTensor[0][yP+overlapSize][xP+overlapSize][1]=inpTRVal inputTensor[0][yP+overlapSize][xP+overlapSize][2]=inpBLVal inputTensor[0][yP+overlapSize][xP+overlapSize][3]=inpBRVal inputTensor[0][yP+overlapSize][xP+overlapSize][4]=inpCbVal inputTensor[0][yP+overlapSize][xP+overlapSize][5]=inpCrVal inputTensor[0][yP+overlapSize][xP+overlapSize][6]=2*(SliceQPY-42) / 6 } } The syntax element nnpfc_complexity_idc indicates that there may be one or more syntax elements indicating the postfiltering complexity associated with nnpfc_id. A value of 0 for nnpfc_complexity_idc indicates that there are no syntax elements indicating the postfiltering complexity associated with nnpfc_id. nnpfc_complexity_idc values ​​MUST be between 0 and 255 inclusive. Values ​​of nnpfc_complexity_idc greater than 1 are reserved for future specification and MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification SHALL ignore SEI messages containing reserved values ​​of nnpfc_complexity_idc.

[0178] The syntax element nnpfc_out_sample_idc is a value between 0 and 3, inclusive, which indicates that the pixel values ​​output by post-filtering are binary16, binary32, binary64, or binary128 floating-point values, respectively, as specified in IEEE754-2019. The functions OutY and OutC, which convert the luminance and chrominance pixel values ​​output by post-processing to integer values ​​of the pixel bit length, respectively, are specified as follows using the pixel bit lengths BitDepthY and BitDepthC, respectively:

[0179] OutY(x)=Clip3(0,(1< <BitDepthY)-1,Round(x*((1<<BitDepthY)-1))) OutC(x)=Clip3(0,(1< <BitDepthC)-1,Round(x*((1<<BitDepthC)-1))) If the value of nnpfc_out_sample_idc is 4, the pixel value output by postfilter processing is Indicates that it is an unsigned integer. The functions OutY and OutC are specified as follows:

[0180] shift=outTensorBitDepth-BitDepthY if(outTensorBitDepth>=BitDepthY) OutY(x)=Clip3(0,(1< <BitDepthY)-1,(x+(1<<(shift-1)))> >shift) else OutY(x)=x<<(BitDepthY-outTensorBitDepth) shift=outTensorBitDepth-BitDepthC if(outTensorBitDepth>=BitDepthC) OutC(x)=Clip3(0,(1< <BitDepthC)-1,(x+(1<<(shift-1)))> >shift) else OutC(x)=x<<(BitDepthC-outTensorBitDepth) In the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, The above formula had a problem in that a negative shift operation occurred when outTensorBitDepth == BitDepthY and when outTensorBitDepth == BitDepthC.

[0181] Therefore, we derive it using the following formula. In this way, when the variable shift is greater than 0, it becomes a right shift operation, so the problem can be solved.

[0182] shift=outTensorBitDepth-BitDepthY if(shift>0) OutY(x)=Clip3(0,(1< <BitDepthY)-1,(x+(1<<(shift-1)))> >shift) else OutY(x)=x<<(BitDepthY-outTensorBitDepth) shift=outTensorBitDepth-BitDepthC if(shift>0) OutC(x) = Clip3(0,(1< <BitDepthC)-1,(x+(1<<(shift-1)))> >shift) else OutC(x)=x<<(BitDepthC-outTensorBitDepth) The variable outTensorBitDepth is derived from the syntax element nnpfc_out_tensor_bitdepth_minus8, which will be described later.

[0183] Also, the value of nnpfc_out_sample_idc must be in the range of 0 to 255. It cannot be greater than 4. Any value of nnpfc_out_sample_idc that is not specified is reserved for future specification and must not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification must ignore SEI messages that contain reserved values ​​of nnpfc_out_sample_idc. This shall be regarded as

[0184] Note that the specified values ​​of nnpfc_inp_sample_idc and nnpfc_out_sample_idc specify that the first dimensions of the input and output tensors, respectively, are used for the batch index, as is practiced in some neural network frameworks. The semantics of the message uses a batch size equal to 1, but the neural network It is up to the post-processing implementation to determine the batch size used as input to the network inference.

[0185] nnpfc_out_tensor_bitdepth_minus8+8 specifies the pixel bit depth of the pixel values ​​of the output integer tensor. The value of outTensorBitDepth is derived as follows:

[0186] outTensorBitDepth=nnpfc_out_tensor_bitdepth_minus8+8 Note that the value of nnpfc_out_tensor_bitdepth_minus8 must be in the range of 0 to 24.

[0187] nnpfc_out_order_idc specifies the output order of pixels obtained from post-filter processing. The semantics of nnpfc_out_order_idc for values ​​0 to 3 inclusive are specified.

[0188] FilteredYPic, FilteredCbPic, and FilteredCrPic are the output pixel arrays of filtered luma Y, chroma Cb, and Cr, respectively. They are the top-left pixel location of the patch of pixels containing the vertical pixel coordinate cTop and the horizontal pixel coordinate cLeft. When they are equal, nnpfc_out_order_idc must not be equal to 3. Value of nnpfc_out_order_idc MUST have a value between 0 and 255, inclusive. Values ​​of nnpfc_out_order_idc greater than 3 are reserved for future specification and MUST NOT be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of the specification SHALL ignore SEI messages containing reserved values ​​of nnpfc_out_order_idc.

[0189] If the value of nnpfc_out_order_idc is 0, only the luminance matrix exists in the output tensor, so the number of channels is 1. The process StoreOutputTensors() that derives pixel values ​​of the filtered output pixel array FilteredYPic from the output tensor with the vertical pixel coordinate cTop and horizontal pixel coordinate cLeft is as follows.

[0190] for(yP=0;yP <outPatchHeight;yP++) for(xP=0;xP <outPatchWidth;xP++){ yY=cTop*outPatchHeight / inpPatchHeight+yP xY=cLeft*outPatchWidth / inpPatchWidth+xP if(yY <nnpfc_pic_height_in_luma_samples && xY <nnpfc_pic_width_in_luma_samples) if(nnpfc_component_last_flag==0) FilteredYPic[yY][xY]=OutY(outputTensor[0][0][yP][xP]) else FilteredYPic[yY][xY]=OutY(outputTensor[0][yP][xP][0]) } If the value of nnpfc_out_order_idc is 1, only the chrominance matrix is ​​present in the output tensor, so the number of channels is 2. The process StoreOutputTensors() that derives pixel values ​​of the filtered output pixel arrays FilteredCbPic and FilteredCrPic from the output tensor at the vertical pixel coordinate cTop and horizontal pixel coordinate cLeft is as follows:

[0191] for(yP=0;yP <outPatchCHeight;yP++) for(xP=0;xP <outPatchCWidth;xP++){ xSrc=cLeft*horCScaling+xP ySrc=cTop*verCScaling+yP if(ySrc <nnpfc_pic_height_in_luma_samples / outSubHeightC && xSrc <nnpfc_pic_width_in_luma_samples / outSubWidthC) if(nnpfc_component_last_flag==0){ FilteredCbPic[ySrc][xSrc]=OutC(outputTensor[0][0][yP][xP]) FilteredCrPic[ySrc][xSrc]=OutC(outputTensor[0][1][yP][xP]) }else{ FilteredCbPic[ySrc][xSrc]=OutC(outputTensor[0][yP][xP][0]) FilteredCrPic[ySrc][xSrc]=OutC(outputTensor[0][yP][xP][1]) } } If the value of nnpfc_out_order_idc is 2, the number of channels is 3 because the luma matrix and the chroma matrix exist in the output tensor. The process StoreOutputTensors() that derives the pixel values ​​of the filtered output pixel arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor at the vertical pixel coordinate cTop and the horizontal pixel coordinate cLeft is as follows:

[0192] for(yP=0;yP <outPatchHeight;yP++) for(xP=0;xP <outPatchWidth;xP++){ yY=cTop*outPatchHeight / inpPatchHeight+yP xY=cLeft*outPatchWidth / inpPatchWidth+xP yC=yY / outSubHeightC xC=xY / outSubWidthC yPc=(yP / outSubHeightC)*outSubHeightC xPc=(xP / outSubWidthC)*outSubWidthC if(yY<nnpfc_pic_height_in_luma_samples && xY<nnpfc_pic_width_in_luma_samples) if(nnpfc_component_last_flag==0){ FilteredYPic[yY][xY]=OutY(outputTensor[0][0][yP][xP]) FilteredCbPic[yC][xC]=OutC(outputTensor[0][1][yPc][xPc]) FilteredCrPic[yC][xC]=OutC(outputTensor[0][2][yPc][xPc]) }else{ FilteredYPic[yY][xY]=OutY(outputTensor[0][yP][xP][0]) FilteredCbPic[yC][xC]=OutC(outputTensor[0][yPc][xPc][1]) FilteredCrPic[yC][xC]=OutC(outputTensor[0][yPc][xPc][2]) } } If the value of nnpfc_out_order_idc is 3, the output tensor has 4 luma matrices and 2 chroma matrices, so the number of channels is 6. This nnpfc_out_order_idc can only be used when the chroma format is 4:2:0. The process StoreOutputTensors() that derives pixel values ​​of the filtered output pixel arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor at vertical pixel coordinate cTop and horizontal pixel coordinate cLeft is as follows:

[0193] for(yP=0;yP <outPatchHeight;yP++) for(xP=0;xP <outPatchWidth;xP++){ ySrc=cTop / 2*outPatchHeight / inpPatchHeight+yP xSrc=cLeft / 2*outPatchWidth / inpPatchWidth+xP if(ySrc <nnpfc_pic_height_in_luma_samples / 2 && xSrc <nnpfc_pic_width_in_luma_samples / 2) if(nnpfc_component_last_flag==0){ FilteredYPic[ySrc*2][xSrc*2]=OutY(outputTensor[0][0][yP][xP]) FilteredYPic[ySrc*2][xSrc*2+1]=OutY(outputTensor[0][1][yP][xP]) FilteredYPic[ySrc*2+1][xSrc*2]=OutY(outputTensor[0][2][yP][xP]) FilteredYPic[ySrc*2+1][xSrc*2+1]=OutY(outputTensor[0][3][yP][xP]) FilteredCbPic[ySrc][xSrc]=OutC(outputTensor[0][4][yP][xP]) FilteredCrPic[ySrc][xSrc]=OutC(outputTensor[0][5][yP][xP]) }else{ FilteredYPic[ySrc*2][xSrc*2]=OutY(outputTensor[0][yP][xP][0]) FilteredYPic[ySrc*2][xSrc*2+1]=OutY(outputTensor[0][yP][xP][1]) FilteredYPic[ySrc*2+1][xSrc*2]=OutY(outputTensor[0][yP][xP][2]) FilteredYPic[ySrc*2+1][xSrc*2+1]=OutY(outputTensor[0][yP][xP][3]) FilteredCbPic[ySrc][xSrc]=OutC(outputTensor[0][yP][xP][4]) FilteredCrPic[ySrc][xSrc]=OutC(outputTensor[0][yP][xP][5]) } } The basic postfiltering process for a decoded image is performed by the filter specified in the first Neural Network Postfilter Characteristics SEI message in decoding order with a particular nnpfc_id value in the CLVS. It is a filter.

[0194] Another neural network with the same nnpfc_id value but with an nnpfc_mode_idc value of 1 Post-filter characteristics The basic post-filter processing when a SEI message is present is The neural network postfilter characteristics are updated by decoding the ISO / IEC 15938-17 bitstream in the SEI message.

[0195] Otherwise, the postfilter PostProcessingFilter() is assigned the same as the base postfilter.

[0196] According to the value of nnpfc_out_order_idc, the decoded image is filtered by a post-filtering process PostProcessingFilter() to generate Y, Cb, and Cr pixel arrays FilteredYPic, FilteredCbPic, and FilteredCrPic as follows.

[0197] The post-filtering process PostProcessingFilter() of the decoded image is as follows.

[0198] if(nnpfc_inp_order_idc==0) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==1) for(cTop=0;cTop <InpPicHeightInLumaSamples / InpSubHeightC; cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples / InpSubWidthC; cLeft+=inpPatchWidth) { DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==2) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==3) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight*2) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth*2){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } Neural network post-filter characteristic SEI message with different contents for the same nnpfc_id If both neural network postfilter characteristics SEI messages are present in the same picture unit, then both neural network postfilter characteristics SEI messages shall be present in the same SEI NAL unit.

[0199] These SEI messages can be used as post-filtering in neural networks. When nnpfc_out_order_idc is specified, the semantics specify the derivation of the luminance pixel array FilteredYPic and the chrominance pixel arrays FilteredCbPic, FilteredCrPic as indicated by the value of nnpfc_out_order_idc that contain the output of the post-filtering process.

[0200] The syntax element nnpfc_reserved_zero_bit shall be equal to 0.

[0201] The syntax element nnpfc_payload_byte[i] contains the i-th byte of an ISO / IEC 15938-17 compliant bitstream. All nnpfc_payload_byte[i] must be fully ISO / IEC 15938-17 compliant bitstreams.

[0202] FIG. 10 shows the syntax of the neural network complexity element nnpfc_complexity_element (nnpfc_complexity_idc) of Non-Patent Document 1. The argument nnpfc_complexity_idc is a value indicating that a syntax element indicating the complexity of the neural network postfilter process associated with nnpfc_id may exist. nnpfc_id includes an identification number that can be used to identify the postfilter process. nnpfc_complexity_element (nnpfc_complexity_idc) is called from the neural network postfilter characteristics SEI message shown in FIG. 9. The neural network postfilter characteristics SEI message specifies the neural network that can be used as the postfilter process. The use of the specified postfilter process for a particular image is indicated in the neural network postfilter activation SEI message shown in FIG. 12.

[0203] Below, the synopsis of nnpfc_complexity_element (nnpfc_complexity_idc) in Non-Patent Document 1 in FIG. Describes a task element.

[0204] A value of 0 for the syntax element nnpfc_parameter_type_flag indicates that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point or integer parameters.

[0205] The syntax element nnpfc_log2_parameter_bit_length_minus3 has values ​​0, 1, 2, and 3 for neural networks larger than 8 bits, 16 bits, 32 bits, and 64 bits, respectively. Indicates that the bit length parameter is not used.

[0206] The syntax element nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for postfiltering, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0, indicates that the maximum number of neural network parameters is not specified. If the value of nnpfc_num_parameters_idc is greater than zero, the variable maxNumParameters is derived as follows:

[0207] maxNumParameters=(2048< <nnpfc_num_parameters_idc)-1 A bitstream compliance requirement is that the number of post-filtering neural network parameters must be less than or equal to maxNumParameters.

[0208] The syntax element nnpfc_num_kmac_operations_idc specifies the number of postfilter operations per pixel. This is a syntax element that indicates the maximum number of multiplication and accumulation operations of nnpfc_num_kmac_operations_idc. A value greater than 0 indicates that the maximum number of per-pixel multiply-accumulate operations for postfiltering is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of network multiply-accumulate operations is unspecified.

[0209] In the specification of the neural network post-filter characteristic SEI in Non-Patent Document 1, There was an issue that the variable maxNumParameters, which indicates the maximum value of network parameters, could overflow with a 32-bit integer or a 64-bit integer value.

[0210] nnpfc_num_parameters_idc is defined as u(8), i.e., 8 bits of a binary number, and can take values ​​between 0 and 255. Using this syntax element, the maximum number of neural network parameters for post-filtering is expressed in powers of 2048, so it can represent a maximum of ((2048<<255)-1). However, this value far exceeds the range of 32-bit integers or 64-bit integers, which are the common ways of representing integers in hardware and software.

[0211] Therefore, when the variable maxNumParameters is a non - negative 32 - bit integer, in order not to overflow during intermediate operations, maxNumParameters must be less than or equal to (2048 << 20) - 1. Therefore, set Description to u(5) and the value range of nnpfc_num_parameters_idc to be from 0 to 20. Alternatively, when the variable maxNumParameters is a non - negative 64 - bit integer, in order not to overflow during intermediate operations, maxNumParameters must be less than or equal to (2048 << 52) - 1. Therefore, set Description to u(6) and the value range of nnpfc_num_parameters_idc to be from 0 to 52. Here, u(x) indicates that the syntax is an unsigned integer and the binary representation has a fixed length of x bits. That is, the value range is from 0 to (1 << x) - 1.

[0212] Alternatively, when the variable maxNumParameters is a non - negative 16 - bit integer, in order not to overflow during intermediate operations, maxNumParameters must be less than or equal to (2048 << 4) - 1. Therefore, use the binary representation indicated by u(3) for Description and set the value range of nnpfc_num_parameters_idc to be from 0 to 4.

[0213] When the variable maxNumParameters is a signed 32 - bit integer, in order not to overflow during intermediate operations, maxNumParameters must be less than or equal to (2048 << 19) - 1. Therefore, use the binary representation of u(5) for Description and set the value range of nnpfc_num_parameters_idc to be from 0 to 19.

[0214] Alternatively, if the variable maxNumParameters is a signed 64-bit integer, it must be equal to or less than maxNumParameters = ( 2048 << 51 ) - 1 to prevent overflow during intermediate calculations. Therefore, use the binarization of Description as u(6) to set the value range of nnpfc_num_parameters_idc to 0 to 51 inclusive. Alternatively, if the variable maxNumParameters is a signed 16-bit integer, it must be equal to or less than maxNumParameters = ( 2048 << 3 ) - 1. Therefore, set Description to u(2) and set the value range of nnpfc_num_parameters_idc to 0 to 3 inclusive.

[0215] The value is not limited to 2048, and the following values ​​may be used. maxNumParameters = 1024< <nnpfc_num_parameters_idc nnpfc_num_parameters_idc is the binarization of u(6). maxNumParameters = 512< <nnpfc_num_parameters_idc nnpfc_num_parameters_idc is the binarization of u(7). That is, it is expressed as follows. maxNumParameters = (1< <N_UNIT)<<nnpfc_num_parameters_idc nnpfc_num_parameters_idc may be the binarization of u(16-N_UNIT), where N_UNIT is an integer between 0 and 11 inclusive.

[0216] Defining binarization in this way solves the problem of the variable maxNumParameters, which indicates the maximum value of the network parameters, overflowing.

[0217] In the above-described embodiment, since the maximum value of the network parameters is defined by shifting the value 2048 used in Non-Patent Document 1, a value range was obtained on the premise that integer arithmetic does not overflow in the derivation of the variable maxNumParameters. However, the same method can be applied to other values than 2048. As another solution, instead of the value 2048, for example, a syntax element named nnpfc_num_parameters_base_value may be added before nnpfc_num_parameters_idc in FIG. 9 and defined by the following formula.

[0218] Let the Description of nnpfc_num_parameters_base_value be u(n), where n is a natural number, and let the variable maxNumParameters be a non-negative 32-bit integer. Let the Description of nnpfc_num_parameters_idc be u(m), where m is

[0219] maxNumParameters=(nnpfc_num_parameters_base_value<<nnpfc_num_parameters_idc)-1 a natural number. Then, the value of (n + m) should be 31 or less. When the variable maxNumParameters is a non-negative 64-bit integer, the value of (n + m) should be 63 or less. When it is a non-negative 16-bit integer, the value of (n + m) should be 15 or less. When it is a signed 32-bit integer, the value of (n + m) should be 30 or less. When it is a signed 64-bit integer, the value of (n + m) should be 61 or less. When it is a signed 16-bit integer, the value of (n + m) should be 14 or less.

[0220] In the specification of the neural network post-filtering characteristic SEI of Non-Patent Document 1, the post There was an issue that the upper limit of nnpfc_num_kmac_operations_idc, which is a syntax element that indicates the maximum number of multiplication and accumulation operations per pixel in the filter processing, was not defined. Therefore, there was an issue that the internal variable may overflow, for example, with a 32-bit integer or a 64-bit integer value.

[0221] Therefore, set the upper limit of nnpfc_num_kmac_operations_idc to an integer range. For example, The upper limit of a non-negative 32-bit integer is set to 2^32-1. It may also be set to 2^64-1 for the upper limit of a non-negative 64-bit integer. Also, it may be set to 2^16-1 for the upper limit of a non-negative 16-bit integer.

[0222] If signed integers are considered, the upper limit of a signed 32-bit integer is 2^31-1, or Alternatively, it may be set to 2^63 - 1, which is the upper limit of a signed 64-bit integer, or 2^15 - 1, which is the upper limit of a signed 16-bit integer. In either case, an upper limit must be determined.

[0223] This definition solves the problem of overflow of nnpfc_num_kmac_operations_idc, which is a syntax element that indicates the maximum number of multiplication and accumulation operations per pixel in post-filter processing. It can be solved.

[0224] (Neural Network Post Filter Activation SEI) FIG. 12 is a diagram illustrating the syntax of the neural network post-filter activation SEI message in Non-Patent Document 1.

[0225] This SEI message is a neural network that can be used to post-filter the current decoded image. Specifies network postfiltering. The neural network postfilter activation SEI message applies only to the current decoded image.

[0226] For example, there may be multiple neural network post-filter activation SEI messages for the same decoded image if the post-filter processing targets different purposes or filters different color components.

[0227] The syntax element nnpfa_id relates to the current decoded image and is a 1 where nnpfc_id is equal to nnfpa_id. One or more neural network postfilter characteristics specified by the SEI Specifies that network postfiltering can be used for postfiltering the current decoded picture.

[0228] In Non-Patent Document 1, the neural network post-filter characteristic SEI is applied for each CVS. The neural network postfilter activation SEI is applied to each decoded image, but it has a problem in that it cannot accommodate changes in the width and height of the decoded image or the quantization parameter for each decoded image.

[0229] Therefore, in the neural network post-filter activation SEI, The neural network postfiltering used for the postfiltering of the model is specified by the SEI index nnfpa_id, which indicates the information of the postfiltering. In addition to identifying the work post-filter characteristic SEI, the current decoded image Specify the width and height of the image and the quantization parameters.

[0230] Specifically, the variable InpPicWidthInLumaSamples is set equal to pps_pic_width_in_luma_samples-SubWidthC*(pps_conf_win_left_offset+pps_conf_win_right_offset) of the current decoded image, and the variable InpPicHeightInLumaSamples is set equal to pps_pic_height_in_luma_samples-SubHeightC*(pps_conf_win_top_offset+pps_conf_win_bottom_offset) of the current decoded image. is set equal to

[0231] The current decoded image is represented by a two-dimensional array of luminance pixels, CroppedYPic[y][x], at vertical coordinate y and horizontal coordinate x, and two-dimensional arrays of chrominance pixels, CroppedCbPic[y][x] and CroppedCrPic[y][x]. Let the y coordinates of the upper left corner of the pixel array be 0 and x coordinates be 0.

[0232] The pixel bit length of the luminance of the decoded image is BitDepthY. The pixel bit length of the chrominance of the decoded image is BitDepthC. Note that both BitDepthY and BitDepthC are set equal to BitDepth.

[0233] The variable InpSubWidthC is the chrominance subsampling ratio for the horizontal luminance of the decoded image, and the variable InpSubHeightC is the chrominance subsampling ratio for the horizontal luminance of the decoded face image. Note that InpSubWidthC is set equal to the variable SubWidthC of the encoded data. InpSubHeightC is set equal to the variable SubHeightC of the encoded data.

[0234] The variable SliceQPY is set equal to the quantization parameter SliceQpY updated at the slice level of the encoded data, where the slice is the quantization parameter of the first slice in the decoded image.

[0235] In this way, in the neural network post-filter activation SEI, Using the information of the current decoded image, the neural network post-filter characteristic SEI is input. By updating the value to be used, appropriate post-filtering becomes possible, thereby resolving the issue.

[0236] (SEI payload) FIG. 13 shows the syntax of the SEI payload, which is a container of the SEI message in Non-Patent Document 1. FIG.

[0237] It is called when nal_unit_type is PREFIX_SEI_NUT. PREFIX_SEI_NUT is the slice data. This indicates that the SEI is located before the data.

[0238] When payloadType is 210, the neural network post-filter characteristic SEI is called. do.

[0239] When payloadType is 211, the Neural Network Post Filter Activation SEI is called. It is called out.

[0240] (SEI decoding and post-filtering) The header decoding unit 3020 reads the SEI payload, which is a container of the SEI message, and decodes the neural network post-filter characteristics SEI message.

[0241] 14 is a diagram showing a flowchart of the process of the NN filter unit 611. The NN filter unit 611 performs the following process in accordance with the parameters of the above SEI message.

[0242] S6001: Read the processing volume and accuracy from the neural network complexity factor.

[0243] S6002: End if the complexity exceeds the processable level of the NN filter unit 611. If not, proceed to S6003.

[0244] S6003: End if the accuracy exceeds the processable accuracy of the NN filter unit 611. If not, proceed to S6004.

[0245] S6004: Identify the network model from the SEI and set the topology of the NN filter unit 611. do.

[0246] S6005: Derive network model parameters from the updated information of SEI.

[0247] S6006: The parameters of the derived network model are read into the NN filter unit 611.

[0248] S6007: The NN filter unit 611 executes the filtering process and outputs the result to the outside.

[0249] However, the SEI is not necessarily required to construct the luma and chroma samples in the decoding process. It won't be done.

[0250] (Example of the configuration of the NN filter unit 611) FIG. 15 shows an interpolation filter using a neural network filter unit (NN filter unit 611). 1 is a diagram illustrating an example of the configuration of a post filter, an interpolation filter, a loop filter, and a post filter.

[0251] The post-processing unit 61 after the video decoding device includes an NN filter unit 611. When the image in the reference picture memory 306 is output, a filter process is performed and the image is output to the outside. The image may be displayed, written to a file, re-encoded (transcoded), transmitted, etc. The NN filter unit 611 filters the input image using a neural network model. At the same time, it may also perform reduction or enlargement at a constant or rational number factor.

[0252] Here, the neural network model (hereafter referred to as NN model) means the elements and connections (topology) of the neural network, and the parameters (weights, biases) of the neural network. Note that the topology may be fixed and only the parameters of the neural network model may be switched.

[0253] (Details of the NN filter unit 611) The NN filter uses the input image inputTensor and input parameters (e.g., QP, bS, etc.) Then, filtering is performed using a neural network model. The input image may be an image for each component, or an image with multiple components as channels. The input parameters may be assigned to a channel different from the image.

[0254] The NN filter unit may repeatedly apply the following process.

[0255] The NN filter part derives the output image outputTensor by convolving the inputTensor with the kernel k[m][i][j] (conv, convolution) and adding bias. Here, nn=0..n-1, xx=0..width-1, yy=0..height-1, and Σ represents the summation for mm, i, and j, respectively.

[0256] outputTensor[nn][xx][yy]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][xx+i-of][yy+j-of]+bias[nn]) In the case of 1x1 Conv, Σ represents the sum of mm=0..m-1, i=0, j=0. In this case, of=0 is set. In the case of 3x3 Conv, Σ represents the sum of mm=0..m-1, i=0..2, j=0..2. , set of=1. n is the number of channels of outSamples, m is the number of channels of inputTensor, width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor. of is the size of the padding area placed around the inputTensor to make the sizes of inputTensor and outputTensor the same. In the following, when the output of the NN filter section is a value (corrected value) rather than an image, the output will be represented as corrNN instead of outputTensor.

[0257] Note that if you write inputTensor and outputTensor in CHW format instead of CWH format, it is equivalent to the following process.

[0258] outputTensor[nn][yy][xx]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][yy+j-of][xx+i-of]+bias[nn]) In addition, a process called Depth-wise Conv, shown in the following formula, may be performed. Here, nn = 0..n-1, xx = 0..width-1, yy = 0..height-1, and Σ represents the summation for i and j, respectively. n is the number of channels of outputTensor and inputTensor, width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor.

[0259] outputTensor[nn][xx][yy]=ΣΣ(k[nn][i][j]*inputTensor[nn][xx+i-of][yy+j-of]+bias[nn]) Also, a nonlinear process called Activate, for example, ReLU, may be used. ReLU(x) = x >= 0 ? x : 0 Alternatively, leakyReLU shown in the following formula may be used.

[0260] leakyReLU(x) = x >= 0 ? x : a * x Here, a is a predetermined value, for example, 0.1 or 0.125. In order to perform integer arithmetic, all the values ​​of k, bias, and a above may be integers, and a right shift may be performed after conv.

[0261] In ReLU, 0 is always output for values ​​less than 0, and the input value is output as is for values ​​greater than or equal to 0. On the other hand, in leakyReLU, linear processing is performed with the gradient set by a for values ​​less than 0. In ReLU, the gradient for values ​​less than 0 disappears, so learning may not progress smoothly. In leakyReLU, the gradient for values ​​less than 0 is left, making the above problem less likely to occur. In addition, among the above leakyReLU(x), PReLU, which uses a parameterized value of a, may be used.

[0262] (NNR) Neural Network Coding and Representation (NNR) is a method for representing neural networks (NNs) This is the international standard ISO / IEC15938-17 for efficient compression. This makes it possible to store and transmit NNs more efficiently.

[0263] The following provides an overview of the NNR encoding and decoding processes.

[0264] FIG. 16 is a diagram showing an NNR encoding device and a decoding device.

[0265] The NN coding device 801 includes a preprocessing unit 8011, a quantization unit 8012, and an entropy coding unit 8013. The NN encoding device 801 receives the uncompressed NN model O, and the quantization unit 8012 quantizes the NN model O. In the NN coding device 801, before quantization, a pre-processing unit 8011 may repeatedly apply a parameter reduction method such as pruning or sparsification. After that, an entropy coding unit 8013 performs entropy coding on the quantization model Q. This is applied to obtain a bitstream S for storing and transmitting the NN model.

[0266] The NN decoding device 802 includes an entropy decoding unit 8021, a parameter restoration unit 8022, and a post-processing unit 8023. The NN decoding device 802 first receives the transmitted bit stream S, and the entropy decoding unit 8021 performs entropy decoding of S to obtain an intermediate model RQ. Operation of the NN model If the environment supports inference using the quantized representation used in the RQ, the RQ may be output and used for inference. If not, the parameters of the RQ are restored to their original representation in the parameter restoration unit 8022, and an intermediate model RP is obtained. If the sparse tensor representation used can be processed in the operating environment of the NN model, the RP may be output and used for inference. If not, the NN model O and A reconstruction NN model R that does not contain different tensors or structural representations is obtained and output.

[0267] The NNR standard specifies the decoding procedures for the numeric representation of specific NN parameters, such as integers and floating point. There is law.

[0268] The decoding method NNR_PT_INT decodes a model with integer-valued parameters. The decoding method NNR_PT_FLOAT extends NNR_PT_INT by adding a quantization step size delta, which is multiplied by the integer value to produce a scaled integer. Delta is derived from the integer quantization parameter qp and the granularity parameter qp_density of delta as follows:

[0269] mul = 2^(qp_density) + (qp & (2^(qp_density)-1)) delta = mul * 2^((qp >> qp_density)-qp_density) (Format of trained NN) The representation of a trained NN consists of two elements: a topological representation, such as the size of layers and the connections between layers, and a parameter representation, such as weights and biases.

[0270] Topological representation is covered in native formats such as Tensorflow and PyTorch. However, to improve interoperability, Open Neural Network Exchange Format (ONNX), Neura l There are exchange formats such as Network Exchange Format (NNEF).

[0271] The NNR standard also transports topology information, nnr_topology_unit_payload, as part of the NNR bitstream that contains the compressed parameter tensors. This allows for the exchange of This enables interoperability not only with existing maps but also with topology information expressed in native formats.

[0272] (Configuration of Image Encoding Device) Next, the configuration of the image encoding device 11 according to this embodiment will be described. 1 is a block diagram showing a configuration of an image encoding device 11 according to the present embodiment. The image encoding device 11 includes a predicted image generating unit 101, a subtraction unit 102, a transform / quantization unit 103, an inverse quantization / inverse transform unit 105, an addition unit 106, a rule The image processing system includes a loop filter 107, a prediction parameter memory (prediction parameter storage unit, frame memory) 108, a reference picture memory (reference image storage unit, frame memory) 109, an encoding parameter determination unit 110, a parameter encoding unit 111, a prediction parameter derivation unit 120, and an entropy encoding unit 104.

[0273] The predicted image generating unit 101 generates a predicted image for each CU.

[0274] The subtraction unit 102 generates a prediction error by subtracting the pixel values ​​of the predicted image of the block input from the predicted image generation unit 101 from the pixel values ​​of the image T. The subtraction unit 102 outputs the prediction error to the transformation and quantization unit 103.

[0275] The transform / quantization unit 103 calculates transform coefficients by frequency transforming the prediction error input from the subtraction unit 102, and derives quantized transform coefficients by quantizing the prediction error. The quantized transform coefficients are output to the parameter coding unit 111 and the inverse quantization and inverse transform unit 105 .

[0276] The inverse quantization and inverse transform unit 105 corresponds to the inverse quantization and inverse transform unit 311 (FIG. 5) in the image decoding device 31. The calculated prediction error is output to the adder 106.

[0277] The parameter coding unit 111 includes a header coding unit 1110, a CT information coding unit 1111, and a CU coding unit 1112 (prediction mode coding unit). The CU coding unit 1112 further includes a TU coding unit 1114. The following describes an outline of the operation of each module.

[0278] The header encoding unit 1110 performs encoding processing of parameters such as header information, division information, prediction information, and quantized transform coefficients.

[0279] The CT information encoding unit 1111 encodes the QT, MT (BT, TT) division information and the like.

[0280] The CU encoding unit 1112 encodes the CU information, prediction information, division information, and so on.

[0281] When a prediction error is included in a TU, the TU encoding unit 1114 encodes the QP update information and the quantized prediction error.

[0282] The CT information encoding unit 1111 and the CU encoding unit 1112 supply syntax elements such as inter prediction parameters and quantized transform coefficients to the parameter encoding unit 111.

[0283] The entropy coding unit 104 receives the quantized transform coefficients and the coding parameters from the parameter coding unit 111. The entropy coding unit 104 entropy codes these to encode them. The encoded data Te is generated and output.

[0284] The prediction parameter derivation unit 120 derives inter-prediction parameters and intra-prediction parameters from the parameters input from the encoding parameter determination unit 110. The derived inter-prediction parameters and intra-prediction parameters are output to the parameter encoding unit 111. will be done.

[0285] The adder 106 generates a decoded image by adding, for each pixel, the pixel value of the predicted block input from the predicted image generation unit 101 and the prediction error input from the inverse quantization and inverse transform unit 105. The adder 106 stores the generated decoded image in a reference picture memory 109.

[0286] The loop filter 107 performs deblocking filtering, SAO, and ALF on the decoded image generated by the adder 106. Note that the loop filter 107 does not necessarily include the above three types of filters. For example, the filter may be configured with only a deblocking filter.

[0287] The prediction parameter memory 108 stores the prediction parameters generated by the coding parameter determination unit 110 in a predetermined location for each current picture and CU.

[0288] The reference picture memory 109 stores the decoded image generated by the loop filter 107 at a predetermined position for each current picture and CU.

[0289] The encoding parameter determination unit 110 determines one of the multiple sets of encoding parameters. The coding parameters are the above-mentioned QT, BT or TT division information, prediction parameters, or parameters to be coded that are generated in relation to these. The predicted image generating unit 101 generates a predicted image using these coding parameters.

[0290] In addition, a part of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment, for example, the entropy decoding unit 301, the parameter decoding unit 302, the loop filter 305, the predicted image generating unit 308, the inverse quantization and inverse transform unit 311, the addition unit 312, the prediction parameter derivation unit 320, the predicted image generating unit 101, the subtraction unit 102, the transform and quantization unit 103, the entropy coding unit 104, the inverse quantization and inverse transform unit 105, the loop filter 107, the coding parameter determination unit 110, the parameter coding unit 111, and the prediction parameter derivation unit 120 may be realized by a computer. In this case, a program for realizing this control function may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into and executed by a computer system. In addition, the "computer system" referred to here is a computer system built into either the image encoding device 11 or the image decoding device 31, and includes hardware such as an OS and peripheral devices. In addition, "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. Furthermore, "computer-readable recording medium" may also include devices that dynamically hold a program for a short period of time, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and devices that hold a program for a certain period of time, such as volatile memory inside a computer system that serves as a server or client in such cases. Furthermore, the above program may be one that realizes part of the above-mentioned functions, or may be one that can realize the above-mentioned functions in combination with a program already recorded in the computer system.

[0291] In addition, a part or the whole of the image encoding device 11 and the image decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the image encoding device 11 and the image decoding device 31 may be individually implemented as a processor, or may be integrated into a processor in part or in whole. The method of implementing the integrated circuit is not limited to LSI. It may be realized by a dedicated circuit or a general-purpose processor. Also, when an integrated circuit technology that can replace LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0292] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.

[0293] The present embodiment will be described with reference to FIG. 1. and a resolution inverse conversion device using a neural network that converts the decoded image to a specified resolution using inverse conversion information, wherein the resolution inverse conversion device decodes information indicating the number of horizontal and vertical pixels of a patch size, which is the unit of processing by the neural network, and sets the maximum values ​​of the number of horizontal and vertical pixels of the patch size as the number of horizontal and vertical pixels of the decoded image.

[0294] The video decoding device also includes an image decoding device that decodes encoded data to generate a decoded image, and a resolution inverse conversion device that uses a neural network to convert the decoded image to a specified resolution using inverse conversion information, and is characterized in that the resolution inverse conversion device decodes information indicating the maximum value of a network parameter of the neural network, and when deriving the maximum value of the network parameter, sets the maximum value to a value that does not overflow in 32-bit integer arithmetic.

[0295] The video coding device also includes an image coding device that codes an image to generate coded data, an inverse conversion information generating device that generates inverse conversion information for inversely converting the resolution of a decoded image when the coded data is decoded, and an inverse conversion information coding device that codes the inverse conversion information as auxiliary extension information, wherein the resolution inverse conversion information coding device codes information indicating the number of horizontal pixels and the number of vertical pixels of a patch size, which is a unit of processing in a neural network, and the maximum values ​​of the number of horizontal pixels and the number of vertical pixels of the patch size are set to the number of horizontal pixels and the number of vertical pixels of the coded image.

[0296] The video coding device also includes an image coding device that codes an image to generate coded data, an inverse conversion information generating device that generates inverse conversion information for inversely converting the resolution of a decoded image when the coded data is decoded, and an inverse conversion information coding device that codes the inverse conversion information as auxiliary extension information, and is characterized in that information indicating the maximum value of a network parameter of a neural network in the resolution inverse conversion information coding device is coded, and when deriving the maximum value of the network parameter, a maximum value that does not overflow in 64-bit integer arithmetic is set.

[0297] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]

[0298] The embodiments of the present invention can be suitably applied to a video decoding device that decodes coded data in which image data is coded, and a video coding device that generates coded data in which image data is coded, and can also be suitably applied to the data structure of coded data that is generated by a video coding device and referenced by the video decoding device. [Explanation of symbols]

[0299] 1. Video transmission system 30 Video Decoding Device 31 Image Decoding Device 301 Entropy Decoding Unit 302 Parameter Decoding Unit 305, 107 Loop Filter 306, 109 Reference Picture Memory 307, 108 Prediction parameter memory 308, 101 Prediction image generation unit 311, 105 Inverse quantization and inverse transformation unit 312, 106 Addition section 320 Prediction parameter derivation part 10 Video Encoding Device 11 Image encoding device 102 Subtraction section 103 Transformation and Quantization Section 104 Entropy coding unit 110 Encoding parameter determination unit 111 Parameter Encoding Unit 120 Prediction parameter derivation part 71 Reverse conversion information creation device 81 Inverse transformation information coding device 91 Inverse transformation information decoding device 611 NN filter section

Claims

1. An image decoding device that decodes encoded data to generate a decoded image, and a resolution inverse conversion device using a neural network that converts the decoded image to a specified resolution using inverse conversion information, wherein the resolution inverse conversion device decodes first information indicating the maximum value of the network parameters of the neural network, and the first information is a value of 0 or more and 52 or less, and the moving image decoding device is characterized by this.

2. An image encoding device that encodes an image to generate encoded data, an inverse conversion information creation device that creates inverse conversion information for inverse-converting the resolution of the decoded image when the encoded data is decoded, and an inverse conversion information encoding device that encodes auxiliary extension information including the inverse conversion information, wherein the inverse conversion information encoding device encodes first information indicating the maximum value of the network parameters of the neural network, and the first information is a value of 0 or more and 52 or less, and the moving image encoding device is characterized by this.

3. The moving image encoding device according to claim 2, wherein the first information is the maximum value that does not overflow when deriving the maximum value of the network parameters.

4. The moving image encoding device according to claim 3, wherein the maximum value that does not overflow is the maximum value in the case of integer arithmetic with 64 bits.

5. The moving image encoding device according to any one of claims 2 to 4, wherein the first information is represented by 6 bits.

6. A computer-readable recording medium that records a program for causing a computer to execute steps of encoding an image to generate encoded data, creating inverse conversion information for inverse-converting the resolution of the decoded image when the encoded data is decoded, encoding first information indicating the maximum value of the network parameters of the neural network, and encoding auxiliary extension information including the inverse conversion information, wherein the first information is a value of 0 or more and 52 or less, and the computer-readable recording medium is characterized by this.