Moving image encoding device, moving image decoding device and program
By transmitting hidden state information in SEI messages, the video encoding and decoding devices facilitate improved image quality through neural network-based super-resolution, addressing the limitations of current video coding standards.
Patent Information
- Application Number
- JP2024091658
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-12-17
AI Technical Summary
Current video coding standards do not allow transmission of hidden state information in Supplemental Enhancement Information (SEI) messages, limiting the ability of video decoding devices to utilize neural network models for super-resolution and improve image quality.
A video encoding device infers super-resolution decoded video and transmits auxiliary enhancement information including syntax elements related to hidden states, while a video decoding device receives and utilizes this information to infer super-resolution decoded video using a learning model.
Enables the utilization of hidden states in video decoding, enhancing image quality through neural network-based super-resolution processing.
Smart Images

Figure 2025183790000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a video encoding device, a video decoding device, and a program. [Background technology]
[0002] There is a technology called super resolution. Super resolution is a technology that increases the resolution of an image. Super resolution makes it possible to obtain a high-resolution output image from a low-resolution input image, thereby improving the image quality of the input image.
[0003] Furthermore, with the recent advances in neural networks, super-resolution using a convolutional neural network (CNN) is gaining attention. Generally, the amount of processing required by a neural network is enormous. Therefore, it may be difficult to perform super-resolution in real time using a neural network. Therefore, a super-resolution technique using a neural network has been proposed in which multiple frames (sliding window-based) that are time-sequential are input, and a hidden state is input recursively in the time direction (see Non-Patent Document 1 below). Non-Patent Document 1 shows that image quality can be improved by utilizing the hidden state.
[0004] Meanwhile, the Joint Video Experts Team (JVET) is currently discussing the use of Supplemental Enhancement Information (SEI) messages in the video coding standard VVC (Versatile Video Coding) (ITU-T H266) to transmit neural-network post-filters (NNPFs) (see Non-Patent Document 2 below). An SEI message is a message that provides specific types of information to assist with processing related to decoding, display, or the like. The SEI message is included in a bitstream along with encoded data and transmitted from a video encoder to a video decoder. An NNPF is a neural network that obtains super-resolution decoded video from decoded video. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Wenyi Lian, Wenjing Lian, “Sliding Window Recurrent Network for Efficient Video”, 2022 / 8 / 24, arXiv:2208.11608. [Non-patent document 2] Sean McCarthy, Takeshi Chujoh, Miska M. Hannuksela, Gary J. Sullivan, Ye-Kui Wang, “Additional SEI messages for VSEI (Draft 5)”, 2023 / 7 / 11-19, JVET-AE2006-v2 Summary of the Invention [Problem to be solved by the invention]
[0006] As mentioned above, in super-resolution techniques using neural networks, the quality of the output image can be improved by utilizing hidden states. However, the current standardization does not allow information about hidden states to be transmitted in the SEI message.
[0007] Therefore, even if a video decoding device can construct a neural network model using SEI messages, it may not be able to use hidden states in the neural network model, and may not be able to improve the image quality of the output video.
[0008] Therefore, an object of the present disclosure is to provide a video encoding device, a video decoding device, and a program that are capable of utilizing hidden states in a video decoding device. [Means for solving the problem]
[0009] A video encoding device according to a first aspect is a learning model that infers super-resolution decoded video having a different resolution from decoded video obtained by decoding encoded video, and has a transmitting unit that transmits to a video decoding device an auxiliary enhancement information message including syntax elements related to hidden states that can be used in the learning model.
[0010] A video decoding device according to a second aspect includes a receiving unit that receives, from a video encoding device, a supplemental enhancement information message including a learning model that infers, from decoded video obtained by decoding encoded video, super-resolution decoded video having a different resolution from the decoded video, the supplemental enhancement information message including syntax elements related to hidden states that can be used by the learning model. The video decoding device also includes a neural network inference unit that infers the super-resolution decoded video from the decoded video using the learning model. The neural network inference unit infers the super-resolution decoded video using the hidden states.
[0011] A program according to a third aspect is a program that causes a computer to function as the video encoding device according to the first aspect or the video decoding device according to the second aspect. [Effects of the Invention]
[0012] According to the present disclosure, it is possible to provide a video encoding device, a video decoding device, and a program that are capable of utilizing hidden states in a video decoding device. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a video transmission system according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of a VVC encoder according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of the configuration of a bitstream according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing types of Non-VCL NALUs and their descriptions according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of an SEI message according to the first embodiment. [Figure 6] FIG. 6 shows an example of syntax elements according to the first embodiment. [Figure 7] FIG. 7 shows an example of syntax elements according to the first embodiment. [Figure 8] FIG. 8 shows an example of syntax elements according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of the configuration of an NNR encoder according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of the configuration of a video encoding device and a video decoding device according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of a neural network learning device according to the first embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the configuration of a neural network model according to the first embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of the configuration of a residual block (Resblock) according to the first embodiment. [Figure 14]FIG. 14 is a diagram illustrating an example of a syntax structure according to the first embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of the configuration of a neural network inference device according to the first embodiment. [Figure 16] FIG. 16 is a diagram illustrating an example of operation according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] [First embodiment] A moving image transmission system according to an embodiment will be described with reference to the drawings. In the description of the drawings, the same or similar parts are denoted by the same or similar reference numerals.
[0015] (Moving image transmission system according to the first embodiment) First, a video transmission system according to the first embodiment will be described.
[0016] FIG. 1 is a diagram illustrating an example of the configuration of a video transmission system 10 according to the first embodiment.
[0017] As shown in FIG. 1, a video transmission system 10 includes a video encoding device 100 and a video decoding device 200.
[0018] The video encoding device 100 encodes input video and obtains encoded video. The video encoding method may be, for example, VVC, or other video encoding methods such as High Efficiency Video Coding (HEVC) (H265) or Advanced Video Coding (AVC) (H.264 / MPEG-4 AVC). The video data of the encoded video is transmitted to the video decoding device 200 as a bitstream.
[0019] The video encoding device 100 according to the first embodiment generates a neural network model. The neural network model is used for super-resolution technology. That is, the neural network model receives video (decoded video of encoded video) as input and can infer super-resolution output video (super-resolution decoded video).
[0020] Furthermore, the video encoding device 100 according to the first embodiment can compress the generated neural network model itself, and can transmit the compressed neural network model together with the video data of the encoded video in a bitstream to the video decoding device 200.
[0021] The video decoding device 200 receives the bitstream transmitted from the video encoding device 100 and extracts the video data of the encoded video and the encoding neural network from the bitstream. The video decoding device 200 decodes the encoded video to obtain the decoded video. The video decoding method may be any decoding method that corresponds to the video encoding method used in the video encoding device 100, and may be, for example, VVC.
[0022] The video decoding device 200 according to the first embodiment decodes a compressed neural network model, and can output super-resolution decoded video from the decoded video using the decoded neural network model.
[0023] (VVC) Next, a description will be given of a video coding method used in the video coding device 100. Here, VVC will be described as a representative example.
[0024] FIG. 2 is a diagram illustrating an example of the configuration of a VVC encoder.
[0025] ·Screen split Similar to HEVC, VVC uses coding tree units (CTUs) for screen division. A CTU consists of a luminance signal CTU and a chrominance signal CTU block. The size of a CTU in VVC can be expanded up to a maximum of 128 x 128 pixels. Intra-prediction, inter-prediction, integer transform, quantization, and other processes are performed in units of coding units (CUs). A CU is a block obtained by recursively dividing a CTU into a tree structure. In addition to quadtrees, ternary and binary trees can also be used as CTU division patterns. VVC can divide a CTU into CUs of a predetermined size and shape according to the local characteristics of the video, thereby improving coding efficiency to a certain extent. Note that prediction units (PUs) and transform units (TUs) are basically the same as CUs, and CUs are used for both prediction and transformation.
[0026] Intra prediction Intra prediction is a prediction process within the same picture. As with HEVC, VVC intra prediction performs intra-picture prediction (spatial prediction) for each prediction block using decoded pixels (prediction reference pixels) near the prediction block, and performs integer conversion and quantization on the prediction difference signal. Compared to HEVC, VVC's prediction modes have been expanded to 67 types, allowing selection of modes with excellent prediction efficiency. For example, VVC newly introduces PDPC (Position Dependent intra Prediction Combination) as a prediction mode, which updates a predicted image generated by normal intra prediction using surrounding prediction reference pixels selected for each pixel position.
[0027] Inter prediction Inter prediction is a prediction process between different pictures (in the temporal direction or between layer images). In VVC inter prediction, as with HEVC, a predicted image is generated by performing motion compensation using multiple reference images stored in a frame memory for each prediction block. A predicted difference signal between the predicted image and the input image is subjected to integer transformation and quantization. As with HEVC, the prediction modes used for inter prediction are basically AMVP (Adaptive Motion Vector Prediction) mode and merge mode. In AMVP mode, a differential motion vector is coded relative to a predicted motion vector derived from a decoded block adjacent to the target block. In merge mode, a differential motion vector is not coded, but a motion vector is determined on the decoding side. The accuracy of the motion vector can be up to 1 / 4 and 1 / 8 pixel accuracy for luma and chroma, respectively, which is higher than that of HEVC.
[0028] Integer transformation and quantization VVC employs linear and quadratic integer transforms. In linear transforms, the predicted difference signal is separated horizontally and vertically and then subjected to integer transforms. In linear transforms, to prevent an increase in the number of operations due to matrix operations, a mechanism is employed in which the coefficient values of high-frequency components are forcibly set to zero for large block sizes. In quadratic transforms, the bias in coefficient distribution remaining in the coefficients after the primary transform is further concentrated in the low-frequency range by performing a retransform, thereby compressing the amount of information.
[0029] In addition to a quantizer with a fixed quantization step, DQ (Dependent Quantization) is introduced. DQ is a quantization method that switches between quantizers with half-shifted quantization step positions according to a state transition table for each coefficient.
[0030] ·filter A deblocking filter reduces block boundary distortion in a reconstructed image. A pixel adaptive offset filter removes ringing distortion or random discrete noise from a reconstructed image processed by a deblocking filter. The adaptive in-loop filter uses a diamond-shaped filter with filter coefficients selected from multiple filter coefficient sets depending on the feature values of the pixel being processed. The adaptive in-loop filter can improve the signal-to-noise (SN) value. Luminance mapping performs a conversion process that biases the pixel values of the luminance component of the input image by varying the step width according to the importance of the information. Chrominance inverse scaling scales the predicted difference signal of the chrominance component so that it is inversely proportional to the bias in the step width after the conversion process, thereby suppressing an increase in the amount of code.
[0031] Entropy coding In entropy coding, CABAC (Context-based Adaptive Binary Arithmetic Coding) coding is applied to coded data for all profiles below CTU. CABAC coding is a method of variable-length coding of syntax elements using a probability model (context) that is adaptively selected according to the type of individual code (syntax element) or the surrounding circumstances. The data after entropy coding is output as coded data.
[0032] The encoded data is transmitted as a bitstream from the video encoding device to the video decoding device. Next, the structure of the bitstream used in VVC will be described.
[0033] (Bitstream) Fig. 3 is a diagram showing an example of the configuration of a bitstream according to the first embodiment. Note that the bitstream shown in Fig. 3 represents an example of the configuration of a bitstream when used in VVC, but in the first embodiment, a bitstream according to an encoding method other than VVC may be applied, and for example, a bitstream used in HEVC or AVC may be applied.
[0034] As shown in FIG. 3, a bitstream is made up of one or more CVSs (Coded Video Sequences) and an EoB (End of Bitstream NAL unit).
[0035] A CVS consists of one or more AUs (Access Units) and an EoS (End of Sequence NAL unit). The first AU of a CVS is called a CVSS AU (Coded Video Sequence Start AU). In a CVS, one or more Non-CVSS AUs are included between the CVSS AU and the EoS. Each layer of a CVS is called a CLVS (Coded Layer Video Sequence).
[0036] An AU consists of one or more Picture Units (PUs). In the case of a non-multi-layer stream, an AU consists of one PU at the same output time, and in the case of a multi-layer stream, an AU consists of multiple PUs at the same output time. In the case of a multi-layer stream, PUs are stored in the AU in order from the layer with the smallest layer number.
[0037] A CLVS is composed of PUs of the same layer. The first PU of a CLVS is called the CLVS PU. CLVS PUs are used exclusively for PUs that are IRAPs (Intra Random Access Points) or GDRs (Gradual Decoding Refreshes). IRAPs represent intra-pictures. GDRs are pictures used during ultra-low latency operation and are used as intra-slice refresh.
[0038] A PU consists of multiple NALUs (Network Abstraction Layer Units). A NALU is the basic access unit of a bitstream. There are two types of NALUs: VCL NALUs (Video Coding Layer NALUs), which are coded data, and Non-VCL NALUs, which are various header information. Figure 4 shows the types of Non-VCL NALUs included in a PU and their descriptions.
[0039] As shown in FIGS. 3 and 4, a bitstream includes a NALU (SEI_NUT) of SEI (Supplemental Enhancement Information). SEI is a specific type of information that assists in processes related to decoding, display, etc. The video decoding device 200 can use the SEI to decode or display coded images. The SEI_NUT may include the SEI as an SEI message. That is, the SEI message is included in a bitstream including coded data. Here, the SEI message according to the first embodiment will be described.
[0040] (SEI Message) An SEI message provides supplemental extension information in a syntax structure that contains one or more syntax elements that occur in a specific order relative to each other in a data bitstream.
[0041] FIG. 5 is a diagram illustrating an example of the configuration of an SEI message according to the first embodiment.
[0042] As shown in FIG. 5, the SEI message is made up of a payload type (payloadType), a payload size (payloadSize), and an SEI payload (sei_payload).
[0043] The payload type indicates the type of SEI included in the SEI payload. In the SEI message, the payload type is represented by a payload type byte (payload_type_byte) that indicates the number of bytes of the payload type.
[0044] The payload size indicates the size of the SEI payload, and in an SEI message, the payload size is represented by the number of bytes of the SEI payload (payload_size_byte).
[0045] The SEI payload contains supplemental extension information (SEI) according to the payload type.
[0046] For example, when the payload type is "0", the SEI payload includes a buffering period (buffering_period) as supplemental enhancement information. The buffering period indicates the initial deletion delay and its offset of the coded picture buffer (CPB) for initializing the hypothetical decoder (HRD) on the decoder side. An SEI message that includes the buffering period in the SEI payload is called a buffering period (BP) SEI message.
[0047] Also, for example, when the payload type is "1", the SEI payload includes picture timing (pic_timing) as supplemental extension information. The picture timing indicates the deletion delay of the CPB and the output delay of the DPB (Decoded Picture Buffer). An SEI message in which picture timing is included in the SEI payload is called a picture timing (PT) SEI message.
[0048] In this way, the SEI message can transmit supplemental enhancement information according to the payload type.
[0049] As mentioned above, JVET is currently discussing SEI messages for transmitting NNPFs (see Non-Patent Document 2). NNPFs are a type of neural network that outputs super-resolution decoded video with different resolutions from decoded video of encoded video. NNPFs are also a type of neural network that performs super-resolution through post-filtering processing. For example, if the payload type is a specific value representing NNPF, the SEI message will include syntax elements related to the NNPF. These syntax elements make it possible to include information about the NNPF in the SEI message as supplemental enhancement information and transmit it as a bitstream.
[0050] Non-Patent Document 2 proposes an NNPF characteristics SEI message (neural-network post-filter characteristics (NNPFC) SEI message) and an NNPF activation SEI message (neural-network post-filter activation (NNPFA) SEI message) as SEI messages for transmitting an NNPF. The NNPF characteristics SEI message is an SEI message that specifies the characteristics of an NNPF. On the other hand, the NNPF activation SEI message is an SEI message that instructs the use of an NNPF for a specific picture.
[0051] In the first embodiment, the description will be focused on the NNPF characteristic SEI message.
[0052] (NNPF Characteristics SEI Message Syntax Elements) 6 to 8 show examples of syntax elements included in the NNPF characteristics SEI message according to the first embodiment. The syntax elements will be described below.
[0053] nnpfc_purpose nnpfc_purpose indicates the purpose of the NNPF, which may be, for example, general video quality improvement, chrominance upsampling (such as conversion from 4:2:0 chrominance format to 4:2:2 or 4:4:4 chrominance format), resolution resampling (increasing or decreasing width and height), picture rate upsampling, bit depth upsampling, or colorization.
[0054] nnpfc_id nnpfc_id contains an identification number that identifies the NNPF. If the NNPF characteristics SEI message is the first NNPF characteristics SEI message in decoding order that has a specific nnpfc_id in the current layer sequence (CLVS), it indicates the following: The SEI message indicates the base NNPF. The SEI message concerns the current decoded picture and all decoded pictures in the current layer that follow it, up to the end of the current layer sequence.
[0055] nnpfc_base_flag If nnpfc_base_flag is "1", the SEI message specifies a base NNPF, whereas if nnpfc_base_flag is "0", the SEI message specifies an update to the base NNPF.
[0056] nnpfc_mode_idc If nnpfc_mode_idc is "0," it indicates that the neural network information is included in the NNPF characteristics SEI message and that the format of the neural network information is the ISO / IEC 15938-17 bitstream format. ISO / IEC 15938-17 specifies the compression format of neural networks. That is, if nnpfc_mode_idc is "0," it indicates that the NNPF characteristics SEI message contains a compressed neural network model. This makes it possible to transmit a compressed neural network model using the NNPF characteristics SEI message. On the other hand, if nnpfc_mode_idc is "1," it indicates that the neural network information is specified by a Uniform Resource Identifier (URI). A URI is an identifier standardized as RFC3896.
[0057] nnpfc_alignment_zero_bit_a nnpfc_alignment_zero_bit_a is equal to "0".
[0058] nnpfc_tag_uri The nnpfc_tag_uri contains a tag URI that identifies the format and associated information about the neural network to be used as a base NNPF, or an update to a base NNPF with the same nnpfc_id value specified in the nnpfc_uri. Tag URIs are specified as RFC4151.
[0059] nnpfc_uri The nnpfc_uri contains a URI that identifies the neural network to be used as a base NNPF or an update to a base NNPF with the same nnpfc_id value specified in the nnpfc_uri. The URI is specified in IETF (Internet Engineering Task Force) Internet Standard 66.
[0060] nnpfc_property_present_flag nnpfc_property_present_flag = "1" indicates that syntax elements related to filter properties, including purpose, input format, output format, and complexity, are present, whereas nnpfc_property_present_flag = "0" indicates that syntax elements related to filter properties are not present.
[0061] nnpfc_num_input_pics_minus1 The value of nnpfc_num_input_pics_minus1 plus "1" indicates the number of pictures (numInputPics) used as input to the NNPF. The value of nnpfc_num_input_pics_minus1 is in the range from "0" to "63". In other words, the maximum number of pictures usable for input to the NNPF is "64".
[0062] ·nnpfc_input_pic_filtering_flag[ i ] nnpfc_input_pic_filtering_flag[ i ] = 1 indicates that the NNPF generates an output picture for the i-th input picture. Conversely, nnpfc_input_pic_filtering_flag[ i ] = 0 indicates that the NNPF does not generate an output picture for the i-th input picture. Each picture generated by the NNPF is stored in the output tensor of the NNPF.
[0063] ·nnpfc_absent_input_pic_zero_flag If nnpfc_absent_input_pic_zero_flag is "1", it indicates that an input picture not present in the bitstream is expected to be represented by a sample array with sample values equal to "0". Conversely, if nnpfc_absent_input_pic_zero_flag is "0", it indicates that an input picture inputPicA not present in the bitstream is expected to be represented by input picture inputPicB present in the bitstream and nearest to input picture inputPicA in output order.
[0064] nnpfc_out_sub_c_flag nnpfc_out_sub_c_flag specifies the variables outSubWidthC and outSubHeightC when the chroma upsampling flag (ChromaUpsamplingFlag) is "1". When nnpfc_out_sub_c_flag is "1", the variables outSubWidthC and outSubHeightC are "1". On the other hand, when nnpfc_out_sub_c_flag is "0", the variables outSubWidthC and outSubHeightC are "2" and "1". When the chroma format IDC (ChromaFormatIdc) is "2" (i.e., when the chroma format is 4:2:2), nnpfc_out_sub_c_flag must exist and its value must be "1".
[0065] nnpfc_out_colour_format_idc When the colorization flag (ColourizationFlag) is "1", nnpfc_out_colour_format_idc specifies the color format of the NNPF-generated picture, and as a result, specifies the variables outSubWidthC and outSubHeightC. When nnpfc_out_colour_format_idc is "1", it specifies that the color format of the NNPF-generated picture is 4:2:0, and that the variables outSubWidthC and outSubHeightC are both "2". When nnpfc_out_colour_format_idc is "2", it specifies that the color format of the NNPF-generated picture is 4:2:2, and that the variables outSubWidthC and outSubHeightC are "2" and "1". When nnpfc_out_colour_format_idc is "3", it specifies that the color format of the NNPF-generated picture is 4:4:4, and that the variables outSubWidthC and outSubHeightC are both "1".
[0066] ·nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 The values of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 plus "1" represent the numerator and denominator, respectively, of the resampling ratio of the NNPF generated picture width to the cropped width (CroppedWidth).
[0067] nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 The values of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 plus "1" represent the numerator and denominator, respectively, of the resampling ratio of the height of the NNPF-generated picture to the cropped height (CroppedHeight).
[0068] ·nnpfc_interpolated_pics[ i ] nnpfc_interpolated_pics[ i ] indicates the number of interpolated pictures generated by the NNPF between the i-th input picture and the (i+1)-th input picture of the NNPF.
[0069] nnpfc_component_last_flag If nnpfc_component_last_flag is "1", it indicates that the last dimension of the NNPF input tensor inputTensor and output tensor outputTensor is used in the current channel. If nnpfc_component_last_flag is "0", it indicates that the third dimension of the NNPF input tensor inputTensor and output tensor outputTensor is used in the current channel.
[0070] nnpfc_inp_format_idc nnpfc_inp_format_idc indicates how to convert the sample values of the input picture into input values for the NNPF. When nnpfc_inp_format_idc is "0", the input values to the NNPF are real numbers, and are converted to input values using a specific formula. On the other hand, when nnpfc_inp_format_idc is "1", the input values to the NNPF are unsigned integers, and are converted to input values using a formula different from when nnpfc_inp_format_idc is "0". nnpfc_inp_format_idc is specified in the range from "0" to "255", with values "2" and above being reserved for future use.
[0071] nnpfc_auxiliary_inp_idc If nnpfc_auxiliary_inp_idc is "0", it indicates that auxiliary input data is not present in the input tensor of the NNPF. On the other hand, if nnpfc_auxiliary_inp_idc is greater than "0", it indicates that auxiliary input data is present in the input tensor of the NNPF.
[0072] nnpfc_inp_order_idc nnpfc_inp_order_idc indicates how to order the sample array of the input picture to form the input tensor to the NNPF. An example of nnpfc_inp_order_idc is shown in Table 1 below.
[0073] [Table 1] ·nnpfc_inp_tensor_luma_bitdepth_minus8 The value of nnpfc_inp_tensor_luma_bitdepth_minus8 plus "8" indicates the bit depth of the luma sample values of the input integer tensor.
[0074] ·nnpfc_inp_tensor_chroma_bitdepth_minus8 The value of nnpfc_inp_tensor_chroma_bitdepth_minus8 plus "8" indicates the bit depth of the chroma sample values in the input integer tensor.
[0075] nnpfc_out_format_idc If nnpfc_out_format_idc is "0", the sample values output by the NNPF are real numbers, with values between "0" and "1" linearly mapped to the unsigned integer range of "0" to "(1 << bit depth bitDepth)-1". On the other hand, if nnpfc_out_format_idc is "1", the luma sample values output by the NNPF are unsigned integer values in the range of "0" to "(1 << output tensor bit depth γ)-1", and the chroma sample values output by the NNPF are unsigned integer values in the range of "0" to "(1 << output tensor bit depth c)-1".
[0076] nnpfc_out_order_idc nnpfc_out_order_idc indicates the sample output order obtained from NNPF. Specifically, it is as shown in Table 2 below.
[0077] [Table 2] ·nnpfc_out_tensor_luma_bitdepth_minus8 The value of nnpfc_out_tensor_luma_bitdepth_minus8 plus "8" determines the bit depth of the luma sample values in the output integer tensor (output tensor bit depth γ(outTensorBitDepth Y )) is shown.
[0078] ·nnpfc_out_tensor_chroma_bitdepth_minus8 The value of nnpfc_out_tensor_chroma_bitdepth_minus8 plus "8" is the bit depth of the chroma sample values in the output integer tensor (output tensor bit depth c(outTensorBitDepth C )) is shown.
[0079] ·nnpfc_separate_colour_description_present_flag If nnpfc_separate_colour_description_present_flag is "1", it indicates that the explicit combination of color primaries, transfer characteristics, matrix coefficients, and scaling and offset values applied in relation to the matrix coefficients for the picture resulting from the NNPF is specified in the syntax configuration of the SEI message, whereas if nnpfc_separate_colour_description_present_flag is "0", it indicates that the explicit combination of color primaries, transfer characteristics, matrix coefficients, and scaling and offset values applied in relation to the matrix coefficients for the picture resulting from the NNPF is implied by the VUI (Video usability Information) parameters.
[0080] nnpfc_colour_primaries nnpfc_colour_primaries specifies the primary colors of the picture resulting from applying the NNPF specified in the SEI message, rather than the primary colors used in the layer sequence (CLVS). If nnpfc_colour_primaries is not present in the NNPF characteristics SEI message, nnpfc_colour_primaries is inferred to be identical to the VUI primary colors syntax element (vui_colour_primaries).
[0081] ·nnpfc_transfer_characteristics nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the NNPF specified in the SEI message, not the transfer characteristics used in the layer sequence (CLVS). If nnpfc_transfer_characteristics is not present in the NNPF characteristics SEI message, the value of nnpfc_transfer_characteristics is inferred to be identical to the VUI transfer characteristics.
[0082] nnpfc_matrix_coeffs nnpfc_matrix_coeffs describes the formula used to derive the luminance and chrominance signals from the green, blue, red, or YZX primary colors. This formula is applied to the picture result after applying the NNPF specified in the SEI message. The formula itself is specified in Matrix Coefficients of ITU-T H.273 (ISO / IEC 23091-2).
[0083] nnpfc_full_range_flag nnpfc_full_range_flag indicates the scaling and offset values to be applied relative to the matrix coefficients specified by nnpfc_matrix_coeffs.
[0084] ·nnpfc_chroma_loc_info_present_flag When nnpfc_chroma_loc_info_present_flag is "1", it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is present in the NNPF characteristics SEI message. On the other hand, when nnpfc_chroma_loc_info_present_flag is "0", it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is not present in the NNPF characteristics SEI message.
[0085] ·nnpfc_chroma_sample_loc_type_frame If nnpfc_chroma_sample_loc_type_frame is not "6" and nnpfc_out_colour_format_idc is "1", specifies the location of the chroma samples in the output picture.
[0086] nnpfc_overlap nnpfc_overlap indicates the horizontal and vertical sample count overlap of adjacent input tensors of the NNPF. nnpfc_overlap can range from '0' to '16383'.
[0087] ·nnpfc_constant_patch_size_flag If nnpfc_constant_patch_size_flag is "1", the NNPF accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. On the other hand, if nnpfc_constant_patch_size_flag is "0", the NNPF accepts as input an extended patch of a predetermined size. The predetermined size is determined based on nnpfc_overlap.
[0088] nnpfc_patch_width_minus1 The value of nnpfc_patch_width_minus1 plus "1" indicates the horizontal sample count of the patch size required for input to the NNPF when nnpfc_constant_patch_size_flag is "1".
[0089] nnpfc_patch_height_minus1 The value of nnpfc_patch_height_minus1 plus "1" indicates the vertical sample count of the patch size required for input to the NNPF when nnpfc_constant_patch_size_flag is "1".
[0090] ·nnpfc_extended_patch_width_cd_delta_minus1 nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap indicates the common divisor for all allowable values of extended patch width required for input to NNPF when nnpfc_constant_patch_size_flag is "0".
[0091] ·nnpfc_extended_patch_height_cd_delta_minus1 nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap indicates the common divisor for all allowable values for the height of the extended patch required for input to the NNPF when nnpfc_constant_patch_size_flag is "0".
[0092] nnpfc_padding_type nnpfc_padding_type indicates the padding method when referencing a sample position outside the input picture boundary, as shown in Table 3 below.
[0093] [Table 3] nnpfc_luma_padding_val nnpfc_luma_padding_val indicates the luminance value used for padding when nnpfc_padding_type is "4".
[0094] nnpfc_cb_padding_val nnpfc_cb_padding_val indicates the Cb value used for padding when nnpfc_padding_type is "4".
[0095] nnpfc_cr_padding_val nnpfc_cr_padding_val indicates the Cr value used for padding when nnpfc_padding_type is "4".
[0096] ·nnpfc_complexity_info_present_flag If nnpfc_complexity_info_present_flag is "1", it indicates that there are one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id, whereas if nnpfc_complexity_info_present_flag is "0", it indicates that there are not one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id.
[0097] nnpfc_parameter_type_idc If nnpfc_parameter_type_idc is "0", it indicates that the neural network uses only integer parameters, if nnpfc_parameter_type_idc is "1", it indicates that the neural network uses either floating-point or integer parameters, and if nnpfc_parameter_type_idc is "2", it indicates that the neural network uses only binary parameters.
[0098] · nnpfc_log2_parameter_bit_length_minus3 When nnpfc_log2_parameter_bit_length_minus3 is the same as any of "0", "1", "2", and "3", it indicates that the neural network does not use parameters with bit lengths longer than "8", "16", "32", and "64" respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, it indicates that the neural network does not use parameters with bit lengths longer than "1".
[0099] · nnpfc_num_parameters_idc nnpfc_num_parameters_idc indicates that the maximum number of neural network parameters in NNPF is shown in units of 2048. When nnpfc_num_parameters_idc is "0", it indicates that the maximum number of neural network parameters is unknown. The value range of nnpfc_num_parameters_idc is from "0" to "52", but values greater than "52" are reserved for future use. The maximum parameter variable maxNumParameters is represented by "(2048 << nnpfc_num_parameters_idc) - 1". It is a requirement for bitstream compliance that the number of neural network parameters in NNPF is less than or equal to the maximum parameter variable maxNumParameters.
[0100] · nnpfc_num_kmac_operations_idc When nnpfc_num_kmac_operations_idc is greater than "0", it indicates that the maximum number of accumulation operations per NNPF sample is less than or equal to nnpfc_num_kmac_operations_idc * 1000. When nnpfc_num_kmac_operations_idc is "0", it indicates that the maximum number of accumulation operations of the neural network is unknown.
[0101] nnpfc_total_kilobyte_size If nnpfc_total_kilobyte_size is greater than "0", it indicates the total size (in kilobytes) required to store the uncompressed parameters of the neural network. nnpfc_total_kilobyte_size indicates the total size in bits divided by "8000" and rounded up. If nnpfc_total_kilobyte_size is "0", it indicates that the total size required to store the parameters of the neural network is unknown.
[0102] ·nnpfc_num_metadata_extension_bits If nnpfc_num_metadata_extension_bits is "0", it indicates that nnpfc_reserved_metadata_extension does not exist. On the other hand, if nnpfc_num_metadata_extension_bits is greater than "0", it indicates the bit length of nnpfc_reserved_metadata_extension.
[0103] ·nnpfc_reserved_metadata_extension nnpfc_reserved_metadata_extension indicates that no bitstreams conforming to this version of this document shall be present. If nnpfc_reserved_metadata_extension is present, its length in bits shall be the same as nnpfc_num_metadata_extension_bits.
[0104] nnpfc_alignment_zero_bit_b nnpfc_alignment_zero_bit_b must be equal to "0".
[0105] nnpfc_payload_byte[ i ] nnpfc_payload_byte[ i ] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[ i ] for all values of i must be a complete bitstream conforming to ISO / IEC 15938-17. ISO / IEC 15938-17 is a specification governing neural network compression methods.
[0106] The last syntax in Figure 8 indicates that information about the compressed neural network is included in the SEI payload of the NNPF Characteristics SEI message, since nnpfc_mode_idc is "0".
[0107] (Compression of neural networks) Next, we will explain neural network compression as defined in ISO / IEC 15938-17. Neural network compression is performed on trained neural network parameters and weights. Neural network compression may be applied to the entire neural network, or to differential updates of the neural network relative to the base network. Differential updates are used, for example, when fine-tuning the model of the base neural network. Neural network compression may also be called neural network coding (NNC). A compressed neural network may also be called a compressed neural network representation (NNR). NNR refers to a neural network representation with model parameters encoded using a compression tool. Note that, hereinafter, the compressed neural network (or compressed neural network model) may be referred to as the "encoded neural network."
[0108] FIG. 9 is a diagram illustrating an example of the configuration of an NNR encoder 400 according to the first embodiment. The NNR encoder 400 receives an original neural network as input and outputs a bitstream including an encoded neural network. The NNR encoder 400 may be included in the video encoding device 100. Note that the NNR encoder 400 illustrated in FIG. 9 represents an example of a compression method for compressing a neural network. In the first embodiment, the neural network may be compressed by an encoder using a compression method other than the NNR encoder 400 illustrated in FIG. 9.
[0109] As shown in FIG. 9, the NNR encoder 400 includes a parameter reduction unit 401, a parameter quantization unit 402, and an entropy coding unit 403.
[0110] The parameter reduction unit 401 reduces the parameters of the neural network. First, the parameter reduction unit 401 may reduce the parameters by sparsification. Sparsification may be performed by replacing some of the weight values used as parameters with "0". Second, the parameter reduction unit 401 may reduce the parameters by unification. Unification is making parameters similar to each other. Unification reduces the entropy of the model parameters. Third, the parameter reduction unit 401 may reduce the parameters by pruning. Pruning is deleting a parameter or a group of parameters. Fourth, the parameter reduction unit 401 may reduce the parameters by decomposition. Decomposition is changing the weight structure of the model by a matrix decomposition operation. The parameter reduction unit 401 may perform a combination of these methods, or apply them in order.
[0111] The parameter reduction unit 401 may output the reduced parameters as a bitstream, or may output the reduced parameters to the parameter quantization unit 402.
[0112] The parameter quantization unit 402 may receive the original neural network as input, or may receive the output from the parameter reduction unit 401. The parameter quantization unit 402 quantizes the parameters to reduce the precision of the parameter representation compared to before quantization. The quantization itself may be performed using a fixed step size (uniform quantization) or using a codebook (codebook-based quantization). The parameters quantized by the parameter quantization unit 402 are output to the entropy coding unit 403.
[0113] The entropy coding unit 403 may receive the original neural network as input, or may receive the output from the parameter quantization unit 402. The entropy coding unit 403 encodes the parameters. DeepCABAC is used for encoding. DeepCABAC is a method for compressing neural network parameters, and is a compression method that minimizes the rate-distortion function while taking into account the impact of quantization on the accuracy of the neural network. The parameters compressed by the entropy coding unit 403 are output from the NNR encoder 400 as a bitstream.
[0114] As described above, the neural network model compressed by the NNR encoder 400 is included in the NNPF characteristic SEI message as an encoded neural network and transmitted from the video encoding device 100 to the video decoding device 200 as a bitstream together with the encoded video data.
[0115] (Configuration examples of a video encoding device and a video decoding device) Next, an example of the configuration of the video encoding device 100 and the video decoding device 200 according to the first embodiment will be described.
[0116] FIG. 10 is a diagram illustrating an example of the configuration of a video encoding device 100 and a video decoding device 200 according to the first embodiment.
[0117] The video encoding device 100 receives video data of an input video and outputs video data of an encoded video and encoded neural network data. Note that, hereinafter, the video data may be simply referred to as "video" and the neural network data may be simply referred to as "neural network."
[0118] On the other hand, the video decoding device 200 receives the coding neural network and coded video, and outputs super-resolution decoded video.
[0119] As shown in FIG. 10, the video encoding device 100 includes a video encoding device 110, a neural network training device 120, a neural network encoding device 130, and a bitstream generation device 140.
[0120] The video encoding device 100 performs encoding processing on input video and outputs encoded video. The encoding processing will be described using VVC as an example, but as mentioned above, encoding processing other than VVC may also be used. The video encoding device 100 may include, for example, a VVC encoder shown in FIG. 2. The video encoding device 100 outputs the encoded video to a bitstream generation device 140.
[0121] Furthermore, the video encoding device 100 generates decoded video by decoding the encoded video. For example, the video encoding device 100 can generate reconstructed video, and the reconstructed video may be output as decoded video. Alternatively, the video encoding device 100 may include a decoding processing unit that performs decoding processing on the encoded video and outputs the decoded video. The video encoding device 100 outputs the decoded video to the neural network training device 120.
[0122] The neural network training device 120 receives the input video and the decoded video, and outputs a neural network model and neural network syntax. The neural network syntax includes one or more syntax elements related to the neural network model. The neural network syntax may be included in the bitstream as a syntax element.
[0123] In the following, an NNPF will be described as an example of a neural network model, but any neural network model other than the NNPF may be used as long as it is capable of outputting super-resolution decoded video from a decoded image. The NNPF is also an example of a learning model (or a trained model) that infers super-resolution decoded video from a decoded video. Therefore, the neural network syntax may include syntax elements of the NNPF (FIGS. 6 to 8).
[0124] FIG. 11 is a diagram illustrating an example of the configuration of a neural network learning device 120 according to the first embodiment.
[0125] As shown in FIG. 11, the neural network training device 120 includes an original picture list creation unit 121, a decoded picture list creation unit, a neural network inference unit 123, a hidden state list creation unit 124, and a neural network training unit 125.
[0126] The original picture list creation unit 121 receives input video and generates an original picture list. The original picture list is a list in which the input video is arranged in a specific order. The original picture list creation unit 121 outputs the original picture list to the neural network training unit 125.
[0127] The decoded picture list creation unit 122 receives the decoded video and generates a decoded picture list. The decoded picture list is a list of decoded videos arranged in a specific order. The decoded picture list creation unit 122 outputs the decoded picture list to the neural network inference unit 123.
[0128] The neural network inference unit 123 receives the decoded video included in the decoded picture list and the hidden states included in the hidden state list, and infers (or generates) a neural network model based on the decoded video and the hidden states.
[0129] Fig. 12 is a diagram showing an example of the configuration of a neural network model according to the first embodiment. The neural network inference unit 123 generates, for example, the neural network model shown in Fig. 12. The structure of the neural network model may be called a topology.
[0130] As shown in Fig. 12, in the neural network model shown in the first embodiment, two decoded images, decoded image A and decoded image B, are input. If decoded image B is the "current image," then decoded image A is the "past image." The number of channels in decoded image A and decoded image B is "1." For example, in decoded image A and decoded image B, only the luminance matrix is used.
[0131] The hidden state may be a feature of an image. Alternatively, the hidden state may represent a learning result in a neural network model. It has been verified in Non-Patent Document 1 that the image quality of a super-resolution image can be improved by using hidden states in a neural network model that infers a super-resolution image from an input image. In the neural network model shown in FIG. 12, a hidden state A related to decoded image A and a hidden state B related to decoded image B are input. In the example shown in FIG. 12, the number of channels in hidden state A and hidden state B is "7". For example, four luminance matrices, two chrominance matrices, and one auxiliary input matrix are used for hidden state A and hidden state B. In the neural network model shown in FIG. 12, the hidden state corresponding to the decoded image is also input.
[0132] In the above-mentioned NNPF syntax element (nnpfc_num_input_pics_minus1), the maximum number of pictures that can be input to the NNPF is currently "64" (nnpfc_num_input_pics_minus1 ranges from "0" to "63"). In the neural network model shown in Fig. 12, the number of input pictures represents an example of "2" pictures, that is, decoded image A and decoded image B.
[0133] In the first embodiment, the topology of the neural network may be any topology that inputs a decoded image and a hidden state and infers (or outputs) a super-resolution decoded image and a hidden state. In the example shown in Fig. 12, the topology has a concatenation unit (concat), six residual blocks (Resblock), and two convolution layers (Conv).
[0134] The concatenation section (concat) combines a total of 16 channels: decoded image A (1 channel), hidden state A (7 channels), decoded image B (1 channel), and hidden state B (7 channels).
[0135] Each residual block (Resblock) extracts features from 16-channel input. FIG. 13 is a diagram showing an example of the configuration of a residual block (Resblock) according to the first embodiment. In the residual block, multiple convolutional layers (Conv) are connected in series, and an output is obtained by connecting the input and the final output of the convolutional layer via a skip connection. While FIG. 13 shows an example including two convolutional layers (Conv), the residual block may also be configured with three or more convolutional layers.
[0136] 12, the convolution layer (Conv) for the intermediate output of the residual block (Resblock) outputs hidden state B. Hidden state B is output to hidden state list creation unit 124.
[0137] The final convolutional layer (Conv) converts 16-channel input into a "1"-channel output. In the topology shown in Figure 12, the final convolutional layer (Conv) adds decoded image A to the super-resolution decoded image B, which is "1" channel. The neural network model outputs super-resolution decoded image B as the super-resolution decoded video.
[0138] The neural network inference unit 123 infers (or generates), for example, a neural network model (NNPF) shown in FIG.
[0139] 11 , the neural network inference unit 123 outputs the neural network model, the super-resolution decoded video, and the hidden state to the neural network learning unit 125. In addition, the neural network inference unit 123 outputs the hidden state to the hidden state list creation unit 124.
[0140] The hidden state list creation unit 124 creates a hidden state list for the hidden states output from the neural network inference unit 123, and outputs the created hidden state list to the neural network inference unit 123. The hidden state list is a list in which the hidden states are arranged in a certain order or the like.
[0141] The neural network training unit 125 receives input video (original image) included in the original picture list, super-resolution decoded video, and a neural network model. Then, the neural network training unit 125 uses the input video and the super-resolution decoded video to train the neural network model so that the difference between the input video and the super-resolution decoded video is minimized. The neural network training unit 125 performs training using a loss function, but any loss function may be used. For example, in the first embodiment, MSE (Mean Squared Error) is used as the loss function. The neural network training unit 125 also uses an optimization algorithm (called an "optimizer") for determining parameters that minimize the loss function, but any optimizer may be used. For example, in the first embodiment, Adam (Adaptive moment estimation) is used. Adam is an optimization algorithm that estimates the bias between the first moment (mean) and second moment (variance) of the gradient by correction.
[0142] Furthermore, depending on the performance of the neural network inference unit 123 and the neural network learning unit 125, the neural network model may be trained by cutting out a portion of the input video and the decoded video. The cutout size (patch size) may be set arbitrarily, but in the first embodiment, the image size is set to 256 x 256 pixels and the patch size is set to "8" for training. Furthermore, to converge the training, it is desirable that the number of iterations (repetitions or updates) is sufficiently large, but any value may be set.
[0143] The neural network training unit 125 generates a neural network syntax using the trained neural network model and the hidden state. The neural network training unit 125 may generate syntax elements of the neural network using parameters of the trained neural network model and combine the syntax elements to form the neural network syntax. Alternatively, the neural network training unit 125 may generate syntax elements of the neural network using the hidden state and combine the syntax elements to form the neural network syntax. The syntax elements may be the syntax elements of the NNPF shown in FIGS. 6 to 8. The neural network training unit 125 outputs the trained neural network model as the neural network model and also outputs the neural network syntax.
[0144] When the neural network model is the neural network model shown in FIG. 12, the neural network learning unit 125 generates the following syntax elements.
[0145] First, with regard to the syntax element nnpfc_num_input_pics_minus1, the value obtained by adding "1" to nnpfc_num_input_pics_minus1 represents the number of pictures input to the neural network model (NNPF). In the case of the neural network model shown in Fig. 12, the number of pictures of the input decoded image is "2". Therefore, the neural network training unit 125 generates a syntax element in which "1" is set to the syntax element nnpfc_num_input_pics_minus1.
[0146] Second, the syntax element nnpfc_input_pic_filtering_flag[ i ] indicates whether or not an output picture is to be output in the NNPF for the i-th input picture. In the neural network model shown in Fig. 12, super-resolution decoded image B is output for decoded image B, but super-resolution decoded image A is not output for decoded image A. Therefore, the neural network training unit 125 generates the syntax element nnpfc_input_pic_filtering_flag[ i ] by setting "0" for nnpfc_input_pic_filtering_flag[ i ] for decoded image A and "1" for nnpfc_input_pic_filtering_flag[ i ] for decoded image B.
[0147] In the first embodiment, a hidden state is used in the neural network model, and therefore, in the first embodiment, a syntax element related to the hidden state is newly defined.
[0148] Fig. 14 is a diagram illustrating an example of a syntax structure according to the first embodiment. As shown in Fig. 14, syntax elements related to the hidden state include "nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag."
[0149] "nnpfc_hidden_state_inp_idc" indicates at least whether or not a hidden state is used as an input to the neural network model (NNPF). That is, when "nnpfc_hidden_state_inp_idc" is "0," it indicates that the hidden state is not used as an input to the neural network model. On the other hand, when "nnpfc_hidden_state_inp_idc" is "1" or greater (or greater than "0"), it indicates that the hidden state is used as an input to the neural network. Also, when "nnpfc_hidden_state_inp_idc" is "1" or greater, "nnpfc_hidden_state_inp_idc" indicates the number of channels of the hidden state. In the neural network model shown in Figure 12, a hidden state is used, and the number of channels is "7," so "nnpfc_hidden_state_inp_idc" is "7." When "nnpfc_hidden_state_inp_idc" is "1" or greater, a larger value may indicate that the video decoding device 200 requires more processing power for the neural network model.
[0150] On the other hand, "nnpfc_hidden_state_absent_input_pic_zero_flag" is a syntax element used to determine the value of a hidden state that has not been calculated. "nnpfc_hidden_state_absent_input_pic_zero_flag" is used when "nnpfc_hidden_state_inp_idc" is equal to or greater than "1" (or greater than "0").
[0151] Specifically, when "nnpfc_hidden_state_absent_input_pic_zero_flag" is "1," it indicates that all values of hidden states that have not been calculated are equal to "0." For example, in the neural network model shown in FIG. 12, if the value of hidden state B of decoded image B has not been calculated, all values of hidden state B are set to "0" if "nnpfc_hidden_state_absent_input_pic_zero_flag" is "1." The same applies to hidden state A; if the value of hidden state A has not been calculated, all values of hidden state A are set to "0" if "nnpfc_hidden_state_absent_input_pic_zero_flag" is "1."
[0152] On the other hand, when "nnpfc_hidden_state_absent_input_pic_zero_flag" is "0," it indicates that the value of the uncalculated hidden state is equal to the value of the hidden state that is closest in output order. For example, in the neural network model shown in FIG. 12, if the value of hidden state B has not been calculated, and "nnpfc_hidden_state_absent_input_pic_zero_flag" is "0," the value is set to a value equal to the value of hidden state A that is closest in output order. Similarly, for hidden state A, if the value has not been calculated, and "nnpfc_hidden_state_absent_input_pic_zero_flag" is "0," the value is set to a value equal to the value of the hidden state that is older than hidden state A.
[0153] In this way, the neural network learning unit 125 may generate "nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag" by setting values to "nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag" based on the hidden state. Note that the neural network learning unit 125 also generates other syntax elements (e.g., FIGS. 6 to 8) based on the neural network model.
[0154] The neural network training unit 125 outputs the neural network model and the neural network syntax to the neural network coding device 130 .
[0155] Returning to Fig. 10, the neural network coding device 130 receives a neural network model and a neural network syntax and outputs an encoded neural network. The neural network coding device 130 compresses the neural network model and the neural network syntax. A known compression method may be used, for example, the NNR encoder shown in Fig. 9. The neural network coding device 130 combines the compressed neural network model and the compressed neural network syntax and outputs the combined result as an encoded neural network.
[0156] The bitstream generating device 140 receives the coding neural network model and the coded video, and generates a bitstream including the coding neural network model and the coded video. The bitstream may have the configuration shown in FIG. 3, for example. The bitstream generating device 140 transmits the bitstream to the video decoding device 200. The bitstream generating device 140 is also an example of a transmitting unit.
[0157] As shown in FIG. 10, the video decoding device 200 includes a neural network decoding device 210, a video decoding device 220, and a neural network inference device 230.
[0158] The neural network decoding device 210 extracts the encoded neural network from the bitstream and decodes the neural network model and neural network syntax from the encoded neural network. The decoding method is the same as the encoding method used in the neural network encoding device 130. For example, in the first embodiment, the decoding method is the same as the encoding method used in the NNR encoder (FIG. 9). The neural network decoding device 210 outputs the neural network model and neural network syntax to the neural network inference device 230.
[0159] The video decoding device 220 performs a decoding process on the coded video included in the bitstream to generate decoded video. The decoding process uses the same coding method as that used by the video coding device 110. In the first embodiment, for example, VVC is used as the decoding method. The video decoding device 220 outputs the decoded video to the neural network inference device 230.
[0160] The neural network inference device 230 receives the neural network model and the decoded image as input, and outputs the super-resolution decoded image.
[0161] FIG. 15 is a diagram illustrating an example of the configuration of a neural network inference device 230 according to the first embodiment.
[0162] As shown in FIG. 15, the neural network inference device 230 includes a decoded picture list creation unit 231, a neural network inference unit 232, and a hidden state list creation unit 233.
[0163] The decoded picture list creation unit 231 receives the decoded video output from the video decoding device 220 and creates a decoded picture list of the decoded video. The decoded picture list is a list of decoded videos arranged in a specific order. The decoded picture list creation unit 231 outputs the decoded picture list to the neural network inference unit 232.
[0164] The neural network inference unit 232 receives the neural network model and neural network syntax output from the neural network decoding device 210, receives the decoded picture list output from the decoded picture list creation unit 231, and performs inference using the neural network model, the neural network syntax, and the decoded video included in the decoded picture list. Specifically, the neural network inference unit 232 inputs the decoded video to the neural network model and causes the neural network to perform inference using the neural network syntax. The neural network syntax includes syntax elements related to hidden states ("nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag"). Therefore, the neural network inference unit 232 can infer super-resolution decoded video using the hidden state. The neural network inference device 230 outputs super-resolution decoded video as an inference result. That is, the video decoding device 200 outputs super-resolution decoded video. The neural network inference unit 232 outputs the hidden states used in the inference to the hidden state list creation unit 233.
[0165] The hidden state list creation unit 233 receives the hidden states and generates a hidden state list. The hidden state list is a list in which the hidden states are arranged in a certain order. The hidden state list creation unit 233 outputs the hidden state list to the neural network inference unit 232. The hidden states included in the hidden state list may be used recursively in the neural network inference unit 232.
[0166] (Operation example according to the first embodiment) FIG. 16 is a diagram illustrating an example of operation according to the first embodiment.
[0167] As shown in FIG. 16, in step S10, the video transmission system 10 starts processing.
[0168] In step S11, the video encoding device 110 receives an input video, performs encoding processing on the input video, and outputs encoded video.
[0169] In step S12, neural network training device 120 receives input video and decoded video output from video encoding device 110, and generates a neural network model based on the input video and the decoded video. Neural network training device 120 also generates syntax elements related to hidden states ("nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag") and outputs them, together with other syntax elements, as neural network syntax.
[0170] In step S13, the neural network coding device 130 receives the neural network model and neural network syntax output from the neural network training device 120, compresses the neural network model and neural network syntax, and outputs the compressed neural network model and neural network syntax as an encoded neural network.
[0171] In step S14, the bitstream generation device 140 receives the coding neural network output from the neural network coding device 130 and the coded video output from the video coding device 110, and generates a bitstream including the coding neural network and the coded video. The bitstream generation device 140 transmits the bitstream to the video decoding device 200. The video decoding device 200 receives the bitstream.
[0172] In step S15, the neural network decoding device 210 performs a decoding process on the coded neural network included in the bitstream to decode the neural network model, and the video decoding device 220 performs a decoding process on the coded video included in the bitstream to decode the coded video.
[0173] In step S16, the neural network inference device 230 receives the neural network model decoded by the neural network decoding device 210 and the decoded video decoded by the video decoding device 220, and infers (outputs) super-resolution decoded video from the neural network model. When inferring the super-resolution decoded video from the neural network model, the neural network inference device 230 performs inference using the hidden state by utilizing syntax elements related to the hidden state included in the neural network syntax.
[0174] Then, in step S17, the video transmission system 10 ends the series of processes.
[0175] As described above, in the first embodiment, a video encoding device (e.g., video encoding device 100) has a transmitting unit (e.g., bitstream generating device 140) that transmits a supplemental enhancement information message (e.g., NNPF characteristic SEI message) including syntax elements (e.g., "nnpfc_hidden_state_inp_idc" and "nnpfc_hidden_state_absent_input_pic_zero_flag") related to hidden states that can be used in a learning model to a video decoding device (e.g., video decoding device 200). The learning model is a learning model (e.g., NNPF, which is a neural network model) that infers, from decoded video obtained by decoding encoded video, super-resolution decoded video having a different resolution from the decoded video.
[0176] As described above, in the first embodiment, the video encoding device 100 transmits an SEI message including syntax elements related to hidden states to the video decoding device 200. This enables the video decoding device 200 to use hidden states in a neural network model that infers super-resolution decoded video from decoded video. The video decoding device 200 can output super-resolution decoded video with improved image quality compared to when the hidden states are not used.
[0177] [Other embodiments] The topology of the neural network model shown in FIG. 12 is an example. Any topology (structure) may be used as long as the model can input a decoded image and a hidden state and infer (or output) a hidden state and a super-resolution decoded image. In the example shown in FIG. 12, the number of pictures of the decoded image input to the neural network model is "2", but the number of pictures of the decoded image may be any number. Furthermore, in the example shown in FIG. 12, the number of channels of the hidden state and the decoded image is "7" and "1", respectively, but the number of channels of the hidden state and the decoded image may be any number.
[0178] Furthermore, in the first embodiment, an example has been described in which a bitstream is generated by the bitstream generating device 140, but the video encoding device 100 may not include the bitstream generating device 140. In that case, the neural network encoding device 130 may have the functions of the bitstream generating device 140. In this case, the bitstream generating device 140 may generate a bitstream and transmit it to the video decoding device 200. The neural network learning device 120 may be an example of a transmitting unit.
[0179] Also, a program may be provided that causes a computer to execute each process performed by the above-mentioned devices (video encoding device 100 and video decoding device 200). The program may be recorded on a computer-readable medium. Using the computer-readable medium, the program can be installed on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a recording medium such as a CD-ROM or a DVD-ROM. Also, circuits that execute each process performed by the above-mentioned devices (video encoding device 100 and video decoding device 200) may be integrated, and the device may be configured as a semiconductor integrated circuit (chip set, SoC).
[0180] As used in this disclosure, the terms "based on" and "depending on / in response to" do not mean "based only on" or "depending only on," unless expressly stated otherwise. The term "based on" means both "based only on" and "based at least in part on." Similarly, the term "depending on" means both "depending only on" and "depending at least in part on." The terms "include," "comprise," and variations thereof do not mean including only the listed items, but may mean including only the listed items or may include additional items in addition to the listed items. Additionally, the term "or," as used in this disclosure, is not intended to mean an exclusive or. Furthermore, any reference to elements using designations such as "first," "second," etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used herein as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed therein or that the first element must precede the second element in some way. In this disclosure, where articles are added by translation, such as a, an, and the in English, these articles shall include the plural unless the context clearly indicates otherwise.
[0181] Although the embodiments have been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes can be made without departing from the spirit of the invention. Furthermore, the operation examples can be combined within a consistent range.
[0182] (Addendum) The above is summarized as follows, but the supplementary notes do not limit the above-described embodiment.
[0183] (Appendix 1) a transmission unit that transmits, to a video decoding device, a supplemental enhancement information message including syntax elements related to hidden states that can be used in a learning model that infers, from a decoded image obtained by decoding an encoded image, a super-resolution decoded image having a different resolution from the decoded image; A video encoding device having the above configuration.
[0184] (Appendix 2) The syntax elements include at least a syntax element indicating whether the hidden state is to be used as an input to the learning model. 2. The video encoding device according to claim 1.
[0185] (Appendix 3) If the syntax element indicates that the hidden state is to be used as an input to the learning model, the syntax element indicates the number of channels of the hidden state. 10. The video encoding device according to claim 1 or 2.
[0186] (Appendix 4) The syntax elements include syntax elements used to determine values of the uncalculated hidden state. 4. A video encoding device according to claim 1.
[0187] (Appendix 5) The hidden state represents the feature of the decoded image. 5. A video encoding device according to any one of claims 1 to 4.
[0188] (Appendix 6) a learning model that infers, from a decoded image obtained by decoding an encoded image, a super-resolution decoded image having a different resolution from the decoded image, the learning model including a supplemental enhancement information message including a syntax element related to a hidden state that can be used in the learning model; and a receiving unit that receives, from a video encoding device, a supplemental enhancement information message including a syntax element related to a hidden state that can be used in the learning model. a neural network inference unit that uses the learning model to infer the super-resolution decoded image from the decoded image, The neural network inference unit infers the super-resolution decoded image using the hidden state. Video decoding device.
[0189] (Appendix 7) A program that causes a computer to function as the video encoding device according to claim 1 or the video decoding device according to claim 6. [Explanation of symbols]
[0190] 100: Video encoding device 110: Video encoding device 120: Neural network learning device 121: Original picture list creation unit 122: Decoded picture list creation unit 123: Neural network inference unit 124: Hidden state list creation unit 125: Neural network learning unit 130: Neural network coding device 140: Bitstream generation device 200: Video decoding device 210: Neural network decoding device 220: Video decoder 230: Neural network inference device 231: Decoded picture list creation unit 232: Neural network inference unit 233: Hidden state list creation unit 400: NNR encoder
Claims
1. a transmission unit that transmits, to a video decoding device, a supplemental enhancement information message including syntax elements related to hidden states that can be used in a learning model that infers, from a decoded image obtained by decoding an encoded image, a super-resolution decoded image having a different resolution from the decoded image; A video encoding device having the above configuration.
2. The syntax elements include at least a syntax element indicating whether the hidden state is to be used as an input to the learning model.
2. The video encoding device according to claim 1.
3. If the syntax element indicates that the hidden state is to be used as an input to the learning model, the syntax element indicates the number of channels of the hidden state.
3. The video encoding device according to claim 2.
4. The syntax elements include syntax elements used to determine values of the uncalculated hidden state.
2. The video encoding device according to claim 1.
5. The hidden state represents the feature of the decoded image.
2. The video encoding device according to claim 1.
6. a learning model that infers, from a decoded image obtained by decoding an encoded image, a super-resolution decoded image having a different resolution from the decoded image, the learning model including a supplemental enhancement information message including a syntax element related to a hidden state that can be used in the learning model; and a receiving unit that receives, from a video encoding device, a supplemental enhancement information message including a syntax element related to a hidden state that can be used in the learning model. a neural network inference unit that uses the learning model to infer the super-resolution decoded image from the decoded image, The neural network inference unit infers the super-resolution decoded image using the hidden state. Video decoding device.
7. A program that causes a computer to function as the video encoding device of claim 1 or the video decoding device of claim 6.