Video data processing device and video data processing method
Patent Information
- Application Number
- PCT/JP2024/036806
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-01
- Filing Date
- 2024-10-16
- Publication Date
- 2025-05-08
AI Technical Summary
In video encoding, neural network-based video encoding technology has the problem of decreasing encoding efficiency, especially when auxiliary information increases.
By introducing packets and quantization processing into the video data processing device, the neural network output is quantized using a common quantization width and inverse quantization and ungrouping processing is performed on the decoding side to maintain encoding efficiency.
The decrease in encoding efficiency is effectively suppressed, and the overall performance of video encoding is improved by reducing the data quantization width transmission of auxiliary information.
Smart Images

Figure JP2024036806_08052025_PF_FP_ABST
Abstract
Description
Video data processing device and video data processing method
[0001] The present disclosure relates to a video data processing device and a video data processing method.
[0002] In order to transmit or record video efficiently, a video encoding device is used to generate a coded representation (hereinafter referred to as a bitstream) of input video, and a video decoding device is used to decode the bitstream to generate decoded video.
[0003] [Video Coding Based on Predictive Coding in Coding Units] Video coding standards include H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High-Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding), which are standardized by ITU-T SG16 and ISO / IEC / SC29. Another recent video coding technology is the technology described in Non-Patent Document 1.
[0004] In these video coding methods, video data is coded and decoded while being managed in a hierarchical structure. The hierarchical structure is made up of, for example, pictures that make up the video data, slices (or tiles) obtained by dividing pictures, coding tree units (CTUs) obtained by dividing slices, and coding units (CUs) obtained by dividing coding tree units.
[0005] H.266 / VVC defines tiles, slices, and subpictures as divisions of a picture. A picture is divided into one or more tiles. A tile is a rectangular area whose constituent unit is a CTU. A slice is a rectangular area whose constituent unit is a tile. Slice scanning orders include raster-scan slice mode and rectangular slice mode. Raster-scan slice mode is a mode in which slices are arranged in raster scan order. Rectangular slice mode is a mode in which the coverage area of a slice is a rectangular area whose unit is a tile or a CTU line within a tile. A subpicture is made up of one or more slices.
[0006] The input image of the target CU is usually encoded earlier than the target CU and is predictively coded based on a predicted image generated based on the decoded image that has been decoded. That is, a prediction error image obtained by subtracting the predicted image from the input image is coded and decoded. Predictive coding includes intra-picture prediction (intra-prediction) that uses a decoded image included in a picture with the same display time as the target CU, and inter-picture prediction (inter-prediction) that uses a decoded image included in a picture with a different display time from the target CU.
[0007] The prediction error image is encoded based on frequency transform, quantization, and entropy coding. The prediction error image is decoded based on entropy decoding, inverse quantization, and inverse frequency transform. The frequency transform value of the quantized prediction error image is called a quantized value.
[0008] [Video Coding Based on Neural Networks] Non-Patent Document 2 describes a new video coding technique that combines an auto-encoder, which is a type of neural network, quantization, and entropy coding.
[0009] An autoencoder compresses input data into a low-dimensional feature vector so that it contains only important features. The autoencoder then generates reconstructed data by reconstructing the low-dimensional feature vector back to its original dimensions. Figure 1 is an explanatory diagram showing the autoencoder algorithm. In Figure 1, the circular parts are called nodes and the arrows are called edges. The process of reducing the data into a low-dimensional feature vector (the first half) is called encoding. The process of generating reconstructed data (the second half) is called decoding.
[0010] The autoencoder is trained to minimize the reconstruction error (the difference between the input data and the reconstructed data). To obtain meaningful features, the autoencoder is designed to impose constraints on the encoding structure and to add regularization terms to the network's loss function.
[0011] "Algorithm description of Enhanced Compression Model 9(ECM 9)", JVET-AD2025, JVET of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 30th Meeting, Antalya, TR, 21-28 April 2023J. Ball'e, V. Laparra, and EP Simoncelli, "End-to-end Optimized Image "Compression", published as a conference paper at ICLR 2017
[0012] 2 is a block diagram showing a general video encoder 110 and a general video decoder 210 that encode and decode each picture constituting video data based on a neural network. Hereinafter, the video encoder will be referred to as an NN video encoder, and the video decoder will be referred to as an NN video decoder.
[0013] The NN video encoder 110 includes an encoder 1001 , a quantizer 1003 A, and an entropy encoder 1005 .
[0014] The NN video decoder 210 includes a decoder 2001 , an inverse quantizer 2003 A, and an entropy decoder 2005 .
[0015] 2 simply indicates the direction of signal (data) flow, but does not exclude bidirectionality. This also applies to other block diagrams.
[0016] [Description on the Encoding Side] In the NN video encoder 110, the encoder 1001 extracts features from the image of the input picture. Specifically, the encoder 1001 obtains a feature vector from the image of the input picture.
[0017] The quantizer 1003A quantizes the feature vector supplied from the encoder 1001 to obtain a quantized value.
[0018] The entropy encoder 1005 entropy encodes the quantized values supplied from the quantizer 1003A to obtain entropy-encoded data, which is output as a bit stream (called an NN bit stream).
[0019] [Description on the Decoding Side] In the NN video decoder 210, the entropy decoder 2005 entropy decodes the entropy-encoded data obtained from the NN bitstream to obtain quantized values.
[0020] The inverse quantizer 2003A inversely quantizes the quantized values supplied from the entropy decoder 2005 to obtain a reconstructed feature vector.
[0021] The decoder 2001 obtains a reconstructed image of a decoded picture (also called an NN decoded picture) from the reconstructed feature vector.
[0022] Generally, a quantization step size (quantization width) is controlled for rate control. Information that can identify the quantization width is transmitted as side information. That is, the NN video encoder 110 includes side information that can identify the quantization width in the NN bitstream.
[0023] For example, the NN video encoder 110 transmits side information that can specify the quantization step size for each quantization value.
[0024] If the NN video encoder 110 changes the quantization width for each quantization value, the amount of data to be transmitted increases due to the increase in side information, resulting in a decrease in coding efficiency, even if video coding efficiency is improved using a neural network. For example, if the width and height of the quantization values for one picture obtained by video coding based on a neural network are W_tensor and H_tensor, respectively, W_tensor × H_tensor quantization widths are required as side information.
[0025] An object of the present invention is to provide a video data processing device and a video data processing method that can suppress a decrease in coding efficiency in video coding based on a neural network.
[0026] A video data processing device based on the present disclosure includes a neural network, a quantization means, and an entropy coding means, and includes a grouping means for grouping outputs of the neural network, and the quantization means performs quantization for each group using a common quantization width.
[0027] Another video data processing device based on the present disclosure includes a neural network, an inverse quantization means, and an entropy decoding means, and includes an input means for inputting a common quantization width for quantization values that constitute a group, the inverse quantization means performing inverse quantization for each group using the common quantization width, and a degrouping means for degrouping the outputs of the inverse quantization means.
[0028] A video data processing method according to the present disclosure is a video data processing method that performs encoding processing, quantization processing, and entropy coding based on a neural network, in which the outputs of the neural network are grouped and quantized for each group using a common quantization width in the quantization processing.
[0029] Another video data processing method based on the present disclosure is a video data processing method that performs a neural network-based decoding process, an inverse quantization process, and an entropy decoding process, in which the inverse quantization process performs inverse quantization for each group using a common quantization width, and ungroups the outputs of the inverse quantization process.
[0030] A video data processing program based on the present disclosure causes a computer to perform encoding processing, quantization processing, and entropy coding based on a neural network, to perform processing to group the outputs of the neural network, and to perform quantization using a common quantization width for each group in the quantization processing.
[0031] Another video data processing program based on the present disclosure causes a computer to execute a neural network-based decoding process, an inverse quantization process, and an entropy decoding process, and in the inverse quantization process, performs inverse quantization using a common quantization width for each group, and ungroups the outputs of the inverse quantization process.
[0032] According to the present invention, it is possible to suppress a decrease in coding efficiency in video coding based on a neural network.
[0033] FIG. 1 is an explanatory diagram showing an algorithm of an autoencoder. FIG. 2 is a block diagram showing a general video encoder and video decoder that encode and decode each picture constituting video data based on a neural network. FIG. 3 is a block diagram showing a video data processing system that encodes and decodes each picture constituting video data based on a neural network. FIG. 4 is an explanatory diagram showing an example of grouping. FIG. 5 is a flowchart showing an example of operation of a NN video data processor on the encoding side. FIG. 6 is a flowchart showing an example of operation of a NN video data processor on the decoding side. FIG. 7 is an explanatory diagram explaining a problem. FIG. 8 is a block diagram showing a video data processing system that encodes and decodes each picture constituting video data based on a neural network. FIG. 9 is an explanatory diagram showing an example of grouping and rearrangement in a NN video data processor on the encoding side, and reverse rearrangement and reverse grouping in a NN video data processor on the decoding side. FIG. 10 is a flowchart showing an example of operation of a NN video data processor on the encoding side. FIG. 11 is a flowchart showing an example of operation of a NN video data processor on the decoding side. FIG. 12 is a block diagram showing an example of the configuration of an encoder and a decoder. FIG. 13 is a block diagram showing an example of the configuration of a video data processor on the encoding side. FIG. 14 is a block diagram showing an example of the configuration of a video data processor on the decoding side. FIG. 15 is an explanatory diagram showing an example of grouping that also takes into account channels. FIG. 16 is a block diagram showing an example of the configuration of an information processing system. FIG. 17 is a block diagram showing main parts of a video data processing device. FIG. 10 is a block diagram showing the main parts of a video data processing device according to another embodiment.
[0034] Hereinafter, an embodiment will be described with reference to the drawings.
[0035] Embodiment 1. Figure 3 is a block diagram showing an encoding-side video data processor 101 and a decoding-side video data processor 201 that encode and decode each picture constituting video data based on a neural network. Hereinafter, the video data processor 101 will be referred to as the NN video data processor 101. The video data processor 201 will be referred to as the NN video data processor 201. The NN video data processor 101 is an encoding-side video data processor that performs video processing (e.g., encoding processing) based on a neural network. The NN video data processor 201 is a decoding-side video data processor that performs video processing (e.g., decoding processing) based on a neural network. A system including the NN video data processors 101 and 201 is called a video data processing system.
[0036] The NN video data processor 101 includes an encoder 1001 , a grouping unit 1002 , a quantizer 1003 , an entropy encoder 1005 , a multiplexer 1006 , and a rate control unit 1007 .
[0037] The NN video data processor 201 includes a decoder 2001 , a degrouping unit 2002 , an inverse quantizer 2003 , an entropy decoder 2005 , and a demultiplexer 2006 .
[0038] [Description of Encoding Side] In the NN video data processor 101, the encoder 1001 extracts features from the image of the input picture. Specifically, the encoder 1001 obtains a feature vector from the image of the input picture.
[0039] The grouping unit 1002 divides the feature vectors of a picture. Hereinafter, the divided regions are referred to as FVGs (Feature Vector Groups).
[0040] The quantizer 1003 obtains a quantized value by quantizing the feature vector supplied from the encoder 1001. The quantizer 1003 performs quantization for each FVG using a quantization width for each group (group quantization width) supplied from the rate control unit 1007.
[0041] That is, the quantizer 1003 quantizes all feature vectors (samples) belonging to one group (FVG) using one quantization step size. In other words, the outputs of the encoder 1001 are grouped. Note that the quantization values to which the quantization step size assigned to one group is applied are consecutive in entropy coding order.
[0042] Fig. 4 is an explanatory diagram showing an example of grouping. In the example shown in Fig. 4, the grouping unit 1002 divides the feature vectors of a picture in units of an externally supplied feature vector group size, i.e., in units of the width FVG_tensor_w and height FVG_tensor_h of the feature vector group.
[0043] The quantizer 1003 supplies the quantization width for each group to the entropy encoder 1005 .
[0044] The entropy encoder 1005 entropy encodes the quantized values supplied from the quantizer 1003 and the group quantization width supplied from the rate control unit 1007 to obtain entropy-coded data. The entropy encoder 1005 supplies the entropy-coded data to the multiplexer 1006. The entropy encoder 1005 entropy encodes the group quantization width as auxiliary information.
[0045] Note that the entropy encoder 1005 may entropy encode prediction error values (prediction error values related to the group quantization width) that have been predictively encoded in FVG units, instead of entropy encoding the group quantization width as is.
[0046] The multiplexer 1006 multiplexes the entropy-encoded data supplied from the entropy encoder 1005 with the feature vector group size, and outputs the multiplexed data as a bit stream (NN bit stream).
[0047] The rate control unit 1007 monitors the amount of output data from the entropy coder 1005 and obtains a quantization step size derived for each FVG. A value appropriate for the time when quantization is performed in rate control is selected as the quantization step size. The rate control unit 1007 supplies the derived quantization step size to the quantizer 1003 and the entropy coder 1005 as a group quantization step size.
[0048] Next, a description will be given of the operation of the NN video data processor 101. FIG.
[0049] In the NN video data processor 101, the encoder 1001 extracts features from an input picture and generates a feature vector (step S101).
[0050] The quantizer 1003 quantizes the feature vector using the above-mentioned group quantization width to generate a quantized value (step S103).
[0051] The entropy encoder 1005 entropy encodes the quantization value and the group quantization width to generate entropy-encoded data (step S104).
[0052] The multiplexer 1006 multiplexes the entropy-encoded data, the feature vector group size, etc., and outputs the multiplexed data as an NN bit stream (step S105).
[0053] [Description on the Decoding Side] In the NN video data processor 201, the demultiplexer 2006 demultiplexes the NN bitstream to obtain entropy-encoded data and a feature vector group size. The feature vector group size is supplied to the degrouping unit 2002.
[0054] The entropy decoder 2005 entropy decodes the entropy-encoded data supplied from the demultiplexer 2006 to obtain a quantization value and a group quantization width.
[0055] In addition, when the information on the group quantization width is predictively coded, the entropy decoder 2005 treats the information on the group quantization width as a prediction error and reconstructs a value by adding a predicted value as the group quantization width.
[0056] The inverse quantizer 2003 inverse quantizes the quantized values supplied from the entropy decoder 2005 for each FVG using a quantization width for each group to obtain an FVG reconstructed feature vector. That is, the inverse quantizer 2003 inverse quantizes the quantized values that make up the group using a common quantization width.
[0057] The degrouping unit 2002 obtains a reconstructed feature vector by degrouping the FVG reconstructed feature vectors supplied from the inverse quantizer 2003 (see FIG. 4). Degrouping is performed by regarding a collection of reconstructed feature vectors as a collection of reconstructed feature vectors that are not related to FVG, as shown on the left side of FIG.
[0058] The decoder 2001 obtains a reconstructed image of a decoded picture (also called an NN decoded picture) from the reconstructed feature vector.
[0059] Next, a description will be given of the operation of the NN video data processor 201. Fig. 6 is a flowchart showing an example of the operation of the NN video data processor 201.
[0060] In the NN video data processor 201, the demultiplexer 2006 demultiplexes the bit stream (step S201). The demultiplexer 2006 obtains entropy-encoded data and feature vector group sizes through demultiplexing.
[0061] The entropy decoder 2005 entropy decodes the entropy-encoded data to obtain a quantization value and a group quantization width (step S202).
[0062] The inverse quantizer 2003 inversely quantizes the quantized value using the group quantization width (step S203). The inverse quantizer 2003 obtains a reconstructed feature vector through the inverse quantization.
[0063] The degrouping unit 2002 degroups the FVG reconstructed feature vectors to obtain reconstructed feature vectors (step S205).
[0064] The decoder 2001 obtains a reconstructed image of the decoded picture from the reconstructed feature vector (step S206).
[0065] In this embodiment, in the NN video data processor 101, the quantizer 1003 generates quantized values belonging to a group using a quantization width (one quantization width) common to the group. Furthermore, in the NN video data processor 201, the inverse quantizer 2003 inversely quantizes all quantized values belonging to the group using a quantization width (one quantization width) common to the group obtained by the entropy decoder 2005. Specifically, in this embodiment, the quantization value width is changed for each FVG, not for each quantization value. This control reduces the amount of coding for auxiliary information related to the quantization width, compared to when each quantization width corresponding to each quantization value is coded. As a result, coding efficiency is improved.
[0066] Referring to the example shown in Fig. 4, when each quantization width corresponding to each quantization value is coded, (W_tensor x H_tensor) quantization widths are transmitted. Specifically, coded data of side information of (W_tensor x H_tensor) quantization widths is transmitted. However, in this embodiment, the number of transmitted quantization widths is reduced to (W_tensor x H_tensor) / (FVG_tensor_w x FVG_tensor_h).
[0067] Embodiment 2 When video coding based on predictive coding in a coding unit and video coding based on a neural network are combined, there is a problem that the processing order of entropy coding for quantized values cannot be integrated.
[0068] Fig. 7 is an explanatory diagram illustrating the above-mentioned problem. The left side of Fig. 7 illustrates an example of the processing order of entropy coding in video coding based on neural networks. The right side of Fig. 7 illustrates an example of the processing order of entropy coding in video coding based on predictive coding in coding units.
[0069] In FIG. 7 , W_tensor indicates the row size of the feature vector. H_tensor indicates the column size of the feature vector. W_img indicates the width (horizontal size) of the input picture. H_img indicates the height (vertical size) of the input picture. CTU_img indicates the width of the CTU. Note that FIG. 7 illustrates a square CTU.
[0070] In neural network-based video coding, quantization values are processed sample by sample in order from top left to bottom right. In coding unit-based video coding based on predictive coding, quantization values are processed block by block (e.g., CTU) in order from top left to bottom right, but within each block, quantization values are processed from top left to bottom right.
[0071] When considering processing within a block, the processing order of entropy coding in neural-network-based video coding is not the same as the processing order of entropy coding in predictive-coding-based video coding in a coding unit. Therefore, when the width (assuming the height) CTU_img of a CTU in predictive-coding-based video coding in a coding unit is greater than 1, even if the width (W_tensor) and height (H_tensor) of the quantized values for one picture obtained by neural-network-based video coding are the same as the width (W_img) and height (H_img) of the quantized values for one picture obtained by predictive-coding-based video coding in a coding unit, the transmission order of the quantized values for the picture will not be the same.
[0072] When using the video data processing system of the second embodiment, the processing order of entropy coding can be aligned when video coding based on predictive coding in the coding unit and video coding based on a neural network are combined.
[0073] 8 is a block diagram showing an NN video data processor 102 on the encoding side and an NN video data processor 202 on the decoding side, which encode and decode each picture constituting video data based on a neural network. The NN video data processor 102 and the NN video data processor 202 constitute a video data processing system.
[0074] The NN video data processor 102 includes an encoder 1001, a grouping / rearranging unit 1002A, a quantizer 1003, an entropy encoder 1005, a multiplexer 1006, and a rate control unit 1007.
[0075] The NN video data processor 202 comprises a decoder 2001 , an inverse sorting / desorupling unit 2002 A, an inverse quantizer 2003 , an entropy decoder 2005 , and a demultiplexer 2006 .
[0076] [Description of the Encoding Side] In the NN video data processor 102, the encoder 1001, quantizer 1003, entropy encoder 1005, and multiplexer 1006 have the same configurations and functions as those in the first embodiment.
[0077] In this embodiment, as an example, a CTU whose width and height are CTU_tensor is used as the grouping unit.
[0078] where CTU_tensor is the width of the coding unit in the encoded feature vector. That is, the ratio between CTU_tensor and CTU_img (see FIG. 9 ), which will be described later, is the same as the ratio between W_tensor and W_img (see FIGS. 7 and 9 ). For simplicity, it is assumed that the sizes of CTU and FVG are the same (CTU_tensor = FVG_tensor_w = FVG_tensor_h).
[0079] The grouping / sorting unit 1002A divides the feature vector of the picture into FVGs having the same size as the CTUs, and then sorts the FVG feature vectors from the top left to the bottom right of the picture to obtain FVG feature vectors.
[0080] 9 is an explanatory diagram showing an example of grouping and rearrangement in the NN video data processor 102, and inverse rearrangement and inverse grouping in the NN video data processor 202. The left side of Fig. 9 illustrates an example of the processing order of entropy coding in neural network-based video coding. The right side of Fig. 9 illustrates an example of the processing order of entropy coding in video coding based on predictive coding in a coding unit.
[0081] 9, the grouping / rearranging unit 1002A rearranges the quantized values in CTU units (which are the same as FVG units in this example) whose width and height are CTU_tensor. As described above, CTU_tensor is the width of the coding unit in the encoded feature vector. Therefore, the ratio between CTU_tensor and CTU_img is the same as the ratio between W_tensor and W_img.
[0082] Although only one FVG in the upper right corner is illustrated in FIG. 9, the quantization values are rearranged in FVG units (which is the same as CTU units in this example) for all FVGs.
[0083] Also, for example, if the size of the FVG (in this example, the same as the size of the CTU) is 4x4, the grouping / rearranging unit 1002A rearranges the quantization values so that the four quantization values in the nth (n: 1 to 3)th row in the FVG are lined up next to the four quantization values in the (n+1)th row.
[0084] The rearrangement method is not limited to this example, and the blocks may be rearranged based on a specified processing order. For example, if the image is processed in a spiral from the periphery to the center, the blocks may be rearranged in accordance with this processing order.
[0085] Next, a description will be given of the operation of the NN video data processor 102. Fig. 10 is a flowchart showing an example of the operation of the NN video data processor 102.
[0086] The process in step S101 is the same as the process in the first embodiment.
[0087] The grouping / rearranging unit 1002A divides (groups) the feature vectors of the picture into FVGs having the same size as the CTUs, and then rearranges the FVG feature vectors from the top left to the bottom right of the picture, as illustrated in Fig. 9 (step S102). By the processing of step S102, the processing order within each FVG becomes the same as the processing order of video coding based on predictive coding in the coding unit.
[0088] The processes in steps S103 to S105 are the same as those in the first embodiment.
[0089] [Description of Decoding Side] The decoder 2001, inverse quantizer 2003, entropy decoder 2005, and demultiplexer 2006 have the same configurations and functions as those in the first embodiment.
[0090] The inverse sorting / undeleting unit 2002A performs a process of sorting and ungrouping the FVG reconstructed feature vectors supplied from the inverse quantizer 2003. In other words, the inverse sorting / undeleting unit 2002A performs the reverse of the process performed by the grouping / sorting unit 1002A.
[0091] Next, a description will be given of the operation of the NN video data processor 202. Fig. 11 is a flowchart showing an example of the operation of the NN video data processor 202.
[0092] The processing in steps S201 to S203 is the same as that in the first embodiment.
[0093] The inverse sorting / degrouping unit 2002A sorts the FVG reconstructed feature vectors supplied from the inverse quantizer 2003 from bottom left to bottom right in sample units (step S204). That is, the inverse sorting / degrouping unit 2002A restores the order of the reconstructed feature vectors to their original order. The inverse sorting / degrouping unit 2002A also operates in the same way as the degrouping unit 2002 in the first embodiment to perform degrouping.
[0094] The processes in steps S205 and S206 are the same as those in the first embodiment.
[0095] In this embodiment, the grouping / rearranging unit 1002A and the de-rearranging / degrouping unit 2002A can align the processing order of entropy coding in neural-network-based video coding with the processing order of entropy coding in entropy coding when video coding based on predictive coding in a coding unit and video coding based on neural-network are combined. Therefore, when video coding based on predictive coding in a coding unit and video coding based on neural-network are combined, the processing order of entropy coding can be aligned.
[0096] Furthermore, the quantization values of neural network-based video coding can be configured in units of CTUs, which are the same size as FVGs, thus enabling picture partitioning to be integrated in a combination of predictive coding-based video coding and neural network-based video coding in a coding unit.
[0097] [Encoder and Decoder Configuration] Fig. 12 is a block diagram showing an example configuration of the encoder 1001 and decoder 2001 in each of the above embodiments. In Fig. 12, a "downward arrow (↓) 2" indicates subsampling by 1 / 2 (also called pooling). An "upward arrow (↑) 2" indicates upsampling by 2.
[0098] In the example shown in Fig. 12, the encoder 1001 is composed of four residual blocks and one convolution block. Each residual block is composed of two convolutional blocks and one shortcut link. Each convolutional block is composed of one convolution layer and one activation function.
[0099] The decoder 2001 consists of four residual blocks and one pixel shuffle convolution layer. The pixel shuffler is a mechanism proposed as sub-pixel convolution. The pixel shuffler rearranges input feature vectors and outputs high-resolution feature vectors.
[0100] As an activation function, a Parametric ReLU (Parametric Rectified Linear Unit) can be used, in which the output value is α times the input value when the input value is below 0 (where α is a parameter determined by learning), and the output value is the same as the input value when the input value is 0 or greater.
[0101] Furthermore, the configuration of the neural network as the encoder 1001 and the decoder 2001 shown in FIG. 12 is just an example, and the configurations of the encoder 1001 and the decoder 2001 are not limited to the configuration shown in FIG.
[0102] Embodiment 3 Fig. 13 is a block diagram showing an example of the configuration of a video data processor on the encoding side, and Fig. 14 is a block diagram showing an example of the configuration of a video data processor on the decoding side.
[0103] The video data processor 100 shown in FIG. 13 includes a switch 3000, an NN video encoder 3001, an NN video decoder 3003, a CP video encoder 3002, a CP video decoder 3004, a decoded picture buffer 3005, and a multiplexer 4000 that performs multiplexing processing of entropy-encoded data and other information.
[0104] Note that "NN" stands for neural network, and "CP" stands for coding unit-based prediction.
[0105] The CP video encoder 3002 performs video encoding processing using a video encoding method based on predictive coding for each coding unit. The CP video decoder 3004 performs decoding processing using a video encoding method based on predictive coding for each coding unit. The decoded picture buffer 3005 is a storage unit that stores decoded pictures (reconstructed pictures). As described above, video encoding methods that comply with H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc. can be used.
[0106] The NN video data processor 101 or 102 in each of the above embodiments can be used as the NN video encoder 3001. The NN video data processor 201 or 202 in each of the above embodiments can be used as the NN video decoder 3003.
[0107] The video data processor 200 shown in FIG. 14 includes a demultiplexer 5000 that demultiplexes a bitstream, an NN video decoder 3003, a CP video decoder 3004, and a decoded picture buffer 3005.
[0108] In other words, video data processor 100 shown in FIG. 13 and video data processor 200 shown in FIG. 14 are NN video data processors in which NN video encoder 3001 is the NN video data processor 101 or 102 described above, NN video decoder 201 or 202 is the NN video decoder 3003, and these are combined with CP video encoder 3002 and CP video decoder 3004 that are based on a video encoding method based on predictive encoding in the encoding unit.
[0109] [Description of the Encoding Side] In the video data processor 100 shown in FIG. 13, a switch 3000 supplies an input picture to either an NN video encoder 3001 or a CP video encoder 3002 .
[0110] The NN video encoder 3001 operates in the same manner as the NN video data processor 101 of the first embodiment or the NN video data processor 102 of the second embodiment to generate an NN bitstream.
[0111] The NN video decoder 3003 receives the NN bitstream supplied from the NN video encoder 3001 and operates in the same manner as the NN video data processor 201 of the first embodiment or the NN video data processor 202 of the second embodiment to obtain decoded pictures (NN decoded pictures). The NN video decoder 3003 stores the NN decoded pictures in the decoded picture buffer 3005.
[0112] The CP video encoder 3002 uses the input picture and the decoded picture stored in the decoded picture buffer 3005 to perform video encoding based on predictive encoding in the encoding unit, and generates a bitstream (also called a CP bitstream).
[0113] The CP video decoder 3004 receives the CP bitstream supplied from the CP video encoder 3002, performs entropy decoding, and then performs decoding based on predictive coding in the coding unit to obtain a decoded picture (also referred to as a CP decoded picture). The CP video decoder 3004 stores the CP decoded picture in the decoded picture buffer 3005.
[0114] The decoded pictures stored in the decoded picture buffer 3005 are used as reference pictures.
[0115] In the configuration shown in FIG. 13 , the NN video decoder 3003 receives an NN bitstream from the NN video encoder 3001 and performs entropy decoding. The NN video decoder 3003 then obtains NN-decoded pictures from the encoded data obtained by entropy decoding. However, the NN video decoder 3003 may also be configured to receive intermediate data (e.g., quantized values) before entropy encoding from the NN video encoder 3001 and obtain NN-decoded pictures from the intermediate data. In this case, the NN video decoder 3003 does not need to perform entropy decoding.
[0116] 13 , the CP video decoder 3004 receives the CP bitstream from the CP video encoder 3002 and performs entropy decoding. The CP video decoder 3004 then obtains CP-decoded pictures from the encoded data obtained by entropy decoding. However, the CP video decoder 3004 may also be configured to receive intermediate data (e.g., quantized values) before entropy encoding from the CP video encoder 3002 and obtain CP-decoded pictures from the intermediate data. In this case, the CP video decoder 3004 does not need to perform entropy decoding.
[0117] 14, the NN video decoder 3003 operates in the same manner as the NN video data processor 201 of the first embodiment or the NN video data processor 202 of the second embodiment, based on the coded data obtained by demultiplexing the NN bitstream, to obtain NN-decoded pictures. The NN video decoder 3003 stores the NN-decoded pictures in the decoded picture buffer 3005.
[0118] The CP video decoder 3004 demultiplexes the CP bitstream, performs entropy decoding, and obtains CP decoded pictures based on the resulting coded data. The CP video decoder 3004 stores the CP decoded pictures in the decoded picture buffer 3005.
[0119] The video data processor 200 outputs the NN decoded picture or the CP decoded picture stored in the decoded picture buffer 3005 as a decoded picture.
[0120] [Variation 1] In each of the above embodiments, the NN video data processors 101, 102 always perform grouping to suppress an increase in the amount of auxiliary information data. However, the NN video data processors 101, 102 may appropriately switch between control of performing grouping and transmitting a common quantization width for each group (the control in each of the above embodiments) and control of transmitting a quantization width for each quantization value.
[0121] For example, in a situation where rate control is desired to take priority over suppressing an increase in the amount of transmission related to the quantization width, the NN video data processors 101 and 102 perform control to transmit the quantization width for each quantization value.In a situation where suppressing an increase in the amount of transmission related to the quantization width is desired to take priority over rate control, the NN video data processors 101 and 102 perform control to perform grouping and transmit a common quantization width for each group.
[0122] By appropriately switching between control of grouping and transmitting a common quantization width for each group and control of transmitting a quantization width for each quantization value, the balance between the amount of auxiliary information transmitted and quantization granularity control can be optimized.
[0123] [Variation 2] In the above embodiment, the size of the CTU and the size of the FVG are the same, but they do not have to be the same. For example, the size of the FVG may be smaller than the size of the CTU, provided that CTU_tensor is a constant multiple of FVG_tensor_w and CTU_tensor is a constant multiple of FVG_tensor_h. Even in this case, the CTU contains the FVG, so the FVGs are arranged in the CTU from bottom left to bottom right.
[0124] [Variation 3] Furthermore, the size of the FVG may be larger than the size of the CTU, provided that FVG_tensor_w and FVG_tensor_h are constant multiples of CTU_tensor. In this case, the FVG encompasses the CTUs, and the CTUs are arranged in the FVG from bottom left to bottom right. Furthermore, by adding the conditions that the heights of the CTUs and FVG are the same and that the width of the FVG is equal to or less than W_tensor, the FVG unit becomes the CTU line unit. As a result, it becomes easier to adapt to both the Raster-Scan Slice mode and the Rectangular Slice mode described above.
[0125] [Variation 4] In a specific FVG, entropy coding / entropy decoding of the quantization width can be enabled for each quantization value. In this case, information indicating whether it is enabled or disabled is entropy coded for each FVG. In the enabled FVG region, for example, it is also possible to code / decode the quantization width for each quantization value based on a neural network.
[0126] [Variation 5] The above-described embodiments can also support highly efficient coding that utilizes correlations between color components (also called inter-channels). For example, the grouping unit 1002 and the ungrouping unit 2002 (or the grouping / rearrangement unit 1002A and the reverse rearrangement / ungrouping unit 2002A) may further group and process FVGs of each color that belong to the same position in space, rather than treating each color component independently.
[0127] Fig. 15 is an explanatory diagram showing an example of grouping that also takes channels into consideration. In the example shown in Fig. 15, the channels are color components (R, G, B or Y, Co, Cr) of a color space (RGB space or YCoCr space). Fig. 15 shows R, G, and B as an example. For example, if the input video is in RGB 4:4:4 format, as shown in Fig. 15, FVGs of each color that belong to the same position in space can be grouped, and the processing in the above embodiment can be performed.
[0128] Alternatively, instead of using the group quantization width for the FVGs of all colors, it is also possible to configure the system so that the group quantization width for only the FVG of a specific color is entropy coded / entropy decoded, and the group quantization width for the FVGs of the remaining colors is not entropy coded / entropy decoded. In this case, the value obtained by adding an offset set in units of picture, tile, slice, or the like to the group quantization width of the FVG of a specific color may be used as the loop quantization width for that color. In such a configuration, the auxiliary information for the quantization width can be reduced by the number of colors for which the group quantization width is not entropy coded / entropy decoded.
[0129] The grouping in the above embodiment can also be applied to color spaces other than RGB, such as YCoCr, and to formats other than the 4:4:4 format (for example, 4:2:0).
[0130] Each of the above embodiments can be configured by hardware, but can also be realized by a computer program.
[0131] The information processing system shown in Fig. 16 includes a processor 701 such as a CPU (Central Processing Unit), a program memory 702, a storage medium 703 for storing video data, and a storage medium 704 for storing a bitstream. The storage medium 703 and the storage medium 704 may be separate storage media or may be storage areas formed by the same storage medium. A magnetic storage medium such as a hard disk can be used as the storage medium.
[0132] In the information processing system, a program memory 702 stores a program (video data processing program) for realizing the functions of each block shown in each of the above embodiments.
[0133] The processor 701 then executes processing in accordance with the program stored in the program memory 702, thereby realizing the functions of the video data processor 100, NN video data processors 101 and 102, video data processor 200, and NN video data processors 201 and 202 shown in each embodiment.
[0134] For example, the functions of the NN video data processors 101, 102 and the video data processor 100 are realized by the processor 701 executing processing in accordance with a video encoding program for realizing the functions of each block (excluding the decoded picture buffer 3005) in the NN video data processors 101, 102 and the video data processor 100 shown in Figures 3, 8 and 13.
[0135] Furthermore, for example, the functions of the NN video data processors 201, 202 and the video data processor 200 are realized by the processor 701 executing processing in accordance with a video decoding program for realizing the functions of each block (excluding the decoded picture buffer 3005) in the NN video data processors 201, 202 and the video data processor 200 shown in Figures 3, 8 and 14.
[0136] At least the program memory 702 is a non-transitory computer-readable medium. However, the program may be stored in various types of transitory computer-readable medium. The program is supplied to the transitory computer-readable medium, for example, via a wired or wireless communication channel, i.e., via an electrical signal, an optical signal, or an electromagnetic wave.
[0137] 17 is a block diagram showing the main components of a video data processing device. The video data processing device 10 shown in Fig. 17 (implemented by a video data processor 100 and NN video data processors 101 and 102 in the embodiment) includes a neural network 11 (implemented by an encoder 1001 in the embodiment), a quantization means 12 (implemented by a quantizer 1003 in the embodiment), and an entropy coding means 13 (implemented by an entropy coder 1005 in the embodiment), and further includes grouping means 14 (implemented by a grouping unit 1002 or a grouping / rearrangement unit 1002A in the embodiment) that groups the outputs of the neural network 11 (as one group), and the quantization means 12 quantizes each group using a common quantization width.
[0138] The video data processing device 10 may include a sorting means (implemented by the grouping / sorting unit 1002A in this embodiment) for sorting the outputs of the neural network.
[0139] 18 is a block diagram showing the main components of another aspect of a video data processing device. A video data processing device 20 shown in Fig. 18 (implemented by a video data processor 200 and NN video data processors 201 and 202 in the embodiment) includes a neural network 21 (implemented by a decoder 2001 in the embodiment), inverse quantization means 22 (implemented by an inverse quantizer 2003 in the embodiment), and entropy decoding means 23 (implemented by an entropy decoder 2005 in the embodiment), and the inverse quantization means 22 includes degrouping means (implemented by a degrouping unit 2002 or a reverse sorting / degrouping unit 2002A in the embodiment) that performs inverse quantization for each group using a common quantization width and degroups the outputs of the inverse quantization means 22.
[0140] The video data processing device 20 may also include an inverse sorting means (implemented by an inverse sorting / degrouping unit 2002A in this embodiment) that sorts the output of the inverse quantization means 22.
[0141] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes.
[0142] (Supplementary Note 1) A video data processing device comprising a neural network, a quantization means, and an entropy coding means, further comprising a grouping means for grouping outputs of the neural network, wherein the quantization means performs quantization for each group using a common quantization width.
[0143] (Supplementary Note 2) The video data processing device according to Supplementary Note 1, further comprising a sorting means for sorting outputs of the neural network.
[0144] (Supplementary Note 3) A video data processing device comprising a neural network, an inverse quantization means, and an entropy decoding means, wherein the inverse quantization means performs inverse quantization for each group using a common quantization width, and the video data processing device further comprises a degrouping means for degrouping outputs of the inverse quantization means.
[0145] (Supplementary Note 4) The video data processing device according to Supplementary Note 3, further comprising an inverse rearrangement means for rearranging the output of the inverse quantization means.
[0146] (Supplementary Note 5) A video data processing method that performs encoding processing, quantization processing, and entropy coding based on a neural network, comprising: grouping outputs of the neural network; and performing quantization for each group using a common quantization width in the quantization processing.
[0147] (Supplementary Note 6) The video data processing method of Supplementary Note 5, further comprising rearranging the outputs of the neural network.
[0148] (Supplementary Note 7) A video data processing method that executes a neural network-based decoding process, an inverse quantization process, and an entropy decoding process, wherein in the inverse quantization process, inverse quantization is performed for each group using a common quantization width, and the outputs of the inverse quantization process are ungrouped.
[0149] (Supplementary Note 8) The video data processing method according to Supplementary Note 7, wherein the outputs of the inverse quantization process are rearranged.
[0150] (Supplementary Note 9) A video data processing program for making a computer execute an encoding process, a quantization process, and an entropy coding process based on a neural network, execute a process of grouping the outputs of the neural network, and perform quantization for each group using a common quantization width in the quantization process.
[0151] (Supplementary Note 10) The video data processing program of Supplementary Note 9, which causes a computer to execute a process of rearranging outputs of the neural network.
[0152] (Supplementary Note 11) A video data processing program for causing a computer to execute a neural network-based decoding process, an inverse quantization process, and an entropy decoding process, causing the inverse quantization process to perform inverse quantization using a common quantization width for each group, and causing the output of the inverse quantization process to be ungrouped.
[0153] (Supplementary Note 12) The video data processing program according to Supplementary Note 11, which causes a computer to execute a process of rearranging outputs of the inverse quantization process.
[0154] (Supplementary Note 13) A storage medium for storing a bitstream generated by a video data processing device comprising a neural network, a quantization means, and an entropy coding means, the storage medium comprising a grouping means for grouping outputs of the neural network, the quantization means performing quantization for each group using a common quantization width.
[0155] (Supplementary Note 14) A storage medium for storing a bitstream generated by a video data processing method that performs encoding processing, quantization processing, and entropy coding based on a neural network, in which outputs of the neural network are grouped and quantization is performed for each group using a common quantization width in the quantization processing.
[0156] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
[0157] This application claims priority based on Japanese Patent Application No. 2023-187757, filed November 1, 2023, the disclosure of which is incorporated herein in its entirety.
[0158] REFERENCE SIGNS LIST 10 Video data processing device 11 Neural network 12 Quantization means 13 Entropy coding means 14 Grouping means 20 Video data processing device 21 Neural network 22 Inverse quantization means 23 Entropy decoding means 24 Input means 100 Video data processor 101, 102 NN video data processor 200 Video data processor 201, 202 NN video data processor 701 Processor 702 Program memory 703, 704 Storage medium 1001 Encoder 1002 Grouping section 1002A Grouping / rearranging section 1003 Quantizer 1005 Entropy coder 1006 Multiplexer 1007 Rate control section 2001 Decoder 2002 Degrouping section 2002A Inverse rearrangement / degrouping section 2003 Inverse quantizer 2005 Entropy decoder 2006 Demultiplexer 3000 Switch 3001 NN video encoder 3002 CP video encoder 3003 NN video decoder 3004 CP video decoder 3005 Decoded picture buffer 4000 Multiplexer 5000 Demultiplexer
Claims
1. A video data processing device comprising a neural network, a quantization means, and an entropy coding means, further comprising a grouping means for grouping outputs of the neural network, wherein the quantization means performs quantization for each group using a common quantization width.
2. The video data processing device according to claim 1, further comprising a sorting means for sorting outputs of said neural network.
3. A video data processing device comprising a neural network, an inverse quantization means, and an entropy decoding means, wherein the inverse quantization means performs inverse quantization for each group using a common quantization width, and the video data processing device further comprises a degrouping means for degrouping the output of the inverse quantization means.
4. The video data processing device according to claim 3, further comprising an inverse sorting means for sorting the output of said inverse quantization means.
5. A video data processing method that performs encoding, quantization, and entropy coding based on a neural network, comprising: grouping outputs of the neural network; and performing quantization for each group using a common quantization width in the quantization process.
6. The video data processing method according to claim 5, further comprising the step of: sorting the outputs of said neural network.
7. A video data processing method for executing a decoding process, an inverse quantization process, and an entropy decoding process based on a neural network, wherein in the inverse quantization process, inverse quantization is performed for each group using a common quantization width, and the output of the inverse quantization process is ungrouped.
8. The video data processing method according to claim 7, further comprising rearranging the outputs of the inverse quantization process.
9. A video data processing program for causing a computer to execute encoding processing, quantization processing, and entropy coding based on a neural network, execute processing to group the outputs of the neural network, and perform quantization using a common quantization width for each group in the quantization processing.
10. A video data processing program for causing a computer to execute a decoding process, an inverse quantization process, and an entropy decoding process based on a neural network, in which the inverse quantization process performs inverse quantization using a common quantization width for each group, and for degrouping the output of the inverse quantization process.
Citation Information
Patent Citations
Quantization and encoder creation method, compressor creation method, compressor creation apparatus, and program
JP2020053820A
Tool selection for feature map encoding VS regular video encoding
WO2022213139A1