Video encoding device, video decoding device, video encoding method, and video decoding method
By constraining encoding processes in a hybrid video encoding device that combines neural network and predictive methods, the issue of interconnectivity and performance mismatch is resolved, ensuring efficient decoding.
Patent Information
- Application Number
- PCT/JP2024/043155
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-10
AI Technical Summary
The interconnectivity between video encoding devices using predictive encoding and neural network-based encoding is compromised due to discrepancies in the maximum amount of codes generated and slice division requirements, leading to potential performance issues in decoding devices.
Implementing a video encoding device that combines neural network-based and predictive encoding methods with constraints on the encoding process, such as limiting the maximum code amount and slice size, to ensure compatibility and performance requirements are met by the decoding device.
Ensures seamless interconnectivity and efficient decoding by constraining the encoding process, preventing excessive performance demands on the decoding device and maintaining video quality.
Smart Images

Figure JP2024043155_10072025_PF_FP_ABST
Abstract
Description
Video encoding device, video decoding device, video encoding method, and video decoding method
[0001] The present disclosure relates to a video encoding device, a video decoding device, a video encoding method, and a video decoding method.
[0002] In order to transmit or record video efficiently, a video encoding device is used to generate a coded representation (hereinafter referred to as a bitstream) of input video, and a video decoding device is used to decode the bitstream to generate decoded video.
[0003] [Video Coding Based on Predictive Coding in Coding Units] Video coding standards include H.264 / AVC (Advanced Video Coding), H.265 / HEVC (High-Efficiency Video Coding), and H.266 / VVC (Versatile Video Coding), which are standardized by ITU-T SG16 and ISO / IEC / SC29 (see, for example, Non-Patent Document 1).
[0004] In these video coding methods, video data is coded and decoded while being managed in a hierarchical structure. The hierarchical structure is made up of, for example, pictures constituting the video data, slices obtained by dividing the pictures, coding tree units (CTUs) obtained by dividing the slices, and coding units (CUs) obtained by dividing the coding tree units.
[0005] In these video coding methods, tiles, slices, and subpictures are defined as units for dividing a picture. A picture is divided into one or more tiles. A tile is a rectangular area whose constituent unit is a CTU. A slice is a rectangular area whose constituent unit is a tile. There are two scan orders for slices: raster-scan slice mode and rectangular slice mode. Raster-scan slice mode is a mode in which slices are arranged in raster scan order. Rectangular slice mode is a mode in which the coverage area of a slice is a rectangular area whose unit is a tile or a CTU line within a tile. A subpicture is made up of one or more slices.
[0006] The input image of the target CU is usually encoded earlier than the target CU and is predictively coded based on a predicted image generated based on the decoded image that has been decoded. That is, a prediction error image obtained by subtracting the predicted image from the input image is coded and decoded. Predictive coding includes intra-picture prediction (intra-prediction) that uses a decoded image included in a picture with the same display time as the target CU, and inter-picture prediction (inter-prediction) that uses a decoded image included in a picture with a different display time from the target CU.
[0007] The prediction error image is encoded based on frequency transform, quantization, and entropy coding. The prediction error image is decoded based on entropy decoding, inverse quantization, and inverse frequency transform. The frequency transform value of the quantized prediction error image is called a quantized value.
[0008] [Video Coding Based on Neural Networks] Non-Patent Document 2 describes a new video coding technique that combines an auto-encoder, which is a type of neural network, quantization, and entropy coding.
[0009] An autoencoder compresses input data into a low-dimensional feature vector that contains only important features. The autoencoder then generates reconstructed data by reconstructing the low-dimensional feature vector back to its original dimensions. Figure 1 is an explanatory diagram showing the autoencoder algorithm. In Figure 1, the circular parts are called nodes and the arrows are called edges. The process of reducing the data into a low-dimensional feature vector (the first half) is called encoding. The process of generating reconstructed data (the second half) is called decoding.
[0010] The autoencoder is trained to minimize the reconstruction error (the difference between the input data and the reconstructed data). To obtain meaningful features, the autoencoder is designed to impose constraints on the encoding structure and to add regularization terms to the network's loss function.
[0011] Recommendation ITU-T H.266 "Versatile video coding", Telecommunication Standardization Sector of ITU, August 2020 J. Ball'e, V. Laparra, and EP Simoncelli, "End-to-end Optimized Image Compression", published as a conference paper at ICLR 2017
[0012] A video encoding device that can realize both video encoding using a video encoding method based on predictive encoding in the encoding unit (predictive encoding processing on a coding unit basis) and video encoding using a video encoding method based on a neural network, and a video decoding device that can realize both video decoding using a video decoding method based on predictive encoding in the encoding unit and video decoding using a video decoding method based on a neural network, are conceivable.
[0013] In a system including such a video encoding device and a video decoding device, there is a risk that the interconnectivity between the video encoding device and the video decoding device may not be ensured. Note that the ability to perform encoding and decoding using both a video coding method based on predictive coding in a coding unit and a video coding method based on a neural network is sometimes referred to as a combination of video coding based on predictive coding in a coding unit and video coding based on a neural network.
[0014] As an example, there is a problem caused by a difference in the maximum code amount of one picture (the maximum code amount generated from one picture by the encoding process). Specifically, there is a risk that the maximum code amount when encoding using a video encoding method based on predictive coding will differ greatly from the maximum code amount when encoding using a video decoding method based on a neural network. If the two maximum code amounts differ greatly, excessive performance is expected to be required of the video decoding device. In this case, if the performance of the video decoding device is insufficient, problems such as the video decoding device being unable to obtain good decoded video may occur. This problem becomes more pronounced when the code amount when encoding using a video decoding method based on a neural network is large.
[0015] Another example is a problem caused by the division of a picture into slices. When performing video encoding based on a neural network, if a predetermined minimum processing size (e.g., 128 × 128) is required, and division allows the generation of slices smaller than the minimum processing size, the video decoding device may be required to perform exceptional processing. In this case, there is also a risk that the interoperability between the video encoding device and the video decoding device may not be ensured.
[0016] The present invention aims to ensure interoperability between a video encoding device and a video decoding device when video encoding based on predictive encoding in a coding unit and video encoding based on a neural network are combined.
[0017] A video encoding device based on the present disclosure is a video encoding device that performs video encoding using a first video encoding method based on a neural network and video encoding using a second video encoding method based on predictive encoding in an encoding unit, and includes a constraint means that imposes constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0018] A video decoding device based on the present disclosure is a video decoding device that performs video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in an encoding unit, and includes a bitstream decoding means that decodes a bitstream generated by imposing constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0019] A video encoding method according to the present disclosure is a video encoding method that performs video encoding using a first video encoding method based on a neural network and video encoding using a second video encoding method based on predictive encoding in a coding unit, and imposes constraints on the encoding process of pictures that are video encoded using the first video encoding method.
[0020] A video decoding method according to the present disclosure performs video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in an encoding unit, and decodes a bitstream generated by imposing constraints on the encoding process of a picture to be video encoded using the first video encoding method.
[0021] A video encoding program according to the present disclosure is a video encoding program that causes a computer to perform video encoding using a first video encoding method based on a neural network and video encoding using a second video encoding method based on predictive encoding in an encoding unit, and causes the computer to perform a process of imposing constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0022] A video decoding program according to the present disclosure is a video decoding program that causes a computer to perform video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in an encoding unit, and causes the computer to perform a process of decoding a bitstream generated by imposing constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0023] According to the present invention, the interconnectivity between the video encoding device and the video decoding device is ensured.
[0024] FIG. 1 is an explanatory diagram showing an algorithm of an autoencoder. FIG. 2 is a block diagram showing an example of the configuration of a video encoding device. FIG. 3 is a block diagram showing an example of the configuration of a video decoding device. FIG. 4 is an explanatory diagram showing an example of the relationship between the image size of an input picture and the size of a feature vector. FIG. 5 is a flowchart showing an example of the operation of a video encoding device. FIG. 6 is a flowchart showing an example of the operation of a video decoding device. FIG. 7 is an explanatory diagram showing an example of slice division. FIG. 8 is an explanatory diagram showing an example of a non-rectangular slice. FIG. 9 is an explanatory diagram showing an example of an expression of a constraint for prohibiting the division of a picture into slices. FIG. 10 is a block diagram showing an example of the configuration of an encoder and a decoder. FIG. 11 is a block diagram showing an example of the configuration of an information processing system. FIG. 12 is a block diagram showing the main parts of a video encoding device. FIG. 13 is a block diagram showing the main parts of a video decoding device.
[0025] Hereinafter, an embodiment will be described with reference to the drawings.
[0026] First Embodiment Fig. 2 is a block diagram showing an example of the configuration of a video encoding device. Fig. 3 is a block diagram showing an example of the configuration of a video decoding device.
[0027] 2 and 3 simply indicate the direction of signal (data) flow, but do not exclude bidirectionality. This also applies to other block diagrams.
[0028] The video encoding device 100 shown in FIG. 2 includes a switch 300, an NN video encoder 301A, an NN video decoder 302A, an NN video encoding controller 400, a CP video encoder 301B, a CP video decoder 302B, a decoded picture buffer 303, and a multiplexer 500 that performs multiplexing processing of entropy-encoded data and other information.
[0029] Note that "NN" stands for neural network, and "CP" stands for coding unit-based prediction.
[0030] Hereinafter, the part of video encoding device 100 that performs video encoding based on a neural network may be referred to as an NN video encoder. The part of video encoding device 100 that performs video encoding based on predictive encoding may be referred to as a CP video encoder.
[0031] The video encoding device 100 shown in FIG. 2 is a video encoder that combines an NN video encoder that performs video encoding based on a neural network and a CP video encoder that performs video encoding based on predictive encoding in a coding unit.
[0032] The video decoding device 200 shown in FIG. 3 includes a demultiplexer 600 that demultiplexes a bitstream, an NN video decoder 401A, a CP video decoder 401B, and a decoded picture buffer 403.
[0033] Hereinafter, the part of video decoding device 200 that performs video decoding based on a neural network may be referred to as an NN video decoder. The part of video decoding device 200 that performs video decoding based on predictive coding may be referred to as a CP video decoder.
[0034] The video decoding device 200 shown in FIG. 3 is a video decoding device that combines an NN video decoder that performs video decoding based on a neural network and a CP video decoder that performs predictive decoding in coding units.
[0035] [Description of the Encoding Side] In the video encoding device 100 shown in FIG. 2, the switch 300 supplies an input picture to either the NN video encoder 301A or the CP video encoder 301B.
[0036] The NN video encoder 301A includes at least an encoder, a quantizer, and an entropy encoder.
[0037] The encoder extracts features from the image of the input picture. Specifically, the encoder obtains a feature vector from the image of the input picture. The quantizer quantizes the feature vector provided from the encoder to obtain a quantized value. The entropy encoder entropy-encodes the quantized value to obtain entropy-encoded data. The entropy-encoded data from the entropy encoder is sometimes called an NN bitstream.
[0038] The NN video decoder 302A includes at least an entropy decoder, an inverse quantizer, and a decoder. The entropy decoder entropy decodes the NN bitstream to obtain quantized values. The inverse quantizer inversely quantizes the quantized values to obtain reconstructed feature vectors. The decoder obtains reconstructed images (NN decoded pictures) from the reconstructed feature vectors. The NN video decoder 302A stores the NN decoded pictures in the decoded picture buffer 303 for subsequent processing.
[0039] The NN video encoding controller 400 appropriately restricts the NN video encoder 301A in order to ensure interconnectivity between the video encoder 100 and the video decoder 200.
[0040] An example of the constraint will be described with reference to Fig. 4. One example of the constraint is a constraint on the maximum code amount (the maximum value of the code amount generated from one picture by the encoding process). Fig. 4 is an explanatory diagram showing an example of the relationship between the image size of an input picture and the size of a feature vector. The constraint described with reference to Fig. 4 is controlled by the NN video encoding controller 400 in the video encoding device 100.
[0041] An example of an input picture is shown on the left side of Fig. 4. In Fig. 4, img_W indicates the width (horizontal size) of the input picture, and img_H indicates the height (vertical size) of the input picture.
[0042] An example of a feature vector obtained by encoding is shown on the right side of Fig. 4. In Fig. 4, tensor_W indicates the row size of the feature vector, and tensor_H indicates the column size of the feature vector.
[0043] The NN video encoding controller 400 specifies the maximum code amount for one input picture processed by the NN video encoder 301A based on the size of the feature vector (tensor_W × tensor_H) rather than the image size (img_W × img_H) of the input picture. Specifically, the maximum code amount is set to a number of bytes proportional to the value of (tensor_W × tensor_H) / MinCrBase. In other words, the maximum code amount is restricted by the size of the feature vector (the dimension of the feature vector).
[0044] MinCrBase is described in, for example, A.4.2 Profile-specific level limits and Table A.2 in Non-Patent Document 1. MinCrBase is a value related to the minimum compression ratio (MinCr).
[0045] As described in A.4.2 of Non-Patent Document 1, the maximum code size of one input picture processed by the CP video encoder 301B is proportional to the value of (img_W×img_H) / MinCrBase, which is determined based on the image size (img_W×img_H) of the input picture and MinCrBase.
[0046] Because the encoder included in the NN video encoder 301A performs dimensional reduction, if the maximum code amount is defined based on img_W × img_H, the compression rate of the NN bitstream will be set low. As a result, the processing performance required of the video decoding device will be high. However, in this embodiment, the NN video encoding controller 400 appropriately restricts the maximum code amount of a picture encoded by neural network-based video encoding. Therefore, the processing performance required of the video decoding device is prevented from becoming excessively high. As a result, interoperability between the video encoding device and the video decoding device is ensured.
[0047] The CP video encoder 301B performs video encoding processing using a video encoding method based on predictive coding for each coding unit. The CP video decoder 302B performs decoding processing using a video encoding method based on predictive coding for each coding unit. The decoded picture buffer 303 is a storage unit that stores decoded pictures (reconstructed pictures). Video encoding methods that can be used include methods compliant with H.264 / AVC, H.265 / HEVC, H.266 / VVC, etc.
[0048] The CP video encoder 301B includes at least a predictor, a frequency transformer, a quantizer, and an entropy encoder.
[0049] The CP video encoder 301B uses the input picture and the decoded picture stored in the decoded picture buffer 303 to perform video encoding based on predictive encoding in the encoding unit to generate a bitstream (also called a CP bitstream).
[0050] CP video decoder 302B includes at least an entropy decoder, an inverse quantizer, an inverse frequency transformer, and a predictive decoder.
[0051] The CP video decoder 302B receives the CP bitstream supplied from the CP video encoder 301B, performs entropy decoding, and then performs decoding based on predictive coding in the coding unit to obtain a decoded picture (also referred to as a CP decoded picture). The CP video decoder 302B stores the CP decoded picture in the decoded picture buffer 303.
[0052] The decoded pictures stored in the decoded picture buffer 303 are used as reference pictures.
[0053] Next, a description will be given of the operation of the video encoding device 100. Fig. 5 is a flowchart showing an example of the operation of the video encoding device 100. Fig. 5 shows the operation of the NN video encoder 301A in the video encoding device 100.
[0054] The encoder included in the NN video encoder 301A extracts features from an input picture and generates a feature vector (step S101).
[0055] As described above, the NN video encoding controller 400 imposes constraints on the encoding process of the NN video encoder 301A. That is, the NN video encoding controller 400 provides a predetermined maximum code amount per picture to the NN video encoder 301A. The NN video encoder 301A generates low-dimensional feature vectors so as not to exceed the maximum code amount.
[0056] The quantizer included in the NN video encoder 301A quantizes the feature vector to generate a quantized value (step S102).
[0057] The entropy encoder included in the NN video encoder 301A entropy encodes the quantized value to generate entropy-encoded data (step S103).
[0058] In the video encoding device 100, the multiplexer 500 outputs the entropy-encoded data as an NN bit stream (step S106).
[0059] In the configuration shown in FIG. 2 , NN video decoder 302A receives an NN bitstream from NN video encoder 301A and performs entropy decoding. Then, NN video decoder 302A obtains NN-decoded pictures from the encoded data obtained by entropy decoding. However, NN video decoder 302A may also be configured to receive intermediate data (e.g., quantized values) before entropy encoding from NN video encoder 301A and obtain NN-decoded pictures from the intermediate data. In this case, NN video decoder 302A does not need to perform entropy decoding.
[0060] 2, the CP video decoder 302B receives the CP bitstream from the CP video encoder 301B and performs entropy decoding. The CP video decoder 302B then obtains CP-decoded pictures from the encoded data obtained by entropy decoding. However, the CP video decoder 302B may also be configured to receive intermediate data (e.g., quantized values) before entropy encoding from the CP video encoder 301B and obtain CP-decoded pictures from the intermediate data. In this case, the CP video decoder 302B does not need to perform entropy decoding.
[0061] [Description on the Decoding Side] In the NN video decoding device 200 shown in FIG. 3, the demultiplexer 600 demultiplexes the NN bitstream to obtain entropy-encoded data.
[0062] The NN video decoder 401A includes at least an entropy decoder, an inverse quantizer, and a decoder. The NN video decoder 401A obtains NN-decoded pictures based on coded data obtained by demultiplexing the NN bitstream. The NN video decoder 401A stores the NN-decoded pictures in the decoded picture buffer 403.
[0063] CP video decoder 401B includes at least an entropy decoder, an inverse quantizer, an inverse frequency transformer, and a predictive decoder.
[0064] The CP video decoder 401B obtains a CP decoded picture based on the coded data obtained by entropy decoding the CP bitstream, and stores the CP decoded picture in the decoded picture buffer 403.
[0065] The video decoding device 200 outputs the NN decoded picture or the CP decoded picture stored in the decoded picture buffer 403 as a decoded picture at the display timing embedded in the bitstream.
[0066] Next, a description will be given of the operation of the video decoding device 200. Fig. 6 is a flowchart showing an example of the operation of the video decoding device 200. Fig. 6 shows the operation of the NN video decoder 401A in the video decoding device 200.
[0067] In the NN video decoding device 200, the demultiplexer 600 demultiplexes the bitstream (step S201). The demultiplexer 600 obtains entropy-encoded data through demultiplexing.
[0068] The entropy decoder included in the NN video decoder 401A entropy decodes the entropy-encoded data to obtain quantized values (step S202).
[0069] The inverse quantizer included in the NN video decoder 401A inversely quantizes the quantized value (step S203). The inverse quantizer obtains a reconstructed feature vector through inverse quantization.
[0070] The decoder included in the NN video decoder 401A obtains a reconstructed image of the decoded picture from the reconstructed feature vector (step S204).
[0071] [Constraint Example 2] In the above embodiment, the NN video encoding controller 400 constrains the maximum code size of the feature vector generated by the NN video encoder 301A (this constraint is referred to as Constraint Example 1). The NN video encoding controller 400 may handle other constraints on the encoding process instead of or in addition to such a constraint.
[0072] For example, the NN video encoding controller 400 specifies a minimum size of a feature vector of a slice region and controls the NN video encoder 301A so that slice division does not result in slices of a size less than the minimum size. That is, the NN video encoding controller 400 provides the NN video encoding controller 400 with a predetermined minimum size. The NN video encoding controller 400 performs slice division so that the size of slices generated by division will not be less than the minimum size.
[0073] FIG. 7 is an explanatory diagram showing an example of slice division.
[0074] 7, img_W indicates the width (horizontal size) of the input picture, img_H indicates the height (vertical size) of the input picture, tensor_W indicates the row size of the feature vector, and tensor_H indicates the column size of the feature vector.
[0075] 7 shows an example in which the input picture is divided vertically equally into three slices. However, this division method is merely an example, and the division method is not limited thereto. For example, the input picture may be divided unevenly. Furthermore, the number of divisions may be two, four, or more.
[0076] When dividing an input picture into slices, the encoder included in the NN video encoder 301A extracts features for each slice obtained by the division and generates a feature vector. The quantizer included in the NN video encoder 301A quantizes the feature vector for each slice to generate a quantized value.
[0077] As described above, when performing neural network-based video encoding, if a predetermined minimum processing size (e.g., 128 × 128) is required, and if the division allows the generation of slices smaller than the minimum processing size, the video decoder may be required to perform exceptional processing, which may result in a risk of not ensuring interoperability between the video encoder and the video decoder.
[0078] However, in this example, the NN video encoding controller 400 controls the NN video encoder 301A so that slice division does not result in slices smaller than the minimum size, thereby ensuring interoperability between the video encoding device and the video decoding device.
[0079] The requirement for a minimum processing size may arise, for example, when the entropy encoder included in the NN video encoder 301A entropy encodes quantized values to obtain entropy-coded data, and uses a process of thinning out the quantized values.
[0080] In constraint example 3, an example of restricting the slice division process (restricting the form of each slice generated by division) in which slices of a size less than the minimum size are not generated has been shown, but constraints on the slice division process are not limited to this. For example, as will be described later, the NN video encoding controller 400 may restrict the shape of the slice division (the shape of each slice generated by division) itself. Note that, in a broad sense, the constraint of not generating slices of a size less than the minimum size is also included in restricting the shape of the slice division.
[0081] [Constraint Example 3] In Constraint Example 1, the maximum code amount for one picture is defined based on the size of the feature vector (tensor_W × tensor_H). Specifically, it is defined based on the value of (tensor_W × tensor_H) / MinCrBase. However, the NN video encoding controller 400 may obtain a value equivalent to (tensor_W × tensor_H) / MinCrBase using another value of MinCrBase (hereinafter referred to as MinCrBase'). For example, the NN video encoding controller 400 may set the maximum code amount to a number of bytes proportional to the value of MinCrBase' = MinCrBase × (tensor_W × tensor_H) / (img_W × img_H).
[0082] In Constraint Example 4, the NN video encoding controller 400 constrains the minimum size of slice division of a picture to be coded by neural network-based video coding. The NN video encoding controller 400 may further constrain the shape of the slices after the division process to be non-rectangular.
[0083] 8 is an explanatory diagram showing an example of a non-rectangular slice, in which the area surrounded by a dashed line corresponds to the non-rectangular slice.
[0084] For example, H.266 / VVC has a mode that allows the use of non-rectangular slices. When a video encoding device performs encoding processing using non-rectangular slices, a video decoding device that inputs a bitstream generated by the encoding processing may be required to perform exceptional processing. In this case, there is a risk that interoperability between the video encoding device and the video decoding device may not be guaranteed.
[0085] However, by restricting the use of non-rectangular slices, interoperability between video encoding devices and video decoding devices is ensured.
[0086] For example, since the above-described Raster-Scan Slice mode allows the use of non-rectangular slices, the NN video encoding controller 400 may prohibit the use of the Raster-Scan Slice mode.
[0087] Another embodiment shows an embodiment in which the NN video coding controller 400 prohibits dividing a picture into slices to be coded by neural network-based video coding. Figure 9 is an explanatory diagram showing an example of an expression of a constraint for prohibiting dividing a picture into slices.
[0088] In this embodiment, the NN video encoding controller 400 controls the NN video encoder 301A illustrated in Fig. 2 so as not to perform slice division. This is because if slice division were permitted, the video decoding device would need to additionally control slice division in pictures encoded by neural network-based video encoding, making it difficult to ensure interconnectivity between the encoder and decoder.
[0089] When VVC is used as the video coding method for the CP video encoder 301B, the prohibition in this embodiment can be expressed by restricting the value of the pps_no_pic_partition_flag syntax of pic_parameter_set_rbsp( ) as shown in Figure 9. Note that in Figure 9, the parts related to the prohibition of dividing a picture into slices are underlined to make them easier to understand.
[0090] In this embodiment, the NN video coding controller 400 eliminates the need to divide pictures into slices when encoding using neural network-based video coding, thereby ensuring interconnectivity between the encoder and decoder. Note that pictures that require slice division can be processed by the CP video encoder 301B instead of the NN video encoder 301A.
[0091] [Encoder and Decoder Configurations] Fig. 10 is a block diagram showing example configurations of the encoder 1001 included in the NN video encoder 301A and the decoder 2001 included in the NN video decoder 302A. In Fig. 10, a "downward arrow (↓) 2" represents subsampling by 1 / 2 (also called pooling). An "upward arrow (↑) 2" represents upsampling by 2. Note that the encoder 1001A and decoder 2001A in the second embodiment can also be configured as shown in Fig. 10.
[0092] In the example shown in Fig. 10, the encoder 1001 is composed of four residual blocks and one convolution block. Each residual block is composed of two convolutional blocks and one shortcut link. Each convolutional block is composed of one convolution layer and one activation function.
[0093] The decoder 2001 consists of four residual blocks and one pixel shuffle convolution layer. The pixel shuffler is a mechanism proposed as sub-pixel convolution. The pixel shuffler rearranges input feature vectors and outputs high-resolution feature vectors.
[0094] As an activation function, a Parametric ReLU (Parametric Rectified Linear Unit) can be used, in which the output value is α times the input value when the input value is below 0 (where α is a parameter determined by learning), and the output value is the same as the input value when the input value is 0 or greater.
[0095] Furthermore, the configuration shown in FIG. 10 is just an example, and the configurations of the encoder 1001 and the decoder 2001 are not limited to the configuration shown in FIG.
[0096] Each of the above embodiments can be configured by hardware, but can also be realized by a computer program.
[0097] The information processing system shown in Fig. 11 includes a processor 701 such as a CPU (Central Processing Unit), a program memory 702, a storage medium 703 for storing video data, and a storage medium 704 for storing a bitstream. The storage medium 703 and the storage medium 704 may be separate storage media or may be storage areas of the same storage medium. A magnetic storage medium such as a hard disk can be used as the storage medium.
[0098] In the information processing system, a program memory 702 stores a program (video encoding program or video decoding program) for implementing the functions of each block shown in each of the above embodiments.
[0099] The processor 701 then executes processing in accordance with the program stored in the program memory 702, thereby realizing the functions of the video encoding device 100 and the video decoding device 200 described in the above embodiments.
[0100] For example, the processor 701 executes processing in accordance with a video encoding program for realizing the functions of the NN video encoder 301A, the CP video encoder 301B, the NN video decoder 302A, the CP video decoder 302B, the NN video encoding controller 400, and the multiplexer 500 shown in FIG. 2, thereby realizing the functions of the video encoding device 100.
[0101] Also, for example, the processor 701 executes processing in accordance with a video decoding program for realizing the functions of the NN video decoder 401A, the CP video decoder 401B, and the demultiplexer 600 shown in FIG. 3, thereby realizing the functions of the video decoding device 200.
[0102] At least the program memory 702 is a non-transitory computer-readable medium. However, the program may be stored in various types of transitory computer-readable medium. The program is supplied to the transitory computer-readable medium, for example, via a wired or wireless communication channel, i.e., via an electrical signal, an optical signal, or an electromagnetic wave.
[0103] Fig. 12 is a block diagram showing the main components of a video encoding device. The video encoding device 10 shown in Fig. 12 (implemented as a video encoding device 100 in the embodiment) performs video encoding using a first video encoding method based on a quantized neural network and video encoding using a second video encoding method based on predictive encoding in a coding unit, and includes a constraint means (constraint unit) 11 (implemented as an NN video encoding controller 400 in the embodiment) that imposes constraints on the encoding process of pictures that are video encoded using the first video encoding method.
[0104] Fig. 13 is a block diagram showing the main components of a video decoding device. The video decoding device 20 shown in Fig. 13 (implemented as a video decoding device 200 in the embodiment) includes bitstream decoding means (bitstream decoding unit) 21 (implemented as a demultiplexer 600 in the embodiment) that performs video decoding using a first video coding scheme based on a neural network and video decoding using a second video coding scheme based on predictive coding in a coding unit, and decodes a bitstream generated by imposing constraints on the coding process of a picture that is video coded using the first video coding scheme.
[0105] Some or all of the above embodiments can be described as, but are not limited to, the following supplementary notes.
[0106] (Supplementary Note 1) A video encoding device that performs video encoding using a first video encoding method based on a neural network and video encoding using a second video encoding method based on predictive encoding in an encoding unit, the video encoding device comprising: a constraint means that imposes constraints on an encoding process of a picture that is video encoded using the first video encoding method.
[0107] (Supplementary Note 2) The video encoding device according to Supplementary Note 1, wherein the restricting means restricts a maximum code amount of a picture encoded in the first video encoding format based on a dimension of a feature vector.
[0108] (Supplementary Note 3) The video encoding device according to Supplementary Note 1 or Supplementary Note 2, wherein the restriction means restricts a shape of slice division of a picture to be video encoded by the first video encoding method.
[0109] (Supplementary Note 4) A video decoding device that performs video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in an encoding unit, the video decoding device comprising: a bitstream decoding means that decodes a bitstream generated by imposing constraints on an encoding process of a picture that is video encoded using the first video encoding method.
[0110] (Supplementary Note 5) The video decoding device according to Supplementary Note 4, wherein the constraint is a constraint on a maximum code amount of a picture to be video-coded in the first video coding format, the constraint being based on a dimension of a feature vector.
[0111] (Supplementary Note 6) The video decoding device according to Supplementary Note 4 or Supplementary Note 5, wherein the constraint is a constraint on the shape of slice division of a picture video-encoded by the first video encoding format.
[0112] (Supplementary Note 7) A video coding method that performs video coding using a first video coding method based on a neural network and video coding using a second video coding method based on predictive coding in a coding unit, wherein the video coding method imposes constraints on a coding process of a picture that is video coded using the first video coding method.
[0113] (Supplementary Note 8) A video decoding method that performs video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in a coding unit, the video decoding method decoding a bitstream generated by imposing constraints on an encoding process of a picture that is video encoded using the first video encoding method.
[0114] (Supplementary Note 9) A video encoding program that causes a computer to perform video encoding using a first video encoding method based on a neural network and video encoding using a second video encoding method based on predictive encoding in an encoding unit, the video encoding program causing the computer to perform a process of imposing constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0115] (Supplementary Note 10) A video decoding program that causes a computer to perform video decoding using a first video encoding method based on a neural network and video decoding using a second video encoding method based on predictive encoding in an encoding unit, the video decoding program causing the computer to perform a process of decoding a bitstream generated by imposing constraints on the encoding process of a picture that is video encoded using the first video encoding method.
[0116] (Supplementary Note 11) The video encoding method according to Supplementary Note 7, wherein a maximum code amount of a picture encoded by the first video encoding system is restricted based on a dimension of a feature vector.
[0117] (Supplementary Note 12) The video coding method according to Supplementary Note 7 or Supplementary Note 11, wherein a shape of slice division of a picture coded by the first video coding system is constrained.
[0118] (Supplementary Note 13) The video decoding method of Supplementary Note 8, wherein the constraint is a constraint on a maximum code amount of a picture video-encoded by the first video encoding method, the constraint being based on a dimension of a feature vector.
[0119] (Supplementary Note 14) The video decoding method according to Supplementary Note 8 or Supplementary Note 13, wherein the constraint is a constraint on the shape of slice division of a picture to be video coded by the first video coding system.
[0120] Some or all of the configurations described in Supplementary Notes 2 and 3, which depend directly or indirectly on Supplementary Note 1, may be made to depend on various hardware, software, various recording means for recording software, or systems, provided that the above-mentioned embodiments are not deviated from. Also, some or all of the configurations described in Supplementary Notes 5 and 6, which depend directly or indirectly on Supplementary Note 4, may be made to depend on various hardware, software, various recording means for recording software, or systems, provided that the above-mentioned embodiments are not deviated from.
[0121] Although the present invention has been described above with reference to the embodiments and examples, the present invention is not limited to the above-described embodiments and examples. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
[0122] This application claims priority based on Japanese Patent Application No. 2024-000581 filed on January 5, 2024, and Japanese Patent Application No. 2024-035448 filed on March 8, 2024, the disclosures of which are incorporated herein in their entireties.
[0123] 10 Video encoding device 11 Restriction means 20 Video decoding device 21 Bitstream decoding means 100 Video encoding device 200 Video decoding device 300 Switch 301A NN video encoder 301B CP video encoder 302A NN video decoder 302B CP video decoder 303 Decoded picture buffer 400 NN video encoding controller 401A NN video decoder 401B CP video decoder 403 Decoded picture buffer 500 Multiplexer 600 Demultiplexer 701 Processor 702 Program memory 703, 704 Storage medium 1001 Encoder 2001 Decoder
Claims
1. A video encoding apparatus that performs video encoding by a first video encoding method based on a neural network and video encoding by a second video encoding method based on predictive encoding in an encoding unit, the video encoding apparatus comprising constraint means for imposing a constraint on the encoding process of a picture encoded by the first video encoding method.
2. The video encoding apparatus according to claim 1, wherein the constraint means constrains the maximum code amount of a picture encoded by the first video encoding method based on the dimension of a feature vector.
3. The video encoding apparatus according to claim 1 or 2, wherein the constraint means constrains the shape of slice division of a picture encoded by the first video encoding method.
4. A video decoding apparatus that performs video decoding by a first video decoding method based on a neural network and video decoding by a second video decoding method based on predictive encoding in an encoding unit, the video decoding apparatus comprising bitstream decoding means for decoding a bitstream generated by imposing a constraint on the encoding process of a picture encoded by the first video encoding method.
5. The video decoding apparatus according to claim 4, wherein the constraint is a constraint regarding the maximum code amount of a picture encoded by the first video encoding method, which is constrained based on the dimension of a feature vector.
6. The video decoding apparatus according to claim 4 or 5, wherein the constraint is a constraint regarding the shape of slice division of a picture encoded by the first video encoding method.
7. A video encoding method that performs video encoding by a first video encoding method based on a neural network and video encoding by a second video encoding method based on predictive encoding in an encoding unit, the video encoding method comprising imposing a constraint on the encoding process of a picture encoded by the first video encoding method.
8. A video decoding method that performs video decoding by a first video decoding method based on a neural network and video decoding by a second video decoding method based on predictive encoding in an encoding unit, the video decoding method comprising decoding a bitstream generated by imposing a constraint on the encoding process of a picture encoded by the first video encoding method.
9. A video encoding program that causes a computer to perform video encoding by a first video encoding method based on a neural network and video encoding by a second video encoding method based on predictive encoding in an encoding unit, the video encoding program causing the computer to execute a process for imposing a restriction on the encoding process of a picture encoded by the first video encoding method.
10. A video decoding program that causes a computer to perform video decoding by a first video decoding method based on a neural network and video decoding by a second video decoding method based on predictive encoding in an encoding unit, the video decoding program causing the computer to execute a process for decoding a bitstream generated by imposing a restriction on the encoding process of a picture encoded by the first video encoding method.
Citation Information
Patent Citations
Case and keyboard instrument
JP2024000581A
Distillation method
JP2024035448A
Image coding method, image decoding method, image coding device and image decoding device
WO2016199330A1