Video compression for both machine and human use using hybrid framework
By introducing the basic layer based on neural networks and traditional enhancement layer into video compression technology, the problem that the existing technology is difficult to optimize the use of human and machine videos at the same time is solved, and an efficient and flexible video compression solution is achieved.
Patent Information
- Application Number
- CN202380074570.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-08-14
- Publication Date
- 2025-06-03
AI Technical Summary
Existing video compression technologies are difficult to optimize video usage requirements for both humans and machines, especially with challenges in compression efficiency and computational complexity.
Video encoding and decoding is adopted based on neural networks, and the hybrid framework of basic layer and enhancement layer is optimized for machine tasks and human use respectively. The base layer uses NN-based codecs to compress machine tasks, while the enhancement layer uses traditional scalable video compression methods for video reconstruction used by humans.
It realizes the flexibility and adaptability of video compression without increasing the complexity of the decoder, and can meet the needs of machine tasks and human viewing at the same time, improving the efficiency and quality of video compression.
Smart Images

Figure CN120092451A_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to a method and apparatus for compressing images and videos for both human and machine use. Background Art
[0002] Traditional video compression standards can achieve low bitrates by transforming and degrading videos based on signal fidelity or visual quality. However, an increasing number of videos are now also "viewed" and analyzed by machines rather than humans, which typically involves neural network-based algorithms.
[0003] Optimizing existing video encoders directly for machine use is not easy because they are all handcrafted coding tools. The performance of neural network (NN)-based computer vision algorithms can be affected by artifacts generated by classical codecs, such as ringing, blocking artifacts, and loss of high spatial frequencies, because these artifacts are considered acceptable to the human visual system.
[0004] A new ad-hoc group at ISO / MPEG, called Video Coding for Machines (VCM), is working on standardizing an efficient way to transmit / store compressed bitstreams that contain information necessary to perform multiple tasks at the receiver, such as segmentation, object tracking, and reconstructing videos for human viewing. Meanwhile, JPEG is standardizing JPEG-AI, which is expected to become an end-to-end neural network-based image compression method that can also be optimized for machine tasks. Other standards and systems may be envisioned in the future as usage has become ubiquitous, such as video surveillance, autonomous vehicles, smart cities, etc. Summary of the Invention
[0005] According to one embodiment, a method for video encoding is provided, including: encoding an image using a neural network-based method to generate a first output; obtaining a first reconstructed version of the image corresponding to the first output; predicting blocks of the image based on the first reconstructed version of the image to form prediction blocks; and encoding the blocks based on the prediction blocks to generate a second output.
[0006] According to another embodiment, a method for video decoding is provided, including: entropy decoding a latent tensor corresponding to an image; obtaining a first reconstructed version of the image based on the latent tensor using a neural network-based method; predicting blocks of the image based on the first reconstructed version of the image to form prediction blocks; and decoding the blocks based on the prediction blocks to obtain a second reconstructed version of the image.
[0007] According to another embodiment, there is provided an apparatus for video encoding, including one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: obtain a first reconstructed version of the image corresponding to the first output; predict blocks of the image based on the first reconstructed version of the image to form prediction blocks; and encode the blocks based on the prediction blocks to generate a second output.
[0008] According to another embodiment, there is provided an apparatus for video decoding, including one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: perform entropy decoding on a latent tensor corresponding to an image; use a neural network-based method to obtain a first reconstructed version of the image based on the latent tensor; predict blocks of the image based on the first reconstructed version of the image to form prediction blocks; and decode the blocks based on the prediction blocks to obtain a second reconstructed version of the image.
[0009] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for video encoding or decoding according to the methods described herein.
[0010] One or more embodiments also provide a computer-readable storage medium having stored thereon video data generated according to the above method. One or more embodiments also provide a method and apparatus for transmitting or receiving video data generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A block diagram of a system in which aspects of the present embodiment can be implemented is illustrated.
[0012] Figure 2 An extensible encoder using a hybrid NN-based base layer and a traditional prediction-based enhancement layer according to an embodiment is illustrated.
[0013] Figure 3 A decoder supporting a base layer optimized for machine tasks and an enhancement layer for reconstructing an input for viewing according to an embodiment is illustrated.
[0014] Figure 4 An extensible video codec having a machine task NN-based base layer and a traditional enhancement layer using a hybrid scheme according to an embodiment is illustrated.
[0015] Figure 5 Illustrates the interaction between a proposed base layer codec and a standard scalable codec for a machine according to an embodiment.
[0016] Figure 6 Illustrates a scalable encoder and decoder according to an embodiment, where the base layer decoder reconstructs the input.
[0017] Figure 7 Illustrates a scalable encoder and decoder according to an embodiment, including an upscaling stage enabling resolution scalability.
[0018] Figure 8 Illustrates a scalable codec with a resolution scalability factor of 2 according to an embodiment. Detailed Description
[0019] Figure 1 Illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 can be implemented as a device including various components described below and configured to perform one or more of the aspects described in the present application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be implemented individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or to other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in the present application.
[0020] System 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing aspects such as those described in the present application. The processor 110 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., volatile memory devices and / or non-volatile memory devices). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.
[0021] System 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of System 100 or may be incorporated into the processor 110 as a combination of hardware and software known to those skilled in the art.
[0022] The program code to be loaded onto the processor 110 or the encoder / decoder 130 to execute the aspects described in the present application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the individual entries during the execution of the processes described in the present application. Such stored entries may include but are not limited to input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.
[0023] In several embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory (such as RAM) is used as the working memory for video encoding and decoding operations, such as for MPEG-2, MPEG-4, HEVC or VVC.
[0024] Input to the elements of the system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to: (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) composite input terminals; (iii) USB input terminals; and / or (iv) HDMI input terminals.
[0025] In various embodiments, the input device of block 105 has corresponding input processing elements associated therewith as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal band to one band), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down converter, a demodulator, an error corrector, and a de-multiplexer. The RF section may include a tuner that performs various functions among these functions, including, for example, down converting a received signal to a lower frequency (e.g., an intermediate frequency or a near baseband frequency) or down converting to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the (and other) elements described above, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0026] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 110 as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 110 as needed. The streams of demodulation, error correction, and de-multiplexing are provided to various processing elements, including, for example, the processor 110 and the encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0027] The various elements of the system 100 may be provided within an integrated housing. Within this integrated housing, the various elements may be interconnected using a suitable connection arrangement 115 (e.g., an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board) and data may be transmitted therebetween.
[0028] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver that is configured to transmit and receive data over the communication channel 190. The communication interface 150 can include, but is not limited to, a modem or a network card, and the communication channel 190 can be implemented, for example, within a wired and / or wireless medium.
[0029] In various embodiments, a Wi-Fi network (such as IEEE 802.11) is used to stream data to system 100. The Wi-Fi signals of these embodiments are received via the communication channel 190 and the communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or a router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streamed data to system 100, and the set-top box delivers the data through the HDMI connection of the input block 105. Yet some other embodiments use the RF connection of the input block 105 to provide streamed data to system 100.
[0030] System 100 can provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. In various examples of the embodiments, the other peripheral devices 185 include one or more of a standalone DVR, a disk player, a stereo system, a lighting system, and other devices based on the output-providing functions of system 100. In various embodiments, signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention) is used to transmit control signals between system 100 and the display 165, the speakers 175, or the other peripheral devices 185. The output devices can be communicatively coupled to system 100 via dedicated connections through the respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 using the communication channel 190 via the communication interface 150. In an electronic device (such as a television, for example), the display 165 and the speakers 175 can be integrated with other components of system 100 in a single unit. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0031] For example, if the RF portion of the input 105 is part of a separate set-top box, the display 165 and the speaker 175 can alternatively be separated from one or more of the other components. In various embodiments where the display 165 and the speaker 175 are external components, output signals can be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.
[0032] In this application, we aim to optimize the compression of bitstreams, including features dedicated to machine use and data for reconstructing image or video frames for human viewing. For machine use only, an NN-based codec can be used to extract and compress features for remote analysis. The advantage of using an NN-based method is that it is possible to train and optimize the system end-to-end and control the compression trade-off relative to the loss of task-based accuracy. Then, the reconstructed features at the decoder side can be used as inputs for computer vision tasks. However, end-to-end compression methods are still challenging for efficient video compression designed for human use. Traditional methods are still more efficient in terms of both compression efficiency and complexity (i.e., memory usage, energy, number of operations, etc.).
[0033] In addition, machine task algorithms may not need to be executed on every frame of the video. For example, object detection or segmentation can be performed every 4 frames to save energy. However, reconstructing the video typically requires processing all or at least a subset of the frames to maintain a satisfactory viewing frame rate, which is less flexible for computational adaptation and energy saving.
[0034] Some methods have been published where a scalable framework produces a bitstream that optimizes different layers for machine tasks or human use. However, they use an NN-based method for each layer, including those for human use, and thus maintain a very high computational complexity at both the encoder and decoder sides.
[0035] In this application, it is proposed to design a hybrid framework with multiple types of outputs to support machine tasks and human use. Such a framework is also referred to as a scalable framework, where the base layer uses an NN-based method to compress content for computer vision machine tasks, and at least one enhancement layer uses traditional scalable video compression methods to compress content for human viewing. We call such a framework a scalable framework because the reconstructed image from the base layer can be used as a predictor for the enhancement layer. However, it should be noted that, unlike traditional temporal, spatial, or quality scalable decoders that target either the base layer output or the enhancement layer output, the decoder according to the proposed hybrid framework can output both the base layer information and the enhancement information simultaneously. It should also be noted that sometimes we refer to the compression of images; however, the method applies to both images and videos.
[0036] Furthermore, like other scalable video compression methods, the proposed scheme is not limited to only one layer for machine tasks and only one layer for human use. It can include several layers optimized for different tasks and several layers for human use with different quality levels, spatial resolutions, temporal resolutions, etc.
[0037] In this context, the compression performance of the scalable framework is evaluated according to bitrate versus reconstructed pixel distortion or bitrate versus machine task accuracy, compared to the case when the input is compressed separately for the target task in each layer (also known as multicast). This can be calculated by measuring the total size of the bitstream containing all the layers in the scalable framework and evaluating the accuracy of the tasks performed for machine vision and the quality metrics regarding the reconstructed image / video for human viewing.
[0038] Figure 2 An example of the encoder of the proposed scheme is shown. In this example, the base layer uses an NN-based codec to perform compression for the machine, and the enhancement layer contains most of the basic operations of a traditional image codec (such as JPEG), but it can use the base layer as a prediction value.
[0039] Specifically, the base layer takes the input image X as input. The compression method for the base layer can consist of the analysis stage of compression analysis g a ((210), which usually consists of NN layers (such as 2D convolution and non-linear activation). This analysis produces the computed latent elements (Y), usually in the form of a 3D tensor consisting of N latent channels, which has a lower spatial dimension than the input image because the convolution used in the encoder usually involves downscaling. If the latent elements Y are optimized for the accuracy of visual tasks, the latent elements Y can be depth feature maps. N is usually much larger than 3 (the 3 channels RGB of the input image), such as 128, 192, 256, 320, etc. The tensor Y is then quantized (220) into and entropy encoded (230) to produce the bitstream corresponding to the base layer, which can be optimized in terms of bitrate and machine task accuracy. Note that the proposed method is not limited to the basic blocks of entropy coding and can include any advanced entropy modeling / transformation.
[0040] The mechanism of the enhancement layer in traditional scalable codecs is based on prediction and residuals, similar to temporal prediction in the context of video coding. The image generated by synthesizing g s (240) from the decoded base layer is used as the prediction value. The synthesis g s usually consists of NN layers (such as transposed 2D convolution and non-linear activation). The source image (X) and the prediction value The difference (250) (referred to as the residual) is then transformed, quantized (260), and entropy encoded (270) to produce a bitstream that enables the reconstruction of the original image. The bitstream is then multiplexed (280) to form a scalable bitstream.
[0041] To reconstruct a video frame, the synthesis module g is used s to transform the quantized latent tensor to produce an image that can be used as a prediction value for the enhancement layer. However, compared to existing methods that aim to simultaneously optimize complex deep autoencoders for both feature encoding and image reconstruction, it is proposed to reuse the compressed information from the base layer to synthesize a frame that can be used as a prediction value for encoding the enhancement layer, for example, using conventional predictive coding of existing or future scalable video compression frameworks for human viewing.
[0042] Figure 3 FIG. shows a basic decoder architecture according to an embodiment. The input bitstream is demultiplexed (310) into bitstream 0 for the base layer and bitstream 1 for the enhancement layer. The first part of the bitstream (bitstream 0) is entropy decoded (320) to produce a reconstructed latent tensor In the case where the task algorithm expects the features to have a different shape from the compressed latent tensor the tensor can optionally be processed by the feature synthesis f s (370) to produce a feature map for performing machine task inference f s can be trained for optimal transformation for the task machine, while g s is trained to produce good prediction values for the enhancement layer. For simplicity, hereinafter, when illustrating the codec, we provide an example in which is directly used as an input by the machine task.
[0043] The second part of the bitstream (bitstream 1) is first entropy decoded (330), then inverse quantized and inverse transformed (340) to reconstruct the residual, which is added (350) to the image generated by the synthesis g using the reconstructed latent tensor as input s (360) Hereinafter, when illustrating the codec, the multiplexing and demultiplexing modules and other communication processes between the encoder and decoder are skipped.
[0044] In this framework, the base layer can typically be end-to-end trained relying on a variational autoencoder, or can be approximated by a differentiable function for training. The machine task algorithm also relies on a differentiable neural network which enables the system to be jointly optimized, so as to update the parameters of the autoencoder and the task algorithm by backpropagating gradients from a loss function that depends on the task accuracy and the size of the compressed bitstream.
[0045] For the enhancement layer, g s is aimed at constructing frames that are good prediction values to be used by the prediction system in a so-called traditional codec. g s The reconstruction can be trained using a loss criterion on it (such as MSE (mean squared error)), but an l 0 norm can also be used because it is commonly used in traditional video compression to generate efficient prediction values, since the goal is not to create a high-fidelity image relative to the source, but to create prediction values that are added to the transmitted residuals, resulting in an efficient bitrate-distortion trade-off. By freezing the parameters of the encoder of the base layer, the training of g s can be carried out separately from the base layer so as not to affect the performance of compression on the task accuracy, or jointly with the base layer, i.e., by using a composite loss function that considers a weighted combination of the task accuracy, the reconstructed prediction values, and the size of the base layer bitstream.
[0046] This framework has the advantages of modularity and flexibility in terms of decoder complexity and usage. The shape of the compressed latent tensor produced by the NN-based autoencoder typically corresponds to a 3D tensor of size where C corresponds to the number of channels, H and W correspond to the height and width of the input image respectively, and n corresponds to the number of factor-2 downsampling operations (such as strided convolution, pooling) in the encoder. Note that the base layer can also use block-based end-to-end compression. In both cases, the reconstructed image can be used as a reference frame for predicting the enhancement layer.
[0047] Frames of lower resolution can be considered because the channels of the latent tensor typically have a lower resolution, thus reducing some of the complex transposed 2D convolutions in the decoder (cut) g s The lighter resolution scalability feature of the scalable codec can handle upsampling and add appropriate residuals.
[0048] In addition, for machine tasks, depth features may not be required at every moment of the video sequence, while for smooth display to the human eye, all frames may be needed. Instead of using computationally expensive synthesis operations, intermediate frames for which no information from the base layer is available can be directly predicted in time. Temporal scalability already exists in traditional scalable codecs. Frames of the enhancement layer can be predicted using temporal reference frames or frames from available lower layers. When the latter do not exist, only temporal prediction is applied.
[0049] This application aims to provide methods that enable the use of compressed data from the base layer for video reconstruction using traditional scalable video compression methods through at least one enhancement layer.
[0050] Compressing the input content into bitstreams optimized for different machine tasks or for human use has been widely studied. However, they always use differential autoencoders for each task (including human viewing purposes) with the goal of optimal transmission of compressed data. In this application, it is proposed to use a codec optimized for machine tasks in combination with a hybrid method that will add the necessary residual information to reconstruct the video for human use. To some extent, the extracted features encoded for machine tasks will be used as the base layer, i.e., the predicted values, in a scalable video codec such as MPEG-4 / SVC, its successor SHVC based on H.265 / HEVC, or any other / future prediction method.
[0051] In order to reference data in the base layer in the context of predictive coding, it may be necessary for the decoder to use the compressed data in the base layer to reconstruct pixels so that the enhancement layer of the scalable codec can reference the reconstructed pixels as a reference. Different cases of reconstruction are described in detail below.
[0052] First, two usage cases are described:
[0053] - Case a: A new scalable compression system that includes a base layer for compressing features optimized for machine tasks and an enhancement layer that relies on traditional predictive coding for viewing.
[0054] - Case b: A combination of the proposed encoder for the base layer and an existing scalable video coding standard for processing the enhancement layer(s).
[0055] Below, we consider an example with one (base) layer for machine tasks and one enhancement layer for viewing, but the proposed method is not limited to having only one layer per case. The system can be extended to multiple enhancement layers for different machine tasks and / or multiple layers for viewing, such as different resolutions, quality levels, temporal resolutions, etc.
[0056] Case a
[0057] In this case, it is proposed that there be a base layer for the machine task that includes NN-based analysis. Similar to H.265 / SHVC, where the base layer can be decoded by any HEVC decoder, i.e., without including scalable extensions, here it is possible to independently process the base layer content for a "video coding for machines" (VCM) encoder / decoder. Reconstructing reference frames from the base layer to predict the enhancement layer will occur at the enhancement layer level, as Figure 4 depicted therein.
[0058] To better understand the context of scalable video compression, Figure 4 a scalable video codec according to an embodiment is shown. Compared to the previous figure, the base layer remains unchanged. However, the enhancement layer now details different prediction options, including the selection of intra prediction modes (420, 425) or inter prediction modes (410, 415). For inter prediction, the prediction can be temporal prediction or layerwise prediction, where temporal prediction involves motion estimation with reference to temporal pictures from the decoded picture buffer (DPB, 430, 435), and layerwise prediction uses the generated prediction values synthesized from the base layer without encoding the motion information. Then, the encoder selects the best prediction value, e.g., using a rate-distortion optimization process. Note that the encoder now includes a decoding module. As Figure 4 illustrated, both the encoder and decoder sides perform inverse transformation and quantization (450, 455), as well as loop filtering (LF, 440, 445), in order to be able to reconstruct (460, 465) and store reference images for temporal prediction. At the decoder, syntax elements are parsed from the bitstream, informing which prediction modes to use and which reference pictures, temporal layers, or base layers to use.
[0059] For case a, when we modify the process for the generated enhancement layer that now includes g s the new system requires syntax for both the base layer process and the enhancement layer process.
[0060] Case b
[0061] Compared to case a, the synthesis stage g s for generating reference frames from the base layer occurs within the proposed codec at the base layer, such that the generated frames can be directly used by existing scalable codecs. In other words, we propose an extension of a new codec or VCM codec that can be coupled with existing traditional scalable video compression systems. Figure 5 shows a scenario where g s is now part of the base layer. The enhancement layer is used as is, thus obtaining the reference pictures generated by the base layer.
[0062] Here, we describe how the proposed method can interact with the existing syntax of traditional scalable codecs. We take the scalability features of H.265 / SHVC described in Appendices F and H of the H.265 / HEVC specification as an example. Appendix F specifies the high-level syntax related to multi-layer bitstreams, namely multi-view, scalability, 3D. Appendix H specifies the process of generating reference pictures for compressing the current enhancement layer based on previously decoded layers, i.e., the layers with lower id (nuh_layer_id).
[0063] In Section 7.4.3.1, the video parameter set syntax elements vps_base_layer_internal_flag and vps_base_layer_available_flag are defined as follows:
[0064] - If vps_base_layer_internal_flag is equal to 1 and vps_base_layer_available_flag is equal to 1, then the base layer exists in the bitstream.
[0065] – Otherwise, if vps_base_layer_internal_flag is equal to 0 and vps_base_layer_available_flag is equal to 1, then the base layer is provided by an external means not specified in this specification.
[0066] – Otherwise, if vps_base_layer_internal_flag is equal to 1 and vps_base_layer_available_flag is equal to 0, then the base layer is not available (neither exists in the bitstream nor is provided by an external means), but the VPS includes information about the base layer as if it existed in the bitstream.
[0067] – Otherwise (vps_base_layer_internal_flag is equal to 0 and vps_base_layer_available_flag is equal to 0), then the base layer is not available (neither exists in the bitstream nor is provided by an external means), but the VPS includes information about the base layer as if it were provided by an external means not specified in this specification.
[0068] In case b, the encoder can, for example, set vps_base_layer_internal_flag = 0 because our base layer does not use HEVC for encoding, and set vps_base_layer_available_flag = 1 because it uses an external means for encoding.
[0069] Then, Section H.8.1.4 specifies the process for exporting inter-layer reference pictures, regardless of whether the inter-layer frames need to be processed (upsampled, color transformed, etc.) before being stored as references.
[0070] Here, we use the H.265 / HEVC syntax as an example. However, the proposed method applies to any traditional multi-layer / scalable codec that relies on predictive coding using reference frames from already reconstructed layers. Additionally, although the codec is described based on VCM, the proposed method is not limited to VCM and applies to other video codec platforms.
[0071] Reconstruction using quality / SNR scalability
[0072] In this embodiment, the decoder includes a synthesis stage that can reconstruct frames from the reconstructed latent tensors (features) at the desired resolution of the output video. In cases corresponding to Figure 2 and Figure 3 , the reconstructed frames from the synthesis stage can be directly used as prediction values for the enhancement layer. This corresponds to what is called SNR (signal-to-noise ratio) scalability in traditional scalable codecs, i.e., when the base layer is encoded at a lower quality.
[0073] Some applications may require the compression system of the base layer to reconstruct images rather than generate feature maps for machine vision networks. Machine task designers who train their algorithms on datasets of videos or images may need a compression system that allows them to extract features on the receiver side. In that case, the base layer also includes a g s that can be trained and optimized together with the encoder to reconstruct images optimized for machine tasks. In that case, the SNR scalable system also enables the reconstruction of images for viewing ( Figure 6 "rec" on s ), because the reconstructed base layer may contain artifacts that are severe for the human visual system but acceptable for machine task accuracy. For example, for image classification, low spatial frequencies or local features may be sufficient but not suitable for viewing. The image reconstructed by the base layer decoder can be directly used as a reference for inter-layer prediction in the enhancement layer. Note that the encoder of the enhancement layer requires a g
[0074] Reconstruction using resolution scalability
[0075] In this embodiment, it is preferable to reconstruct frames from the synthesis module at a lower resolution and use the resolution scalability module from a traditional scalable codec. This option can be useful when the decoder device is limited in terms of deep learning capabilities (such as a graphics unit for convolutional layers, for example), but includes a hardware implementation of a traditional scalable codec with upsampling / downsampling separable filters on it.
[0076] The shape of the compressed latent tensor produced by the NN-based autoencoder typically corresponds to a 3D tensor of size where C corresponds to the number of channels, H and W correspond to the height and width of the input image respectively, and n corresponds to the number of downsampling operations (such as strided convolution, pooling) with a factor of 2 included in the encoder.
[0077] Figure 7 A codec according to an embodiment is described, where in addition to the Figure 3 modules described in, the upsampling process (710, 720) takes as input a synthetic image with a resolution lower than X. Note that other processes can be added or replace upsampling, for example, bit-depth increase, color format conversion, resampling, etc.
[0078] To provide a clearer example of the proposed resolution scalability, Figure 8 shows the case where the encoder contains an analysis stage g a which analysis stage g a includes 3 convolutions with a stride of 2 (810). The resulting latent tensor will have dimensions The codec can use a synthesis g s that includes only two transposed convolutions (820) to synthesize an image, and the last output dimension of the tensor, that is, an image with a lower dimension that can finally be upsampled (830) using specific tools from a scalable codec for the compression enhancement layer.
[0079] Note that for training g s , we now use the downsampled version of the input image with a standard (such as MSE) and use the same filter as we use in the upsampling process.
[0080] This extension is also compatible with case b, where synthesis occurs in the compressed base layer and for a codec that produces a reference frame for a scalable codec. In this case, the processing (upsampling, for example, bit-depth increase, color format conversion, resampling, etc.) can still occur at the enhancement layer, that is, within the traditional scalable codec.
[0081] Discarding frames at the base layer
[0082] Many computer vision tasks do not need to be performed at every frame of a video sequence. Especially when feature extraction can be computationally intensive for low-end encoders, it is often possible to compute the necessary features every n frames. a Feature extraction is performed for the target frame, which can still be used as a base layer in the proposed model. Other frames are encoded in an enhancement layer without referring to the base layer as a prediction value. In terms of high-level syntax, no base layer prediction value is added to the reference picture list, such that for the current frame, only temporal pictures can be used for motion compensation.
[0083] Feature extraction is performed for the target frame, which can still be used as a base layer in the proposed model. Other frames are encoded in an enhancement layer without referring to the base layer as a prediction value. In terms of high-level syntax, no base layer prediction value is added to the reference picture list, such that for the current frame, only temporal pictures can be used for motion compensation.
[0084] In this embodiment, temporal scalability can be envisioned if frames corresponding to a regular period are discarded (not analyzed), and the base layer will correspond to a lower frame rate. Another syntax can also be used for irregular temporal re-partitioning of the frames in the base layer, which will indicate inter-layer prediction only when frames from the base layer are available.
[0085] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0086] Various methods are described herein, and each of the methods includes one or more steps or actions for implementing the method. Unless the correct operating method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as for example "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and can occur, for example, before, during, or overlapping with the second decoding.
[0087] Various implementations relate to decoding. As used in this application, "decoding" can cover, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as, for example, entropy decoding, inverse quantization, and inverse transform. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context described and is considered to be well understood by those skilled in the art.
[0088] The various implementations involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can cover, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream.
[0089] The implementations and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, an apparatus such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the transfer of information between end users.
[0090] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations thereof that occur throughout this application do not necessarily all refer to the same embodiment.
[0091] Additionally, this application can be related to "determining" various pieces of information. Determining information can include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.
[0092] Furthermore, this application can be related to "accessing" various pieces of information. Accessing information can include, for example, one or more of the following: receiving information, retrieving information (e.g., retrieving information from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0093] Additionally, this application can be related to "receiving" various pieces of information. Similar to "accessing", receiving is intended to be a broad term. Receiving information can include, for example, one or more of the following: accessing information or retrieving information (e.g., retrieving information from a memory). Furthermore, "receiving" is generally involved in one way or another during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0094] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely, " / ", "and / or", and "at least one of...", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of the first-listed option and the second-listed option (A and B), or the selection of the first-listed option and the third-listed option (A and C), or the selection of the second-listed option and the third-listed option (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended to as many listed items as possible.
[0095] As will be apparent to those of ordinary skill in the art, the implementation can generate a variety of signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal can be formatted to carry the bitstream of the described embodiment. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
Claims
1. A method for video encoding, comprising: encoding an image using a neural network-based method to generate a first output; obtaining a first reconstructed version of the image corresponding to the first output; predicting a block of the image based on the first reconstructed version of the image to form a prediction block; and encoding the block based on the prediction block to generate a second output.
2. The method according to claim 1, wherein encoding to generate the first output comprises: obtaining a latent tensor corresponding to the image based on the neural network-based method; quantizing the latent tensor; and performing entropy encoding on the quantized latent tensor to generate the first output.
3. The method according to claim 2, wherein obtaining the first reconstructed version of the image comprises performing a second neural network-based method on the quantized latent tensor.
4. The method according to claim 3, further comprising: upsampling the output from the second neural network-based method to obtain the first reconstructed version of the image.
5. The method according to any one of claims 1-4, wherein encoding to generate the second output comprises: obtaining the difference between the image and the first reconstructed version of the image; transforming the difference and quantizing the transformed difference; and performing entropy encoding on the quantized transformed difference to generate the second output.
6. The method according to any one of claims 1-5, wherein encoding to generate the second output comprises: selecting a prediction mode for the block from intra prediction, temporal prediction, and inter-layer prediction, wherein in response to selecting inter-layer prediction, performing the prediction on the block based on the first reconstructed version of the image.
7. The method according to any one of claims 1-6, wherein performing the encoding of the block according to a video coding standard or recommendation, wherein the second output is generated as an enhancement layer in scalable video coding.
8. The method according to claim 7, wherein the first output and the first reconstructed version of the image are generated by a means other than the video coding standard or recommendation.
9. The method according to claim 8, wherein a syntax element indicates that the first output is generated by a means other than the video coding standard or recommendation.
10. The method according to any one of claims 1-9, wherein the image is from a video sequence, and wherein the encoding for generating the first output is performed on a subset of the first images of the video sequence, and the encoding for generating the second output is performed on a subset of the second images of the video sequence, wherein the subset of the second images includes more images than the subset of the first images.
11. The method according to claim 10, wherein a syntax element is used to indicate that inter-layer prediction is included only when the image is available at the base layer.
12. A method for video decoding, comprising: performing entropy decoding on a latent tensor corresponding to an image; Using a neural network-based method, obtain a first reconstructed version of the image based on the latent tensor; Predict blocks of the image based on the first reconstructed version of the image to form predicted blocks; And Decode the blocks based on the predicted blocks to obtain a second reconstructed version of the image.
13. The method according to claim 12, Wherein, Obtaining the first reconstructed version of the image includes performing a second neural network method based on the latent tensor.
14. The method according to claim 13, further comprising upsampling the output from the second neural network method to obtain the first reconstructed version of the image.
15. The method according to any one of claims 12-14, Wherein, Decoding to generate the second reconstructed version of the image includes: Performing entropy decoding to obtain the quantized transform difference of the blocks; and Performing dequantization and inverse transform to decode the prediction residuals of the blocks, wherein the second reconstructed version of the image is obtained based on the first version of the reconstructed image and the decoded prediction residuals.
16. The method according to any one of claims 12-15, further Comprising: Determine a prediction mode for the blocks from intra-frame prediction, temporal prediction, and inter-layer prediction, wherein in response to determining inter-layer prediction, perform the prediction on the blocks based on the first reconstructed version of the image.
17. The method according to any one of claims 12-16, Wherein, Perform the decoding for obtaining the second reconstructed version of the image according to a video coding standard or recommendation, wherein the second reconstructed version of the image is generated as an enhancement layer in scalable video coding.
18. The method according to claim 17, Wherein, The first reconstructed version of the image is generated by a means other than the video coding standard or recommendation.
19. The method according to claim 18, wherein a syntax element indicates that the first reconstructed version of the image is generated by a means other than the video coding standard or recommendation.
20. The method according to any one of claims 12-19, wherein the image is from a video sequence, and wherein the latent tensor is decoded for a first subset of images of the video sequence, and the second reconstructed version of the image is performed for a second subset of images of the video sequence, wherein the second subset of images includes more images than the first subset of images.
21. The method according to claim 20, wherein a syntax element indicates that inter-layer prediction is included only when the image is available at the base layer.
22. An apparatus comprising one or more processors, and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method of any one of claims 1-21.
23. A signal comprising video data, which is formed by performing the method of any one of claims 1-11.
24. A computer-readable storage medium having stored thereon instructions for encoding or decoding a video according to the method of any one of claims 1-21.