Deep neural network-based image compression using latent shift based on the gradient of latent entropy
By utilizing deep neural networks to exploit gradients of latent entropy, the method improves video compression efficiency by inferring and encoding decoder-side information, optimizing quantization and entropy modeling for enhanced encoding and decoding processes.
Patent Information
- Application Number
- JP2025514384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-20
- Filing Date
- 2023-09-14
- Publication Date
- 2025-10-01
AI Technical Summary
Conventional video compression methods, including both traditional and fully neural network-based approaches, fail to effectively utilize the information available at the decoder side to enhance compression efficiency, limiting the potential improvements in encoding and decoding processes.
The method employs deep neural networks to utilize gradients of latent entropy for encoding and decoding, involving deep hyper-prior neural networks and factorized entropy models to infer and encode side and main latents, optimizing the compression process through quantization and arithmetic encoding based on probability mass functions and step sizes derived from entropy gradients.
This approach enhances compression efficiency by leveraging decoder-side information, improving the encoding and decoding processes through optimized quantization and entropy modeling, leading to more efficient video data representation and transmission.
Smart Images

Figure 2025532521000001_ABST
Abstract
Description
[Technical Field]
[0001] At least one of the present embodiments relates generally to methods and devices for encoding and decoding picture data using deep neural networks, and more particularly to methods and devices that exploit gradients of latent entropy to improve compression efficiency. [Background technology]
[0002] To achieve high compression efficiency, conventional video compression schemes typically employ prediction and transformation to exploit spatial and temporal redundancy in video content. During encoding, a picture of video content is divided into blocks of samples (i.e., pixels), which are then divided into one or more sub-blocks, hereinafter referred to as original sub-blocks. Intra- or inter-prediction is then applied to each sub-block to exploit intra- or inter-image correlation. Regardless of the prediction method (intra- or inter-) used, a predictor sub-block is determined for each original sub-block. Sub-blocks representing the difference between the original and predictor sub-blocks, often referred to as prediction error sub-blocks, prediction residual sub-blocks, or simply residual sub-blocks, are then transformed, quantized, and entropy coded to generate an encoded video stream. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to the transform, quantization, and entropy coding.
[0003] Recently explored video coding solutions have investigated neural network (NN)-based processing. Several NN-based solutions have been proposed, ranging from hybrid solutions where NNs are used to implement some of the tools of the traditional video compression schemes mentioned above, to fully NN-based compression solutions.
[0004] An important aspect of any video (or image) compression solution is to use as much information available at the decoder side as possible to improve compression efficiency. The information available at the decoder side naturally includes data explicitly coded in the bitstream (i.e., in the coded video data), but also includes data that is not explicitly coded but can be inferred from the explicitly coded video data. Such inferred data has been primarily used in traditional video compression schemes. However, fully NN-based compression methods are not as mature as methods based on traditional video compression schemes. Therefore, the possibility of utilizing explicitly coded data to infer other data is still an aspect to be explored.
[0005] It is desirable to determine what data can be inferred from explicitly coded data available at the decoder side in a fully NN-based video (or image) compression method, and how these inferred data can be used to improve the compression efficiency of the fully NN-based video compression method. Summary of the Invention
[0006] In a first aspect, one or more of the present embodiments provides a method, the method comprising: A method for encoding a video signal includes: obtaining an input picture; applying a deep neural network-based encoder to the input picture to obtain a main latent; applying a deep hyper prior neural network-based encoder to the input picture to obtain a side latent; quantizing the side latent using a first quantization method to obtain a side code; applying a factorized entropy model to the side code to determine a first probability mass function of the side code under the first quantization method; arithmetically encoding the side code based on the first probability mass function in the video data; dequantizing the side code to obtain a reconstructed side latent; determining a first step size based on a gradient of entropy of the side information relative to the reconstructed side latent; encoding information representing the first step size in the video data; shifting the reconstructed side latent using the first step size and the gradient of entropy of the side information relative to the reconstructed side latent; applying a deep hyper-ahead decoder to the reconstructed side latent to obtain information representing a probability distribution modeling the main latent; obtaining a second probability mass function of the main code using the probability distribution modeling the main latent; quantizing the main latent to obtain the main code; and arithmetically encoding the main code based on the second probability mass function in the video data.
[0007] In one embodiment, the method further includes dequantizing the main code to obtain a reconstructed main latent; determining a second step size based on a gradient of entropy of the main information with respect to the main latent; and encoding information representing the second step size in the video data.
[0008] In a second aspect, one or more of the present embodiments provides a method, the method comprising: obtaining an input picture, applying a deep neural network-based encoder to the input picture to obtain a main latent, applying a deep hyper prior neural network-based encoder to the input picture to obtain a side latent, quantizing the side latent using a first quantization method to obtain a side code, applying a factorized entropy model to the side code to determine a first probability mass function of the side code under the first quantization method, arithmetically encoding the side code based on the first probability mass function in the video data, and arithmetically encoding the side code to obtain a reconstructed side latent. to obtain a second probability mass function of the main code using the probability distribution modeling the main latent; arithmetically encoding the main code based on the second probability mass function in the video data; dequantizing the main code to obtain a reconstructed main latent; determining a second step size based on a gradient of entropy of the main information with respect to the main latent; and encoding information representing the second step size in the video data.
[0009] In a third aspect, one or more of the present embodiments provides a method, the method comprising: the first step size based on a gradient of entropy of the side information with respect to the reconstructed side latent; shifting the reconstructed side latent using the first step size and the gradient of entropy of the side information with respect to the reconstructed side latent; applying a deep hyper-ahead decoder to the reconstructed side latent to obtain information representing a probability distribution modeling the main latent; obtaining a second probability mass function of the main code using the probability distribution modeling the main latent; arithmetically decoding the main code from the video data using the second probability mass function;
[0010] In one embodiment, the method further includes decoding information representing a second step size based on a gradient of entropy of the main information with respect to the main latent from the video data, and shifting the reconstructed main latent using the second step size and the gradient of entropy of the main information with respect to the main latent, wherein a deep neural network-based decoder is applied to the shifted reconstructed main latent.
[0011] In a fourth aspect, one or more of the present embodiments provides a method, the method comprising: the main information is decoded from the video data using the second probability mass function; the main information is decoded using the second step size based on a gradient of entropy of the main information with respect to the main latent from the video data; the main information is decoded using the second step size based on a gradient of entropy of the main information with respect to the main latent; and the deep neural network-based decoder is applied to the shifted reconstructed main latent to obtain an output picture.
[0012] In a fifth aspect, one or more of the present embodiments provides a device comprising an electronic circuit, the electronic circuit comprising: Take the input picture, To obtain the main latent, a deep neural network-based encoder is applied to the input picture, To obtain the side latents, a deep hyper-prior neural network based encoder is applied to the input picture, quantizing the side latents using a first quantization method to obtain side codes; applying the factorized entropy model to the side code to determine a first probability mass function of the side code under a first quantization method; arithmetically encoding a side code based on a first probability mass function in the video data; Dequantize the side codes to obtain the reconstructed side potentials. determining a first step size based on a gradient of the entropy of the side information relative to the reconstructed side latent; encoding information representing a first step size in the video data; shifting the side latents using a first step size and a gradient of the entropy of the side information with respect to the reconstructed side latent; Apply a deep hyper-prior decoder to the reconstructed side latents to obtain information representing the probability distribution that models the main latent. Using the probability distribution that models the main latent, we obtain a second probability mass function for the main code, Quantize the main latent to obtain the main code, arithmetically encoding the main code based on a second probability mass function in the video data; It is configured for:
[0013] In one embodiment, the electronic circuitry comprises: Dequantize the main code to obtain the reconstructed main latent. determining a second step size based on the gradient of the entropy of the main information relative to the main latent; encoding information representing the second step size in the video data; It is further configured for:
[0014] In a sixth aspect, one or more of the present embodiments provides a device comprising an electronic circuit, the electronic circuit comprising: Take the input picture, To obtain the main latent, a deep neural network-based encoder is applied to the input picture, To obtain the side latents, a deep hyper-prior neural network based encoder is applied to the input picture, quantizing the side latents using a first quantization method to obtain side codes; applying the factorized entropy model to the side code to determine a first probability mass function of the side code under a first quantization method; arithmetically encoding a side code based on a first probability mass function in the video data; Dequantize the side codes to obtain the reconstructed side potentials. Apply a deep hyper-prior decoder to the reconstructed side latents to obtain information representing the probability distribution that models the main latent. Using the probability distribution that models the main latent, we obtain a second probability mass function for the main code, Arithmetically encoding a main code based on a second probability mass function in the video data. Dequantize the main code to obtain the reconstructed main latent. determining a second step size based on the gradient of the entropy of the main information relative to the main latent; encoding information representing the second step size in the video data; It is configured for:
[0015] In a seventh aspect, one or more of the present embodiments provides a device comprising an electronic circuit, the electronic circuit comprising: applying arithmetic decoding to the side information included in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model; Dequantize the side codes to obtain the reconstructed side potentials. decoding information representing a first step size based on a gradient of the entropy of the side information with respect to the reconstructed side latent; shifting the reconstructed side latent using a first step size and a gradient of the entropy of the side information with respect to the reconstructed side latent; Apply a deep hyper-prior decoder to the reconstructed side latents to obtain information representing the probability distribution that models the main latent. Using the probability distribution that models the main latent, we obtain a second probability mass function for the main code, arithmetically decoding a main code from the video data using a second probability mass function; Dequantize the main code to obtain the reconstructed main latent. Apply a deep neural network-based decoder to the reconstructed main latent to obtain the output picture. It is configured for:
[0016] In one embodiment, the electronic circuitry comprises: Decoding information representing a second step size based on a gradient of entropy of the main information relative to the main latent from the video data; Shifting the reconstructed main latent using a second step size and the gradient of the entropy of the main information with respect to the main latent. Further constructed, a deep neural network based decoder is applied to the shifted reconstructed main latent.
[0017] In an eighth aspect, one or more of the present embodiments provides a device comprising an electronic circuit, the electronic circuit comprising: applying arithmetic decoding to the side information included in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model; Dequantize the side codes to obtain the reconstructed side potentials. Apply a deep hyper-prior decoder to the reconstructed side latents to obtain information representing the probability distribution that models the main latent. Using the probability distribution that models the main latent, we obtain a second probability mass function for the main code, arithmetically decoding a main code from the video data using a second probability mass function; Dequantize the main code to obtain the reconstructed main latent. Decoding information representing a second step size based on a gradient of entropy of the main information relative to the main latent from the video data; Shifting the reconstructed main latent using a second step size and the gradient of the entropy of the main information with respect to the main latent; Apply a deep neural network-based decoder to the shifted reconstructed main latent to obtain the output picture. It is configured for:
[0018] In a ninth aspect, one or more of the present embodiments provide a computer program comprising program code instructions for implementing a method according to the first, second, third, fourth, fifth or sixth aspect.
[0019] In a tenth aspect, one or more of the present embodiments provide a non-transitory information storage medium storing program code instructions for implementing a method according to the first, second, third, fourth, fifth, or sixth aspect.
[0020] In an eleventh aspect, one or more of the present embodiments provides a signal generated by the method of the first, second or third aspect, or by the device of the seventh, eighth or ninth aspect. [Brief explanation of the drawings]
[0021] [Figure 1] 1 illustrates generally the context in which the embodiments may be implemented; [Figure 2A] 1 illustrates schematically an example of a hardware architecture of a processing module capable of implementing an encoding module or a decoding module in which various aspects and embodiments are implemented. [Figure 2B] 1 illustrates a block diagram of a first example system in which various aspects and embodiments may be implemented. [Figure 2C] 1 illustrates a block diagram of a second example system in which various aspects and embodiments may be implemented. [Figure 3]We illustrate the training phase of a deep NN-based video compression system. [Figure 4] We illustrate a video encoder based on a trained deep NN-based video compression system. [Figure 5] We illustrate a video decoder based on a trained deep NN-based video compression system. [Figure 6] 1 illustrates a video encoding process based on a trained deep NN-based video compression system of one embodiment. [Figure 7] 1 illustrates a video decoding process based on a trained deep NN-based video compression system of one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0022] FIG. 1 illustrates schematically an example of a context in which embodiments may be implemented.
[0023] 1, system 11, which may be a camera, a storage device, a computer, a server, or any device capable of delivering video data, transmits video data to system 13 using communication channel 12. The video data is encoded and transmitted by system 11, or received and / or stored by system 11 and then transmitted. Communication channel 12 may be a wired (e.g., Internet or Ethernet) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.
[0024] The system 13, which may be, for example, a set-top box, receives and decodes the video data to produce a sequence of decoded pictures.
[0025] The obtained sequence of decoded pictures is then transmitted using a communication channel 14, which may be a wired or wireless network, to a display system 15. The display system 15 then displays the pictures.
[0026] In an embodiment, system 13 is included in a display system 15. In that case, system 13 and display 15 are included in a TV, computer, tablet, smartphone, head-mounted display, etc.
[0027] 2A, 2B, and 2C illustrate example devices, apparatuses, and / or systems that may implement various embodiments.
[0028] Figure 2A illustrates schematically an example of a hardware architecture of a processing module 200 capable of implementing an encoding module or a decoding module capable of implementing the encoding method of Figure 6 and the decoding method of Figure 7, respectively. The encoding module is included in system 11, for example, if this system is responsible for encoding video data. The decoding module is included in system 13, for example.
[0029] The processing module 200 includes a processor or central processing unit (CPU) 2000, including, by way of non-limiting example, one or more microprocessors, general purpose computers, special purpose computers, and processors based on multi-core architectures, connected by a communication bus 2005; a random access memory (RAM) 2001; a read only memory (ROM) 2002; and a memory array (RAM) 2003, including, by way of example only, an electrically erasable programmable read-only memory (EEPROM), a read only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), flash, a magnetic disk drive, and / or an optical disk drive, or a secure digital (SD) card reader and / or a hard disk drive. The system includes a storage unit 2003, which may include non-volatile and / or volatile memory, including, but not limited to, a storage media reader and / or a network-accessible storage device such as a hard disk drive (HDD), and at least one communication interface 2004 for exchanging data with other modules, devices, or systems. The communication interface 2004 may include, but is not limited to, a transceiver configured to send and receive data over a communication channel. The communication interface 2004 may include, but is not limited to, a modem or a network card.
[0030] If processing module 200 implements a decoding module, communication interface 2004 may, for example, enable processing module 200 to receive encoded video data and provide a sequence of decoded pictures. If processing module 200 implements an encoding module, communication interface 2004 may, for example, enable processing module 200 to receive and encode a sequence of original pictures and provide encoded video data.
[0031] The processor 2000 can execute instructions loaded into the RAM 2001 from the ROM 2002, an external memory (not shown), a storage medium, or a communication network. When the processing module 200 is powered on, the processor 2000 can read instructions from the RAM 2001 and execute them. These instructions form a computer program that causes the processor 2000 to perform, for example, the decoding method described in relation to Figure 7 or the encoding method described in relation to Figure 6.
[0032] All or part of the algorithms and steps of the methods of Figures 6 and 7 may be implemented in software form by execution of an instruction set by a programmable machine such as a DSP (digital signal processor) or a microcontroller, or may be implemented in hardware form by a machine or dedicated component such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit).
[0033] As can be seen, microprocessors, general purpose computers, special purpose computers, processors based or not based on multi-core architectures, DSPs, microcontrollers, FPGAs, and ASICs are electronic circuits adapted to at least partially implement the methods of Figures 6 and 7.
[0034] FIG. 2C illustrates a block diagram of an example system 13 in which various aspects and embodiments can be implemented. System 13 can be embodied as a device including various components, described below, configured to perform one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected consumer electronics, and head-mounted displays. Elements of system 13, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, system 13 includes a processing module 200 that implements a decoding module. In various embodiments, system 13 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 13 is configured to implement one or more of the aspects described herein.
[0035] Input to processing module 200 may be provided via various input modules, as shown in block 231. Such input modules may include, but are not limited to, (i) a radio frequency (RF) module, for example, receiving RF signals transmitted over the air from a broadcast station, (ii) a component (COMP) input module (or set of COMP input modules), (iii) a Universal Serial Bus (USB) input module, and / or (iv) a High Definition Multimedia Interface (HDMI) input module. Another example, not shown in FIG. 2C, is composite video.
[0036] In various embodiments, the input modules of block 231 have associated respective input processing elements, as known in the art. For example, the RF module may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band, which in certain embodiments may be referred to (for example) as a channel, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF module of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF module and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, omit some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF module includes an antenna.
[0037] Additionally, the USB module and / or HDMI module may include respective interface processors for connecting system 13 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of the input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or, if desired, in processing module 200. Similarly, aspects of the USB or HDMI interface processing may be implemented, if desired, in a separate interface IC or in processing module 200. The demodulated, error corrected, and demultiplexed streams are provided to processing module 200.
[0038] The various elements of system 13 may be provided within a unitary housing, where the various elements may be interconnected and transmit data between them using any suitable connection arrangement, e.g., internal buses known in the art, including an inter-IC (I2C) bus, wiring, and printed circuit boards. For example, in system 13, processing module 200 is interconnected to the other elements of system 13 by bus 2005.
[0039] The communication interface 2004 of the processing module 200 enables the system 13 to communicate over the communication channel 12. As already mentioned above, the communication channel 12 can be implemented, for example, in a wired and / or wireless medium.
[0040] In various embodiments, data is streamed or otherwise provided to system 13 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via communication channel 12 and communication interface 2004 adapted for Wi-Fi communication. Communication channel 12 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the RF connection of input block 231 is used to provide streaming data to system 13. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as a cellular network or a Bluetooth network.
[0041] System 13 can provide output signals to various output devices, including display system 15, speakers 235, and other peripheral devices 236. Display system 15 in various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display system 15 can be for a television, a tablet, a laptop, a mobile phone, a head-mounted display, or other device. Display system 15 can also be integrated with other components (e.g., like a smartphone) or separate (e.g., an external monitor for a laptop). In various example embodiments, other peripheral devices 236 include one or more of a standalone digital video disc (or digital versatile disc) (both terms referred to as a DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 236 to provide functionality based on the output of system 13. For example, a disc player performs the function of playing the output of system 13.
[0042] In various embodiments, control signals are communicated between system 13 and display system 15, speakers 235, or other peripheral devices 236 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices with or without user intervention. Output devices can be communicatively coupled to system 13 via dedicated connections through respective interfaces 232, 233, and 234. Alternatively, output devices can connect to system 13 using communication channel 12 via communication interface 2004 or using a dedicated communication channel corresponding to communication channel 12 of FIG. 2C via communication interface 2004. Display system 15 and speakers 235 can be integrated into a single unit with other components of system 13 in an electronic device such as, for example, a television. In various embodiments, display interface 232 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0043] Display system 15 and speakers 235 may alternatively be separate from one or more of the other components. In various embodiments in which display system 15 and speakers 235 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0044] FIG. 2B illustrates a block diagram of an example system 11 in which various aspects and embodiments can be implemented. System 11 is very similar to system 13. System 11 can be embodied as a device including various components, described below, configured to perform one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, cameras, and servers. The elements of system 11, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components. For example, in at least one embodiment, system 11 includes a processing module 200 that implements an encoding module. In various embodiments, system 11 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 11 is configured to implement one or more of the aspects described herein.
[0045] Input to processing module 200 may be provided via various input modules as shown in block 231, previously described with respect to FIG. 2C.
[0046] The various elements of system 11 may be provided within a unitary housing, where the various elements may be interconnected and transmit data between them using any suitable connection arrangement, e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards. For example, in system 11, processing module 200 is interconnected to the other elements of system 11 by bus 2005.
[0047] The communication interface 2004 of the processing module 200 allows the system 11 to communicate over the communication channel 12 .
[0048] In various embodiments, data is streamed or otherwise provided to system 11 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal in these embodiments is received via communication channel 12 and communication interface 2004 adapted for Wi-Fi communication. Communication channel 12 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the RF connection of input block 231 is used to provide streaming data to system 11.
[0049] As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0050] The data provided to system 11 can be provided in different formats. In various embodiments, these data are encoded and conform to known video compression formats such as AV1, VP9, VVC, HEVC, AVC, etc. In various embodiments, these data are raw data provided, for example, by picture and / or audio capture modules connected to or included in system 11. In that case, processing module 200 is responsible for encoding these data.
[0051] System 11 can provide output signals to various output devices, such as system 13, that can store and / or decode the output signals.
[0052] Various implementations involve decoding. "Decoding," as used herein, can encompass all or part of the processing performed on received video data to generate a final output suitable for display, for example. In various embodiments, such processes include deep neural network-based decoding processes and deep hyper-prior neural network-based decoding processes.
[0053] Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.
[0054] Various implementations involve encoding. Similar to the above discussion regarding "decoding," as used herein, "encoding" can encompass, for example, all or part of a process performed on an input video sequence to generate encoded video data. In various embodiments, such processes include deep neural network-based encoding processes and deep hyper-prior neural network-based encoding processes.
[0055] Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process as a whole will be clear based on the context of the specific description and will be well understood by one of ordinary skill in the art.
[0056] Where a figure is presented as a flow diagram, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that the figure also provides a flow diagram of the corresponding method / process.
[0057] Various embodiments refer to a rate-distortion tradeoff. In particular, a balance or tradeoff between rate and distortion is usually considered during training of deep NN-based compression methods. Rate-distortion optimization is usually formulated to minimize a rate-distortion function, which is a weighted sum of rate and distortion. NN-based compression methods have trainable parameters found by gradient-based optimization and hyperparameters that are initially defined. To minimize the rate-distortion optimization problem, there are various approaches to finding these hyperparameters. For example, these approaches may be based on extensive testing of all deep neural network parameter values, involving a thorough evaluation of their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. Also, to reduce computational complexity, faster approaches may be used, particularly calculation of approximate distortion based on a predicted or predicted residual signal rather than the reconstructed signal. A hybrid of these two approaches may also be used, such as by using approximate distortion for only some of the possible deep NN network parameters and full distortion for other deep NN network parameters. Other approaches evaluate only a subset of possible coding options. More generally, many approaches employ any of a variety of techniques to perform optimization, but the optimization is not necessarily a complete assessment of both the coding cost and the associated distortion.
[0058] Implementations and aspects described herein may be implemented as, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed feature may also be implemented in other forms (e.g., an apparatus or program). For example, an apparatus may be implemented in appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, where processor refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include, for example, communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0059] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation," as well as other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with that embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment," or "in one implementation" or "in an implementation" in various places throughout this application, as well as other variations, are not necessarily all referring to the same embodiment.
[0060] Additionally, the application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, retrieving information from memory, or retrieving information from, for example, another device, module, or user.
[0061] Additionally, the application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0062] Additionally, the application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" typically involves in some manner, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0063] Use of any of " / ", "and / or", "at least one of", "one or more of", e.g., "A / B", "A and / or B", "at least one of A and B", "one or more of A and B" should be understood to be intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of both alternatives (A and B). As a further example, "A, B, and / or C" and "at least one of A, B, and C", "one or more of A, B, and C" are intended to encompass selection of only the first listed alternative (A), or selection of only the second listed alternative (B), or selection of only the third listed alternative (C), or selection of only the first and second listed alternatives (A and B), or selection of only the first and third listed alternatives (A and C), or selection of only the second and third listed alternatives (B and C), or selection of all three alternatives (A, B, and C). This can be expanded to include as many items as listed, as would be apparent to one skilled in this and related arts.
[0064] Also, as used herein, the term "signaling" specifically refers to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a parameter or step size. In this way, embodiments may use the same parameters on both the encoder and decoder sides. Thus, for example, an encoder may transmit a specific parameter to a decoder (explicit signaling) so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling without transmission (implicit signaling) may be used to simply allow the decoder to know and select the specific parameter. By avoiding transmitting any actual function, bit savings are realized in various embodiments. It will be appreciated that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. While the above relates to the verb form of the term "signal," the term "signal" may also be used herein as a noun.
[0065] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the encoded video data of the described embodiments. For example, such a signal can be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding the video data and modulating a carrier wave with the encoded video data. The information carried by the signal can be, for example, analog or digital information. As is known, the signal can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
[0066] Figure 3 illustrates the training phase of a deep NN-based video compression system.
[0067] The training phase of FIG. 3 is performed, for example, by the processing module 200.
[0068] In step 301, the processing module 200 generates a deep NN-based encoding g a Let x ∈ R be an input picture of size n × n with three components. n×n×3 The output of the deep NN-based coding, y(y ∈ R m×m×o ) is called the main latent (or main embedding) of the input picture x.
[0069] In step 303, the processing module 200 applies a quantization Q(.) to the main latent y to obtain the main code of the input picture x.
[0070]
number
[0071]
number
[0072] In step 304, the processing module 200 performs inverse quantization (Q -1 ()) as the main code
[0073]
number
[0074]
number
[0075]
number
[0076] In step 302, the processing module 200 generates a deep hyper-priori NN-based encoding h a Applying to the main potential y, z ∈ R k×k×f We obtain the side latent (or side embedding) z, where z is the deep hyper-prior NN-based coding h a is the hyper-prior entropy model P h Based on.
[0077] In step 305, the processing module 200 applies quantization (Q()) to the side latent z to obtain the side code
[0078]
number
[0079] In step 306, the processing module 200 performs inverse quantization (Q -1()) as a side code
[0080]
number
[0081]
number
[0082] Restructured side potential
[0083]
number
[0084]
number
[0085]
number
[0086] In step 307, the processing module 200 generates a deep hyper-ahead decoder h s Reconfigured side potential
[0087]
number
[0088]
number
[0089] In step 308, the processing module 200 generates a deep NN-based decoder g s The reconfigured main potential
[0090]
number
[0091]
number
[0092] During the training phase illustrated in Figure 3, the side codes
[0093]
number
[0094]
number
[0095]
number
[0096]
number
[0097]
number
[0098] In step 309, the processing module 200 calculates the factorized entropy model p f , the reconstructed side potential
[0099]
number
[0100]
number
[0101]
number
[0102]
number
[0103]
number
[0104]
number
[0105] Meanwhile, the main code
[0106]
number
[0107]
number
[0108]
number
[0109]
number
[0110] In step 310, the processing module 200 uses the PMF of the trained Gaussian to calculate the main code.
[0111]
number
[0112]
number
[0113]
number
[0114] actually,
[0115]
number
[0116] Main Code
[0117]
number
[0118]
number
[0119] In the process of Figure 3, deep NN-based encoding g a (or (g a (.;φ)))), deep NN-based decoding g s (or (g s(.;θ))), Deep hyper-prior NN-based coding h a (or (h a (.;φ))), Deep Hyper-prior NN based decoding h s (or h s (.;Θ)), and the factorized entropy model p f (or p f A neural network (.;ω)) consists of multiple neural network layers, such as convolutional layers. Each neural network layer can be described as a function that first multiplies the input by a tensor, adds a vector called a bias, and then applies a nonlinear function to the resulting value. The shape (and other properties) of the tensor and the type of nonlinear function are called the network's architecture. The values of the tensor and bias are expressed in terms of weights. The weights and, if applicable, the parameters of the nonlinear function are called parameters. The architecture and parameters define the model. Deep neural network-based coding g a (or (g a (.;φ))), deep NN-based decoding g s (or (g s (.;θ))), Deep hyper prior NN based coding h a (or (h a (.;φ))), Deep Hyper-prior NN based decoding h s (or h s (.;Θ)), and the factorized entropy model p f or p f p indicated by (.;ω) f (or p f The parameters of the model used in (.;ω)) are denoted by φ, θ, φ, Θ, and ω, respectively.
[0120] The model must be trained on a large database of pictures D to learn its parameters. Typically, the model parameters are optimized to minimize the training loss LOSS, given by Equation 1.
[0121]
number
[0122]
number
[0123]
number
[0124]
number
[0125]
number
[0126] FIG. 4 illustrates the video encoding process based on the trained deep NN-based video compression system.
[0127] For example, the trained deep NN-based video compression system is that of Figure 3 (i.e., the deep NN-based video compression system with parameters φ, θ, φ, Θ, and ω according to the method of Figure 3). The video encoding process is performed, for example, by a processing module of system 11. In Figure 4, the same reference numbers are used for steps already described in relation to Figure 3.
[0128] In step 301, the processing module 200 generates a deep NN-based encoding g a is applied to the input picture x.
[0129] In step 303, the processing module 200 applies a quantization Q(.) to the main latent y to obtain the main code
[0130]
number
[0131] In step 302, the processing module 200 generates a deep hyper-priori NN-based encoding h a Apply to the main potential y to obtain the side potential z.
[0132] In step 305, the processing module 200 applies quantization to the side latent z to obtain the side code
[0133]
number
[0134] In step 306, the processing module 200 performs the dequantization on the sidecode.
[0135]
number
[0136]
number
[0137] As already mentioned above, the reconstructed side potential
[0138]
number
[0139]
number
[0140]
number
[0141]
number
[0142]
number
[0143]
number
[0144] In step 401, the processing module 200 performs arithmetic encoding (AE) on the main code.
[0145]
number
[0146] In step 309, the processing module 200 calculates the factorized entropy model p f The side cord
[0147]
number
[0148]
number
[0149] In step 402, the processing module 200 side codes the AE.
[0150]
number
[0151] FIG. 5 illustrates a video decoder based on the trained deep NN-based video compression system.
[0152] For example, the trained deep NN-based video compression system is that of Figure 3 (i.e., the deep NN-based video compression system with parameters φ, θ, φ, Θ, and ω according to the method of Figure 3). The video decoding process is performed, for example, by a processing module of system 13. In Figure 5, the same reference numbers are used for steps already described in relation to Figure 3.
[0153] In step 501, the processing module 200 applies arithmetic decoding (AD) to the side information contained in the video data (i.e., the encoded video stream) to obtain side codes using the learned PMF table provided by the entropy model factorized in step 309.
[0154] In step 306, the processing module 200 performs the dequantization on the sidecode.
[0155]
number
[0156]
number
[0157] In step 307, the processing module 200 generates a deep hyper-ahead decoder h s Reconfigured side potential
[0158]
number
[0159]
number
[0160] In step 502, the processing module 200 applies the AD to the encoded main information contained in the video data using the learned PMFs determined in step 307. These distributions (and the learned PMF table) tell the AD how to read the main codes from the bitstream.
[0161] In step 304, the processing module 200 performs the dequantization on the main code.
[0162]
number
[0163]
number
[0164] In step 308, the processing module 200 performs deep NN-based decoding g s The reconfigured main potential
[0165]
number
[0166] As evoked in the introduction of this specification, the purpose of the following embodiments is to study what data can be inferred from explicitly encoded data available at the decoding side in deep NN-based video compression methods (also called full NN-based video compression methods), and how these inferred data can be used to improve the compression efficiency of the deep NN-based video compression methods.
[0167] One first example of data that can be inferred from explicitly coded data contained in video data is the reconstructed side potential.
[0168]
number
[0169]
number
[0170]
number
[0171] A second example of data that can be inferred from explicitly coded data contained in video data is the reconstructed main latent
[0172]
number
[0173]
number
[0174]
number
[0175] To date, these two gradients remain unused in the literature of deep NN-based video compression methods.
[0176] In the following embodiments, these gradients of entropy for the side latent and main latent are used at the decoding side to improve compression efficiency. These embodiments exploit the correlation that exists between these two gradients and other useful gradients.
[0177] More specifically, in these embodiments, the following is demonstrated: Restructured side potential
[0178]
number
[0179]
number
[0180]
number
[0181]
number
[0182]
number
[0183]
number
[0184]
number
[0185]
number
[0186] In the following, the best first and second step sizes are found on the encoding side, and the two step sizes are coded in the video data so that they can be used in the decoding process.
[0187] In the following, we provide some definitions, observations, and theorems that allow us to clarify various embodiments.
[0188] State-of-the-art deep NN-based video compression methods generally use the loss function expressed in Equation 1. This loss function can be viewed as an unconstrained multi-objective optimization problem, where the objective is to find the minimum bit length of the side information.
[0189]
number
[0190]
number
[0191]
number
[0192] The following definitions give an idea of what the optimal solution in multi-objective optimization is.
[0193] Definition: A Pareto optimal solution is an optimal solution for which no objective can be made better without making at least one objective worse. Different solutions can be obtained by using different importance weights for the objectives. If these solutions are Pareto optimal, they create a curve called the Pareto frontier curve.
[0194] Therefore, the goal of multi-objective optimization is to find Pareto-optimal points on the Pareto frontier curve, where different points on the curve are obtained by a given weighting term λ. In Equation (1), the first bit length term (i.e., the minimum bit length of the side information
[0195]
number
[0196]
number
[0197] The following observations demonstrate useful properties of solutions to unconstrained multi-objective optimization problems.
[0198] Remark: A solution to a multi-objective optimization problem is Pareto optimal if and only if it satisfies the Karush-Kuhn-Tucker (KKT) condition. More specifically, in the case of an unconstrained multi-objective optimization problem, if the objective of the problem is:
[0199]
number
[0200]
number
[0201]
number
[0202] In other words, this observation states that at the optimal solution point, all objective gradients compete and no one can win anymore. All gradient-driven forces cancel each other out and the solution reaches a saddle point. This property is used to test the optimality of candidate solutions.
[0203] This observation also holds for end-to-end compression models. The following theorem shows how to use the KKT conditions for end-to-end video compression systems, such as the deep NN-based compression systems of Figures 3 and 4.
[0204] Theorem: A λ-tradeoff optimized end-to-end compression model is Pareto optimal if and only if the following two conditions are met:
[0205]
number
[0206] Proof: To show the similarity between the end-to-end loss and the unconstrained multi-objective optimization loss, the loss can be rewritten as:
[0207]
number
[0208]
number
[0209]
number
[0210]
number
[0211]
number
[0212]
number
[0213]
number
[0214]
number
[0215]
number
[0216] Since this problem has two sets of variables to be optimized, both variables should satisfy the KKT conditions.
[0217]
number
[0218]
number
[0219]
number
[0220]
number
[0221]
number
[0222] Since α1 = α2, these cancel each other out,
[0223]
number
[0224] Second variable
[0225]
number
[0226]
number
[0227]
number
[0228]
number
[0229]
number
[0230] Replace the objectives with their definitions and define the two edges as α2 and
[0231]
number
[0232] The proof is easy and almost trivial. However, the consequence of this theorem is important: if one of the existing end-to-end models is optimal (at least in terms of the training procedure, not in terms of compression performance), then the side latency
[0233]
number
[0234]
number
[0235]
number
[0236]
number
[0237] The following assumptions are important to the following embodiments. Assumption: Assume that all existing end-to-end video compression models are well trained and their solutions are Pareto optimal. The sum of the gradients in the theorem is 0 over the training set, and they are correlated for a given single picture.
[0238] To verify the above hypothesis, the correlation between pairs of gradients in the theorem is calculated. Tests show that the correlation between gradients in equation (Equation 2) is not very clear for a given single picture, and therefore weakly affects performance. However, in the latter case of equation (Equation 3), the correlation is very clear. These test results motivate us to use gradients available at the decoding side as a substitute for useful gradients that are unknown at the decoding side.
[0239] Side Potential
[0240]
number
[0241]
number
[0242]
number
[0243]
number
[0244]
number
[0245]
number
[0246]
number
[0247]
number
[0248] Proposal 1: Decode the side code and reconstruct the side potential
[0249]
number
[0250]
number
[0251]
number
[0252]
number
[0253]
number
[0254] ρ f can be found by explicit force from a small number of candidates or by any optimization method to find the optimum such that:
[0255]
number
[0256] In addition, the main potential
[0257]
number
[0258]
number
[0259]
number
[0260]
number
[0261]
number
[0262]
number
[0263]
number
[0264]
number
[0265] Proposal 2. Decrypt the main code and reconstruct the main latent
[0266]
number
[0267]
number
[0268]
number
[0269]
number
[0270]
number
[0271] ρ h can be found by explicit force from a small number of candidates or by any optimization method to find the optimum such that:
[0272]
number
[0273] In the following embodiments, suggestions 1 and 2 can be implemented together or independently.
[0274] In one embodiment, finding the best step size is done by identifying the best step size from some predefined candidate list of step sizes, or by performing a nonlinear optimization to find the best step size. h =p f This is done by finding the best step size starting with =0.
[0275] FIG. 6 illustrates a video encoding process based on a trained deep NN-based video compression system of one embodiment.
[0276] The process of Figure 6 is based on Proposals 1 and 2. The process of Figure 6 is performed by, for example, processing module 200 of system 11.
[0277] In step 601, the processing module 200 obtains an input picture x.
[0278] In step 602, which is identical to step 301, the processing module 200 generates a deep NN-based encoding g a is applied to the input picture x to obtain the main potential y.
[0279] In step 603, which is identical to step 302, the processing module 200 generates a deep hyper-priori NN-based encoding h aApply to the main potential y to obtain the side potential z.
[0280] In step 604, which is identical to step 305, the processing module 200 applies quantization to the side latent z to obtain the side code
[0281]
number
[0282] In step 605, which is identical to step 309, the processing module 200 calculates the factorized entropy model p f The side cord
[0283]
number
[0284] In step 606, which is identical to step 402, the processing module 200 side codes the AE.
[0285]
number
[0286] In step 607, which is identical to step 306, the processing module 200 performs the dequantization as a sidecode.
[0287]
number
[0288]
number
[0289] In step 608, the processing module 200 calculates the reconstructed side potentials as follows:
[0290]
number
[0291]
number
[0292]
number
[0293] In step 609, the processing module 200 performs a first step size
[0294]
number
[0295] In step 610, the processing module 200 calculates a first step size,
[0296]
number
[0297]
number
[0298]
number
[0299] In step 611, the processing module 200 generates a deep hyper-ahead decoder h s The shifted reconstructed side potential
[0300]
number
[0301]
number
[0302] In step 612, the processing module 200
[0303]
number
[0304] In step 613, which is identical to step 303, the processing module 200 applies a quantization Q(.) to the main latent y to obtain the main code
[0305]
number
[0306] In step 614, the processing module 200 converts the AE into a main code.
[0307]
number
[0308] In step 615, which is identical to step 304, the processing module 200 performs the dequantization on the main code.
[0309]
number
[0310]
number
[0311] In step 616, the processing module 200 calculates the main side potential as follows:
[0312]
number
[0313]
number
[0314]
number
[0315] In step 617, the processing module 200 applies a second step size to the video data.
[0316]
number
[0317] In Figure 6, the steps associated with the first proposal are steps 608, 609, and 610, and the steps associated with the second proposal are steps 615, 616, and 617. Figure 6 represents an embodiment in which proposals 1 and 2 are implemented.
[0318] If only proposal 1 is implemented, steps 615, 616 and 617 are skipped.
[0319] If only proposal 2 is implemented, steps 608, 609, and 610 are skipped and step 611 is
[0320]
number
[0321]
number
[0322] FIG. 7 illustrates a video decoding process based on a trained deep NN-based video compression system of one embodiment.
[0323] The process of Figure 7 is based on Proposals 1 and 2. This process is performed, for example, by processing module 200 of system 13.
[0324] In step 701, which is identical to step 501, the processing module 200 applies arithmetic decoding (AD) to the side information contained in the video data to obtain side codes using the trained PMF table provided by the entropy model factorized in step 309.
[0325] In step 702, which is identical to step 306, the processing module 200 performs the dequantization on the sidecode.
[0326]
number
[0327]
number
[0328] In step 703, the processing module 200 performs a first step size
[0329]
number
[0330] In step 704, the processing module 200 calculates a first step size as follows:
[0331]
number
[0332]
number
[0333]
number
[0334] In step 705, the processing module 200 generates a deep hyper-ahead decoder h s The shifted side potential
[0335]
number
[0336]
number
[0337] In step 706, the processing module 200
[0338]
number
[0339] In step 707 , which is identical to step 502 , the processing module 200 applies AD to the encoded main information contained in the video data using the learned PMF table determined in step 706 .
[0340] In step 708, which is identical to step 304, the processing module 200 performs the dequantization on the main code.
[0341]
number
[0342]
number
[0343] In step 709, the processing module applies a second step size to the video data.
[0344]
number
[0345] In step 710, the processing module 200 calculates a second step size,
[0346]
number
[0347]
number
[0348]
number
[0349] In step 711, the processing module 200 performs deep NN-based decoding g s The shifted reconstructed main potential
[0350]
number
[0351] In Figure 6, the steps associated with the first proposal are steps 703 and 704, and the steps associated with the second proposal are steps 709 and 710. Figure 7 represents an embodiment in which proposals 1 and 2 are implemented.
[0352] If only proposal 1 is implemented, steps 709 and 710 are skipped and step 711 is executed to generate the reconstructed main potential.
[0353]
number
[0354]
number
[0355] If only proposal 2 is implemented, steps 703 and 704 are skipped and step 705 is executed to generate the reconstructed side potentials.
[0356]
number
[0357]
number
[0358] Several embodiments have been described above. The features of these embodiments may be provided alone or in any combination. Furthermore, the embodiments may include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types: A bitstream or signal comprising one or more of the main information, side information and first step size and / or second step size described, or variants thereof. Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal comprising one or more of the main information, side information, and first step size and / or second step size described, or variants thereof. A television, set-top box, mobile phone, tablet, or other electronic device that executes at least one of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that performs at least one of the described embodiments and displays the resulting picture (e.g., using a monitor, screen, or other type of display). A television, set-top box, mobile phone, tablet, or other electronic device that tunes to a channel (e.g., using a tuner) to receive a signal including an encoded video stream and that performs at least one of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that receives a signal containing an encoded video stream wirelessly (e.g., using an antenna) and that performs at least one of the described embodiments. A server, camera, mobile phone, tablet, or other electronic device that receives wirelessly (e.g., using an antenna) a signal containing the encoded video stream and that performs at least one of the described embodiments. A server, camera, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to transmit a signal including an encoded video stream and performs at least one of the described embodiments.
Claims
1. 1. A method comprising: Obtaining an input picture (601); applying a deep neural network based encoder to the input picture to obtain a main latency (602); applying a deep hyper-prior neural network based encoder to the main latent to obtain side latents (603); quantizing (604) the side latents to obtain side codes using a first quantization method; applying 605 the factorized entropy model to a sidecode to determine a first probability mass function for the sidecode under the first quantization method; arithmetically encoding the sidecodes based on the first probability mass function in video data (606); dequantizing (607) the side codes to obtain reconstructed side latents; Determining 608 a first step size based on a gradient of the entropy of the side information relative to the reconstructed side latent; encoding (609) information indicative of the first step size in the video data; Shifting the reconstructed side latent (610) using the first step size and the gradient of the entropy of side information with respect to the reconstructed side latent; applying a deep hyper-ahead decoder to the shifted reconstructed side latents (611) to obtain information representing a probability distribution that models the main latent; Obtaining 612 a second probability mass function of a main code using the probability distribution that models the main latent; quantizing (613) the main latent to obtain the main code; arithmetically encoding (614) the main code based on the second probability mass function in the video data; A method comprising:
2. Dequantizing (615) the main code to obtain a reconstructed main latent; Determining 616 a second step size based on a gradient of the entropy of the main information with respect to the main latent; encoding (617) information indicative of the second step size in the video data; The method of claim 1 further comprising:
3. 1. A method comprising: Obtaining an input picture (601); applying a deep neural network based encoder to the input picture to obtain a main latency (602); applying a deep hyper-prior neural network based encoder to the main latent to obtain side latents (603); quantizing (604) the side latents to obtain side codes using a first quantization method; applying 605 the factorized entropy model to a sidecode to determine a first probability mass function for the sidecode under the first quantization method; arithmetically encoding the sidecodes based on the first probability mass function in video data (606); dequantizing (607) the side codes to obtain reconstructed side latents; applying 611 a deep hyper-ahead decoder to the reconstructed side latents to obtain information representing a probability distribution that models the main latent; Obtaining 612 a second probability mass function of a main code using the probability distribution that models the main latent; arithmetically encoding (614) the main code based on the second probability mass function in the video data; Dequantizing (615) the main code to obtain a reconstructed main latent; Determining 616 a second step size based on a gradient of the entropy of the main information with respect to the main latent; encoding (617) information indicative of the second step size in the video data; A method comprising:
4. 1. A method comprising: applying arithmetic decoding to side information contained in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model (701); dequantizing (702) the side codes to obtain reconstructed side latents; Decoding (703) information representing the first step size based on a gradient of the entropy of the side information with respect to the reconstructed side latent; Shifting the reconstructed side latent (704) using the first step size and the gradient of the entropy of side information with respect to the reconstructed side latent; applying 705 a deep hyper-predecoder to the shifted reconstructed side latents to obtain information representing a probability distribution that models a main latent; Obtaining (706) a second probability mass function of a main code using the probability distribution that models the main latent; arithmetically decoding (707) a main code from the video data using the second probability mass function; Dequantizing (708) the main code to obtain a reconstructed main latent; applying a deep neural network based decoder to the reconstructed main latent (711) to obtain an output picture; A method comprising:
5. decoding (709) information representing a second step size based on a gradient of entropy of main information relative to the main latency from the video data; Shifting the reconstructed main latent (710) using the second step size and the gradient of the entropy of main information with respect to the main latent; wherein the deep neural network based decoder is applied to the shifted reconstructed main latent.
6. 1. A method comprising: applying arithmetic decoding to side information contained in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model (701); dequantizing (702) the side codes to obtain reconstructed side latents; applying 705 a deep hyper-ahead decoder to the reconstructed side latents to obtain information representing a probability distribution that models the main latent; Obtaining (706) a second probability mass function of a main code using the probability distribution that models the main latent; arithmetically decoding (707) a main code from the video data using the second probability mass function; Dequantizing (708) the main code to obtain a reconstructed main latent; decoding (709) information representing a second step size based on a gradient of entropy of main information relative to a main latency from the video data; Shifting the reconstructed main latent (710) using the second step size and the gradient of the entropy of main information with respect to the main latent; applying a deep neural network based decoder to the shifted reconstructed main latent (711) to obtain an output picture; A method comprising:
7. A device comprising an electronic circuit, the electronic circuit comprising: Get an input picture (601); applying 602 a deep neural network based encoder to the input picture to obtain a main latency; applying 603 a deep hyper-prior neural network based encoder to the main latent to obtain side latents; quantizing (604) the side latents to obtain side codes using a first quantization method; applying 605 the factorized entropy model to the sidecode to determine a first probability mass function for the sidecode under the first quantization method; arithmetically encoding (606) the sidecodes based on the first probability mass function in the video data; dequantizing 607 the side codes to obtain reconstructed side latents; determining 608 a first step size based on a gradient of the entropy of the side information relative to the reconstructed side latent; encoding (609) information indicative of the first step size in the video data; shifting 610 the reconstructed side latent using the first step size and the gradient of the entropy of side information with respect to the reconstructed side latent; applying 611 a deep hyper-ahead decoder to the shifted reconstructed side latents to obtain information representing a probability distribution that models the main latent; obtaining 612 a second probability mass function for a main code using the probability distribution that models the main latent; quantizing (613) the main latent to obtain the main code; arithmetically encoding (614) the main code based on the second probability mass function in the video data; The device that is configured for
8. The electronic circuit dequantizing 615 the main code to obtain a reconstructed main latent; determining 616 a second step size based on the gradient of the entropy of the main information relative to the main latent; encoding (617) information indicative of the second step size in the video data; The device of claim 7 further configured for:
9. A device comprising an electronic circuit, the electronic circuit comprising: Get an input picture (601); applying 602 a deep neural network based encoder to the input picture to obtain a main latency; applying 603 a deep hyper-prior neural network based encoder to the main latent to obtain side latents; quantizing (604) the side latents to obtain side codes using a first quantization method; applying 605 the factorized entropy model to the sidecode to determine a first probability mass function for the sidecode under the first quantization method; arithmetically encoding (606) the sidecodes based on the first probability mass function in the video data; dequantizing 607 the side codes to obtain reconstructed side latents; applying 611 a deep hyper-ahead decoder to the shifted reconstructed side latents to obtain information representing a probability distribution that models the main latent; obtaining 612 a second probability mass function for a main code using the probability distribution that models the main latent; arithmetically encoding (614) the main code in the video data based on the second probability mass function; dequantizing 615 the main code to obtain a reconstructed main latent; determining 616 a second step size based on the gradient of the entropy of the main information relative to the main latent; encoding (617) information indicative of the second step size in the video data; The device that is configured for
10. A device comprising an electronic circuit, the electronic circuit comprising: applying arithmetic decoding to side information contained in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model (701); dequantizing 702 the side codes to obtain reconstructed side latents; Decoding 703 information representing the first step size based on a gradient of the entropy of the side information with respect to the reconstructed side latent; shifting 704 the reconstructed side latent using the first step size and the gradient of the entropy of side information with respect to the reconstructed side latent; applying 705 a deep hyper-ahead decoder to the shifted reconstructed side latents to obtain information representing a probability distribution that models the main latent; obtaining 706 a second probability mass function for a main code using the probability distribution that models the main latent; arithmetically decoding (707) a main code from the video data using the second probability mass function; Dequantizing 708 the main code to obtain the reconstructed main latent; applying 711 a deep neural network based decoder to the reconstructed main latent to obtain an output picture; The device that is configured for
11. The electronic circuit decoding (709) information representing a second step size based on a gradient of entropy of main information relative to a main latency from the video data; Shifting 710 the reconstructed main latent using the second step size and the gradient of the entropy of the main information with respect to the main latent; 11. The device of claim 10, further configured for: wherein the deep neural network based decoder is applied to the shifted reconstructed main latent.
12. A device comprising an electronic circuit, the electronic circuit comprising: applying arithmetic decoding to side information contained in the video data to obtain side codes using a first probability mass function provided by the factorized entropy model (701); dequantizing 702 the side codes to obtain reconstructed side latents; applying 705 a deep hyper-ahead decoder to the reconstructed side latents to obtain information representing a probability distribution that models the main latent; obtaining 706 a second probability mass function for a main code using the probability distribution that models the main latent; arithmetically decoding (707) a main code from the video data using the second probability mass function; Dequantizing 708 the main code to obtain the reconstructed main latent; decoding (709) information representing a second step size based on a gradient of entropy of main information relative to a main latency from the video data; shifting 710 the reconstructed main latent using the second step size and the gradient of the entropy of the main information with respect to the main latent; applying 711 a deep neural network based decoder to the reconstructed main latent to obtain an output picture; The device that is configured for
13. A computer program comprising program code instructions for implementing the method according to any of the preceding claims 1 to 6.
14. A non-transitory information storage medium storing program code instructions for implementing the method according to any of the preceding claims 1 to 6.
15. A signal generated by a method according to any preceding claim 1 to 3 or a device according to any preceding claim 7 to 9.
Citation Information
Patent Citations
Data compression using conditional entropy models
US20200027247A1
Methods And Apparatuses For Learned Image Compression
US20200160565A1
Method and apparatus for variable rate compression with a conditional autoencoder
US20200304147A1
Image encoding and decoding, video encoding and decoding: methods, systems and training methods
WO2022084702A1
Progressive data compression using artificial neural networks
WO2022159897A1