Training method of compression system based on end-to-end neural network

By quantizing and freezing the decoder parameters periodically during the training process and compensating the low complexity of the decoder in the encoder, the problems of decoder complexity and floating-point operation dependence in the video compression method based on neural networks in low-end devices are solved, and high-efficiency video decoding and reducing computational complexity are achieved.

CN120051780APending Publication Date: 2025-05-27INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071871.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing video compression methods based on neural networks are difficult to achieve high-efficiency video decoding in low-end devices, mainly due to the complexity of the decoder and its dependence on floating-point operations, making it difficult to implement computing complexity and storage management.

Method used

By quantizing and freezing decoder parameters periodically during the training process, the complexity of the decoder is reduced, and the low complexity of the decoder is compensated in the encoder, avoiding the use of floating point operations, using integer convolution and ReLU activation functions to simplify the decoder.

Benefits of technology

High-efficiency video decoding in low-end devices is achieved, maintaining rate distortion performance similar to traditional decoders, while reducing computational complexity and storage management requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051780A_ABST
    Figure CN120051780A_ABST
Patent Text Reader

Abstract

A method is disclosed that includes training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, where the method includes quantizing and freezing the learned decoder parameters layer by layer at different cycles during the training.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 414,977, filed on October 11, 2022, which is incorporated herein by reference in its entirety. Technical field

[0003] At least one of the present embodiments generally relates to a method for training an encoder neural network and a decoder neural network. Background art

[0004] In recent years, new neural - network - based image and video compression methods have been developed. Contrary to traditional methods that apply predefined prediction patterns and transforms, ANN - based methods rely on many parameters learned on large datasets by iteratively minimizing a loss function during the training phase. In the case of compression, the loss function is defined, for example, by a rate - distortion cost, where the bitrate represents an estimate of the bitrate of the encoded bitstream, and the distortion measures the quality of the decoded video relative to the original input. Traditionally, the quality of the decoded input image is optimized, for example, based on measurements of mean - square error or approximations of visual quality perceived by humans.

[0005] The Joint Video Exploration Team (JVET) between ISO / MPEG and ITU is currently researching ANN - based tools to replace some modules in the latest video coding standard H.266 / VVC and to replace the entire structure by an end - to - end auto - encoder approach. Summary of the invention

[0006] In one embodiment, a method is disclosed that includes training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, wherein the method includes, during training, quantizing and freezing the learned decoder parameters layer - by - layer for different epochs.

[0007] In another embodiment, a method is disclosed that includes training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, wherein the method includes, during training, quantizing and freezing the learned encoder parameters layer - by - layer for different epochs.

[0008] Additional embodiments are described herein that can be used alone or in combination.

[0009] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to implement an encoding method or a decoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for video encoding or decoding according to the methods described herein.

[0010] One or more embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above method. One or more embodiments also provide methods and apparatuses for transmitting or receiving a bitstream generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 A block diagram of a system in which aspects of the present embodiment can be implemented is illustrated;

[0012] Figure 2 An example of a compression system based on an end-to-end neural network is illustrated;

[0013] FIG. 3 illustrates an example of an end-to-end autoencoder of a neural network for image compression, referred to as a factorization prior, according to the prior art;

[0014] Figure 4A and Figure 4B Examples of end-to-end autoencoders of neural networks for image compression according to various embodiments are illustrated;

[0015] Figure 5 An example of a rectified linear unit (ReLU) activation function is depicted;

[0016] Figure 6 A flowchart of a method for training an end-to-end autoencoder of a neural network according to another embodiment is illustrated; and

[0017] Figure 7 A flowchart of a method for training an end-to-end autoencoder of a neural network according to another embodiment is illustrated. DETAILED DESCRIPTION

[0018] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and are typically described in a way that may sound restrictive, at least to show individual characteristics. However, this is for the purpose of clear description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide additional aspects. In addition, these aspects can also be combined and interchanged with the aspects described in earlier applications.

[0019] Aspects described and contemplated in this application may be implemented in many different forms. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These aspects and other aspects may be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing a bitstream generated according to any of the described methods.

[0020] In this application, the terms "reconstruct" and "decode" may be used interchangeably, the terms "encoded" or "coded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Generally but not necessarily, the term "reconstruct" is used on the encoder side, while "decode" is used on the decoder side.

[0021] Various methods are described herein, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first", "second", etc. may be used in various embodiments to modify elements, components, steps, operations, etc., such as for example "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and the first decoding may occur, for example, before, during, or in a time period overlapping with the second decoding.

[0022] Figure 1 A block diagram illustrating an example of a system in which various aspects and embodiments may be implemented. System 100 may be embodied as a device including various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 may be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0023] System 100 includes at least one processor 110, which is configured to execute instructions loaded therein for implementing various aspects described, for example, in the present application. The processor 110 may include an embedded memory, an input / output interface, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0024] System 100 includes an encoder / decoder module 130, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents the (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of System 100, or may be incorporated within the processor 110, as a combination of hardware and software known to those skilled in the art.

[0025] The program code to be loaded onto the processor 110 or the encoder / decoder module 130 to execute the various aspects described in the present application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in the present application. Such stored items may include but are not limited to input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0026] In some embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations.

[0027] As indicated in block 105, input can be provided to the elements of the system 100 through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 1 Other examples not shown include composite video.

[0028] In various embodiments, the input device of block 105 has associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting the narrower frequency band again to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as, for example, a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various functions of these, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or a baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0029] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor 110 as needed. Similarly, various aspects of the USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 110 as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, the processor 110 and the encoder / decoder module 130, which operate in combination with memory and storage elements to process the data stream as needed for presentation on the output device.

[0030] The various elements of system 100 may be provided within an integrated housing, within which the various elements may be interconnected using a suitable connection arrangement 115 and data may be transmitted therebetween, the connection arrangement 115 being, for example, an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.

[0031] System 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0032] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, and the set-top box transmits data through an HDMI connection of the input box 105. Still other embodiments use an RF connection of the input box 105 to provide streaming data to system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0033] System 100 may provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The display 165 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 may be used for a television, a tablet, a laptop, a cellular phone (mobile phone), or other devices. The display 165 may also be integrated with other components (e.g., in a smart phone) or be standalone (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 185 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of system 100. For example, a disc player performs the function of playing the output of system 100.

[0034] In various embodiments, control signals are transmitted between system 100 and display 165, speaker 175, or other peripheral device 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 via communication interface 150 using communication channel 190. Display 165 and speaker 175 can be integrated into a single unit with other components of system 100 in an electronic device (e.g., a television). In various embodiments, display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0035] For example, if the RF portion of input 105 is part of a stand-alone set-top box, display 165 and speaker 175 can alternatively be separate from one or more other components. In various embodiments where display 165 and speaker 175 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0036] These embodiments can be implemented by computer software implemented by processor 110, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments can be implemented by one or more integrated circuits. Memory 120 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, as non-limiting examples, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 110 can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0037] Figure 2 An example of an end-to-end neural network-based compression system is shown. The input X to the encoder part of the network can include:

[0038] - images or frames of a video;

[0039] - a portion of an image;

[0040] - a tensor representing a set of images;

[0041] - a tensor representing a portion (crop) of a set of images.

[0042] In each case, the input can have one or more components or channels, such as monochrome, RGB, or YCbCr components. AsFigure 2 As shown, the input X is fed into the encoder neural network g a () (210, also referred to as the analysis transform). g a () is typically a series of downsampling convolutions followed by an activation function. A large stride in the convolution can be used to reduce the spatial resolution. The stride is the step size between the current position and the next position of the filter kernel (defined, for example, in terms of the number of pixels). When the stride is 1, the filter typically moves one pixel at a time in both the horizontal and vertical directions. When the stride is 2, the filter moves 2 pixels at a time. This produces a smaller output, e.g., C out ×H / 2×W / 2 —— if the input has the dimension C in ×H / 2×W / 2, where H and W are the spatial dimensions and C out and C in are the number of channels. The higher the stride, the smaller the output tensor. In other words, the encoder neural network (210) typically consists of a series of convolutional layers with strides, thus allowing the spatial resolution of the input to be reduced while increasing the depth, i.e., the number of channels of the input. Pooling (e.g., average pooling, max pooling, etc.) or squeezing operations (from spatial to depth via reshaping and permutation) can also be used instead of the stride convolutional layers. The encoder neural network can be regarded as the learned transform g a ().

[0043] The output of the analysis transform is Z = g a (X), which is a three-dimensional tensor (referred to as a tensor), also known as the latent tensor or latent representation. From a broader perspective, a set of latent variables constructs the latent space, which is also often used in the context of neural network-based end-to-end compression. In this article, the terms latent variable and latent coefficient can be used interchangeably. The latent representation Z is quantized (Q) and entropy encoded (EC) (220) into a binary stream (bitstream) for storage or transmission. Entropy coding utilizes the probability distribution of the symbols to be encoded. Hereinafter, it is assumed that EC embeds the quantization operation (Q). The bitstream is a set of encoded syntax elements and binary payloads representing the quantized symbols, which can be transmitted to the decoder or stored on a storage medium.

[0044] The decoder first decodes (ED, 230) the quantized symbols from the bitstream to obtain a quantized version of Z The decoder network g s () (240, also referred to as the synthesis transform) generates the reconstructed input: An approximation of the original X from the quantized latent representation . g s() is typically a series of upsampling convolutions (e.g., "deconvolution" or convolution followed by an upsampling filter) or depth-to-space operations. The decoder network (240) can be seen as a learned inverse transform g that operates on the quantized coefficients s (), or denoising and generative transformations. The output of the decoder is the reconstructed image or set of images

[0045] The encoder and decoder neural networks of the encoder consist of multiple layers, such as convolutional layers. Each layer can be described as a function that first multiplies the input by weights, adds a vector called a bias, and then applies a non-linear function (activation function) to the resulting values. The values of the weights and biases are denoted by the term "neural network parameters". In such a compression system, the encoder and decoder are fixed based on a predetermined model assumed to be known during encoding and decoding (inference phase). For this purpose, the encoder neural network and the decoder neural network are trained simultaneously (i.e., learning the neural network parameters) so that they are compatible. In practice, to learn the parameters of the encoder and decoder (e.g., weights and biases), the end-to-end neural network is trained on a large-scale database D of images. The learning phase includes a forward pass and a backward pass. The forward pass indicates the flow direction from "input" to "output". The backward pass indicates the flow direction from "output" to "input", during which the gradient of the loss function propagates backward. The purpose of the backward pass is to distribute the total error back to the network in order to update the parameters and thus minimize the cost function (loss function). The update is determined by the gradient of the cost function with respect to those parameters. The parameters are updated in such a way that when the next forward pass utilizes the updated parameters, the total error is reduced by a certain margin (until a minimum is reached).

[0046] The transform g a () and g s (), more precisely their parameters, such as weights and biases, are learned by minimizing a loss function (also called a cost function) that compares the output of the network with a known data set (e.g., the input image). The first type of loss function can be based on an "objective" metric, typically the mean squared error (MSE), or on structural similarity (SSIM).

[0047] The second type of loss function can be based on "subjective" (or proxy subjective), typically using a generative adversarial network (GAN) or an advanced visual metric via a proxy NN during the training phase.

[0048] The results obtained using the first type of loss function may not be as perceptually good as those of the second type of loss function, but they have higher fidelity to the original signal (image).

[0049] Several types of training datasets can be used to train the encoder and decoder networks. The same networks can first be trained on a general training set, allowing for satisfactory performance across a wide range of content types, and then it is possible to fine-tune the model using a specific training set for a particular purpose, thereby improving performance on domain-specific content.

[0050] Above, the probability distribution used by the entropy encoder EC, or a simpler distribution, is learned once (one distribution for each channel of the latent representation) and does not change depending on the input image. There are more complex architectures called "hyper autoencoders" where an additional NN (hyper prior) is added to the network to jointly learn the parametric probability distribution of the latent representation variables as the output of the encoder.

[0051] Once the parameters of the encoder neural network and the decoder neural network are learned, the network can be effectively used to encode a specific image X. Inference refers to the effective use of the neural network defined by the learned parameters. In other words, inference applies the trained neural network model and uses it to infer the result. Inference occurs after training because it requires the trained neural network model.

[0052] Figure 3 shows an example of an end-to-end autoencoder for image compression called a factorial prior, which is described in the document "Variational image compression with a scale hyperprior" by Ballé et al., ArXiv 1802.01436 Cs Eess Math, May 2018. This end-to-end autoencoder includes an encoder neural network 310 (associated with the transform g a ()) and a decoder neural network 320 (associated with the transform g s ()). In this end-to-end autoencoder, g a (310) includes 4 convolutional layers Conv and 3 non-linear generalized divisive normalization (GDN), and g s (320) includes 4 deconvolutional layers Deconv and 3 non-linear inverse generalized divisive normalization (iGDN). The parameters of the convolutional layers are labeled as the number of filters × kernel support height × kernel support / downsampling or upsampling stride, for example as Mx5x5. In the example, M = 192 or 128, and the stride is equal to 2. Downsampling is applied on the encoder side, and upsampling is applied on the decoder side.

[0053] The GDN consists of a linear transformation followed by division normalization in a generalized form. This activation / normalization layer includes division and thus requires the use and generation of intermediate and output floating-point values. The iGDN includes square roots and thus requires the use and generation of intermediate and output floating-point values. In the autoencoder described above, the quantizer (Q) rounds the floating-point values of the latent variables, i.e., the output of g a to integer values and then feeds these integer values to the entropy encoder (EC). The entropy codec (entropy encoder and entropy decoder) uses the cumulative distribution function (CDF) as a prior to compress the quantized latent tensor. These CDFs (p ψ ) are learned during training. In this version, the CDFs are frozen after the model training is completed. A CDF is computed for each channel in the latent representation, i.e., all the latent variables (coefficients) of the channel share the same prior for entropy coding. The parameters (e.g., weights and biases) of the encoder neural network and the decoder neural network are learned during training by minimizing a loss function defined as R + λ.D, where R is the rate and D is the distortion between the input image x and the reconstructed image . For example, the rate R is defined as and the distortion D is defined as where x is the input image, is the reconstructed image, y is the latent representation (i.e., the output of the encoder), is the quantized latent, and ψ represents the parameters of the cumulative distribution function (CDF).

[0054] Video compression systems - especially decoders embedded in low-end devices such as, for example, smartphones or set-top boxes - need to be able to handle videos with continuously increasing resolution and frame rate, which involves extremely challenging computational complexity and memory management. Despite the rise and rapid improvement of graphics processing units - which enable highly parallel floating-point operations in deep neural networks - lightweight decoder architectures are key for potential deployment in the foreseeable future.

[0055] The fully factorized prior model and its variants described above with respect to Figure 3 are widely used for end-to-end image and video compression. Almost all models use GDN (where division is performed) at the encoder and iGDN (where square roots are performed) at the decoder. Such floating-point operations have high computational complexity and memory management, which may be an obstacle for large-scale deployment in low-end devices.

[0056] When using a neural network (in the inference phase), quantization is used to avoid using floating-point operations because the training process cannot be fully executed using integer values. PyTorch is a machine learning framework based on the Torch library. It exposes three options for quantizing neural networks to reduce their complexity.

[0057] In the first option, called post-training dynamic quantization, the parameters of the network are dynamically quantized during the inference phase. The main drawback is that the quantization is performed after training. Without feedback, there is no guarantee that the performance estimates computed during training will be preserved during inference.

[0058] In the second option, called post-training static quantization, the model parameters, i.e., weights, biases, and activations, are quantized at the end of the training phase. Then, a dataset is used to adjust the parameters and reduce the distortion between the results of performing inference with and without quantization. The main advantage is that there are no changes during training as this is done after quantization. However, for complex networks with many layers, the performance suffers some accuracy loss.

[0059] In the third option, quantization-aware training called static quantization, the model is trained in floating-point form but uses pseudo-quantization modules, i.e., the operations are still performed in floating-point form but include clamping and rounding to simulate integer conversion. In this method, the entire network is quantized with the same quantization bits. In terms of performance, this is the best method, but it is not flexible. Additionally, the pseudo-quantization method is not optimal for representing the actual exact inference that will be performed.

[0060] In the document "Integer Networks for Data Compression with Latent-Variable Models" by Ballé et al. published in ICLR in 2019, the authors explain that floating-point-based ANNs can cause serious malfunctions when deployed on heterogeneous platforms such as embedded platforms. They suggest using integer arithmetic in these ANNs. To this end, they expose a quantized network, the so-called integer network. When the parameters are quantized, the gradients are computed and kept in floating-point form. The main drawback is that integer networks are used for both the encoder and the decoder, while in some cases, an integer decoder is sufficient and tuning for heterogeneous platforms is required.

[0061] The embodiments described below may aim to reduce the complexity of the decoder by discarding division, square root, and floating point in the decoder while maintaining a similar rate-distortion cost. In these embodiments, the encoder compensates for the low complexity of the decoder during the training phase. To avoid deviation of the decoded output from the encoding expectation, the encoder needs to know the bit-exact behavior of the decoder. Therefore, the encoder neural network and the decoder neural network are trained simultaneously.

[0062] The complexity of the decoder is reduced by replacing the computationally expensive inverse GDN of the normalization layer with a basic activation layer and performing integer convolution by quantizing the weights and biases of the convolutional layers in the decoder. The reduction in decoder complexity is compensated by the encoder during training. This makes it possible to maintain a rate-distortion performance similar to that of the decoder disclosed in Figure 3. After a predetermined number of epochs of the training set (during which the entire system is "normally" trained end-to-end), the decoder parameters are quantized and further frozen / fixed, i.e., no longer updated. Then, the training continues with backpropagation of the gradients to update the encoder parameters while the quantized decoder parameters are no longer updated.

[0063] The disclosed method is not limited to the decoder and can be implemented in the encoder in addition to or only in the encoder. In the following description, although the method is not limited to the decoder, the decoding process implemented in an embedded low-end platform is focused on. In fact, in the decoder, complexity is generally more critical. However, the same principle can be implemented on the encoder side.

[0064] Figure 4A An example of a neural network end-to-end autoencoder for image compression according to an embodiment is shown. The method proposed herein is not limited to using an autoencoder. Any end-to-end differentiable codec can be considered, such as a codec using a video compression transformer.

[0065] The end-to-end autoencoder includes an encoder neural network 410 (associated with the transform g a ()) and a decoder neural network 420 (associated with the transform g s ()). In this end-to-end autoencoder, g a (310) includes n convolutional layers Conv and m non-linear generalized division normalization (GDN), and g s (320) includes p deconvolutional layers Deconv i (i ∈ {0, 1, 2, 3}) and q ReLU (representing rectified linear unit). Other low-complexity activation functions can be used instead of ReLU, such as leaky ReLU, etc. In the example depicted in Figure 4, n = p = 4, m = q = 3. However, other values can be used.

[0066] The encoder neural network architecture 410 can be the same as or similar to the encoder neural network architecture 310 of the encoder in FIG. 3. The GDN includes a linear transformation followed by a generalized form of divisive normalization. This activation / normalization layer includes division, thus requiring the use and generation of intermediate floating-point values and output floating-point values. In this end-to-end autoencoder, before being fed into the entropy encoder, the quantizer (Q) rounds the floating-point values of the latent variables, i.e., the output of g a to integer values. The entropy codec (entropy encoder and entropy decoder) uses the cumulative distribution function (CDF) as a prior to compress the quantized latent tensor. These CDFs (p ψ ) are learned during training. In this version, the CDFs are frozen after the model training is completed. One CDF is computed for each channel in the latent representation, i.e., all the latent variables (coefficients) of the channel share the same prior for entropy coding. The parameters of the encoder and decoder networks are learned by minimizing the loss function defined as R + λ.D, where R is the rate and D is the distortion between the input image x and the reconstructed image For example, the rate R is defined as and the distortion D is defined as where x is the input image, is the reconstructed image, y is the latent representation (i.e., the output of the encoder), is the quantized latent, and ψ represents the parameters of the distribution function (CDF). However, the method disclosed with respect to FIG. 4 is not limited to the type of loss function.

[0067] The decoder neural network architecture 420 is simplified with respect to the decoder neural network architecture 320 to avoid square roots and divisions. For this purpose, each iGDN activation function is replaced by a ReLU activation function, such as the function depicted in Figure 5 which outputs the maximum value between the input value and 0. Other low-complexity activation functions can be used instead of ReLU, such as leaky ReLU, etc. In addition, integer deconvolution replaces each iGDN. This decoder neural network architecture does not involve any square roots or divisions, but only integer operations (e.g., addition and multiplication), and is thus simplified.

[0068] Integer deconvolution is obtained by quantizing and freezing (also known as setting to fixed values or fixing) the decoder parameters (e.g., the weights and biases of the deconvolution layer) during training after a given number of epochs, while continuing to train the encoder so that it adapts to the quantized frozen / fixed decoder parameters. At the end of the learning phase, the parameters (e.g., weights and biases) of the deconvolution layer of the decoder are thus integer parameters. The parameters of the convolutional layer of the encoder can be floating-point parameters. However, these floating-point parameters are learned knowing that the decoder uses integer parameters and ReLU activation functions.

[0069] In other words, the encoder neural network and the decoder neural network are trained together, and the encoder takes into account the simplification of the decoder during training (i.e., learning the encoder parameters). The encoder neural network can thus compensate for the distortion caused by quantizing the parameters (e.g., weights and biases) of the decoder neural network. A decoder neural network with lower complexity can be embedded in a low-end device, such as a smart phone or a set-top box, while retaining the rate-distortion performance.

[0070] In Figure 4B depicted in Figure 4A a variant of, the encoder neural network 412 is modified such that each GDN is replaced by a ReLU.

[0071] Figure 6 Illustrated is an example of a flowchart of a method for training a neural network end-to-end autoencoder according to an embodiment.

[0072] Consider multiple epochs number_of_epochs for training, i.e., when the current epoch number epoch_curr is equal to number_of_epochs, training stops. Both epoch_curr and number_of_epochs are integers. In terms of an artificial neural network, an epoch refers to one pass through the complete training dataset. In one epoch, all the data in the training dataset is used exactly once. Thus, the current epoch number epoch_curr is incremented by one for each training pass through the complete training dataset. An epoch consists of one or more batches, where a portion of the training dataset is used to train the neural network. The current epoch number epoch_curr is first initialized to a value, such as 0.

[0073] In step S600, the encoder neural network and the decoder neural network are trained to learn their parameters (e.g., weights and biases) within one epoch (one pass through the complete training dataset). These parameters are floating-point parameters.

[0074] In step S602, the current epoch number epoch_curr is compared with a value epoch_freeze, e.g., epoch_freeze = number_of_epochs / 2. In the case where the current epoch number epoch_curr is lower than epoch_freeze, the method continues at step S604. At step S604, the current epoch number epoch_curr is incremented by one, and the method continues to the next epoch at step S600.

[0075] In the case where the current epoch number epoch_curr is greater than or equal to epoch_freeze, the method continues at step S606.

[0076] At step S606, the learned decoder parameters of all decoder layers (i.e., all deconvolution layers) are quantized and frozen / fixed. In other words, the parameters (e.g., weights and biases) of the deconvolution layer deconv of the decoder neural network are no longer updated until the end of training. i of the decoder neural network are no longer updated until the end of training.

[0077] The parameters of the decoder neural network are quantized to n-bit precision but are still temporarily stored as floating-point for floating-point operations during the training phase. The quantization range is [-2 nbits-1 , +2 nbits+1 .

[0078] quantizer = 2 n

[0079] Then, the matrix X = {deconv i} i=0..3 is quantized as follows:

[0080]

[0081] where newround() = (torch.round(x) - x).detach() + x is defined using the Torch backend. In fact, uniform quantization is thus defined, which is differentiable in the case of forward or backward propagation. The classical round() function used in uniform quantization has a gradient of NULL. A function that returns a NULL gradient cannot be used for backward propagation. In other words, in the case of returning a NULL gradient, the network will not learn anything. A workaround is to use the straight-through estimator (STE) such that the gradient becomes non-NULL. The STE ignores the derivative of the round() function and passes the incoming gradient as if the function were an identity function.

[0082] In the first embodiment, the same quantizer is thus used for all deconvolution layers deconv i .

[0083]

[0084] where

[0085] quantizer = 2 n

[0086] n = nbits - 1 - ceil(log 2 (max(abs(X))))

[0087] In this case, the quantizer is stored and transmitted to the decoder. The same quantizer (deconv i *quantizer) is used to quantize the weights and biases of each deconv i and transmit them to the decoder for inference. During inference, all operations are performed in integers, and the output of each deconv i is clamped to [-2 nbits -1 , +2 nbits+1 . At the end of decoding, the output is finally rescaled to the initial range using the quantizer (e.g., [0, 256] if the output image is in 8 bits).

[0088] In a variant, a quantizer is set for each deconvolution layer deconv i .

[0089]

[0090] where

[0091]

[0092] n i = nbits - 1 - ceil(log 2 (max(abs(deconv i ))))

[0093] In this case, {quantizer i} is stored and transmitted to the decoder.

[0094] During the inference decoding phase, the decoder uses the transmitted associated quantizer quantizer i to apply a dequantizer to each deconvolution layer deconv i .

[0095] In step S608, the encoder neural network is trained, that is, the parameters of the encoder neural network are updated until a stop criterion is reached to continue learning the encoder parameters (e.g., weights and biases). The stop criterion can be reached when epoch_curr is equal to number_of_epochs, or when the loss is below a predetermined value.

[0096] The parameters of the deconvolution layers in the decoder neural network are frozen while continuing to learn, that is, updating the parameters of the convolution layers in the encoder neural network, making it possible to compensate for the distortion caused by the quantization of the network parameters (weights and biases). Thus, the loss continues to decrease until the end of training.

[0097] Figure 7 An example of a flowchart of a method for training a neural network end-to-end autoencoder according to another embodiment is illustrated. In Figure 6 the method, all parameters of the decoder neural network are quantized and frozen in the same epoch (i.e., when epoch_curr is greater than or equal to epoch_freeze).

[0098] In Figure 7 the embodiment, the parameters of the decoder neural network are quantized and frozen in different epochs. The index i is initialized to the index value associated with the deconvolution layer closest to the output, e.g., value 3. i is the index used to identify the deconvolution layers in the decoder neural network. In this description, i varies from 0 to 3. However, alternatively, i can vary from 1 to 4. This is just a matter of convention.

[0099] In step S700, the encoder neural network and the decoder neural network are trained to learn their parameters (e.g., weights and biases) within one epoch (going through the complete training dataset once). These parameters are floating-point parameters.

[0100] In step S702, the current epoch number epoch_curr is compared with the value epoch_freeze[i]. For example, epoch_freeze[3]=number_of_epochs / 2, epoch_freeze[2]=5*number_of_epochs / 8, epoch_freeze[1]=6*number_of_epochs / 8, epoch_freeze[0]=7*number_of_epochs / 8. In the case where the current epoch number epoch_curr is lower than epoch_freeze[i], the method continues at step S704. At step S704, the current epoch number epoch_curr is incremented by one, and the method continues to the next epoch at step S700.

[0101] In the case where the current epoch number epoch_curr is greater than or equal to epoch_freeze[i], the method continues at S706.

[0102] In step S706, the learned decoder parameters of the decoder layer of index i, i.e., the deconvolution layer deconv i are quantized and frozen. The various embodiments disclosed for quantization are similarly applied at step S706. In other words, the deconvolution layer deconv of the decoder neural network Figure 6 i ​The parameters (e.g., weights and biases) are no longer updated until the end of training. Since the last layer, i.e., the layer closest to the decoder output, is the most important layer in terms of its impact on the reconstruction performance, the parameters of this layer are first frozen. Regarding Figure 7 , first, the parameters of deconv3 are quantized and fixed, then the parameters of deconv2 are quantized and frozen, then the parameters of deconv1 are quantized and frozen. Finally, the parameters of deconv0 are quantized and frozen.

[0103] In step S708, the index i is decreased. At step S710, i is compared with zero (compared with 1 in the case where i varies from 1 to 4). In the case where i is strictly lower than zero, then the method continues at step S712, otherwise the method continues at step S700. It should be noted that the deconvolution layers can be indexed differently, i.e., the layer closest to the decoder output is indexed as 0, while the layer closest to the input is indexed as 3. In the latter case, i is initialized to the value 0, incremented by 1 at step S708, and compared with 3 at step S710. Thus, the parameters of the deconvolution layers are quantized and frozen layer by layer, from the layer closest to the decoder output to the layer closest to the decoder input.

[0104] In step S712, the encoder neural network is trained, i.e., the parameters of the encoder neural network are updated until a stopping criterion is reached to continue learning the encoder parameters (e.g., weights and biases). The stopping criterion can be reached when epoch_curr is equal to number_of_epochs, or when the loss is lower than a predetermined value.

[0105] In Figure 7 a variant of the method, the moment at which the learned decoder parameters of the decoder layer with index i are quantized and frozen is determined relative to the loss evolution rather than a fixed number of epochs.

[0106] In fact, the selection of the number of epochs is a problem in training neural networks. Too many epochs may cause overfitting of the training dataset, while too few epochs may result in an underfitted model. Early stopping is a method of stopping training once the model performance stops improving on the validation dataset. The early stopping technique consists in stopping training when the best validation error has passed at least Ne epochs, i.e., when there is a plateau.

[0107] Early stopping with plateau detection is described as follows:

[0108]

[0109] This early stopping is applied to each quantization and freezing of the deconvolution layers. First, the learned decoder parameters of the deconvolution layer deconv3 are quantized and frozen. Then, when a plateau is detected, the learned decoder parameters of the deconvolution layer deconv2 are quantized and frozen, and so on until deconv0. Thus, the parameters of the deconvolution layers are quantized and frozen layer by layer, from the layer closest to the decoder output to the layer closest to the decoder input.

[0110] As previously explained, the method disclosed with respect to Figure 6 and Figure 7 is not limited to the decoder and can also be implemented in the encoder, where the encoder includes GDN or ReLU activation functions as depicted in Figure 4A and Figure 4B In another embodiment, the method disclosed with respect to Figure 6 and Figure 7 can also be applied to both the encoder and the decoder during training. In the latter case, the deconvolution layers of the decoder are first quantized and frozen layer by layer at different epochs, while the deconvolution layers of the encoder are gradually quantized and frozen layer by layer from the layer closest to the encoder output (conv3) to the layer closest to the encoder input (conv0) at different epochs. This latter solution provides a lightweight encoder / decoder architecture.

[0111] In one embodiment, a method is disclosed that includes training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, where the method includes, during training, quantizing and freezing the learned decoder parameters layer by layer at different epochs.

[0112] In an example, quantizing and freezing the learned decoder parameters layer by layer at different epochs includes: quantizing and freezing the learned decoder parameters layer by layer at different epochs from the layer closest to the output of the decoder neural network to the layer closest to the input of the decoder neural network.

[0113] In an example, the same quantizer is used to quantize the learned decoder parameters of all decoding layers.

[0114] In an example, a specific quantizer is associated with each decoding layer of the decoder neural network and is used to quantize the learned decoder parameters of the decoding layer.

[0115] In an example, the decoder neural network includes deconvolution layers, each followed by a rectified linear unit.

[0116] In one embodiment, a decoder neural network is also disclosed, which includes deconvolution layers, each followed by a rectified linear unit, where the parameters of the decoder neural network are integer parameters learned by the above method.

[0117] In one embodiment, a method is disclosed, which includes training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, where the method includes, during training, quantizing and freezing the learned encoder parameters layer by layer in different epochs.

[0118] In an example, quantizing and freezing the learned encoder parameters layer by layer in different epochs includes: quantizing and freezing the learned encoder parameters layer by layer in different epochs from the layer closest to the output of the encoder neural network to the layer closest to the input of the encoder neural network.

[0119] In an example, the same quantizer is used to quantize the learned encoder parameters of all encoding layers.

[0120] In an example, a specific quantizer is associated with each encoding layer of the encoder neural network and is used to quantize the learned encoding parameters of the encoding layer.

[0121] In an example, the encoder neural network includes deconvolution layers, each followed by a rectified linear unit.

[0122] In one embodiment, an encoder neural network is also disclosed, which includes deconvolution layers, each followed by a rectified linear unit, where the parameters of the encoder neural network are integer parameters learned by the above method.

[0123] A computer program is also disclosed, which includes program code instructions for implementing the method disclosed above when executed by a processor.

[0124] A computer-readable storage medium is disclosed, on which instructions for implementing the method disclosed above when executed by a processor are stored.

[0125] Unless otherwise indicated or technically excluded, the aspects described in this application can be used alone or in combination.

[0126] Various numerical values are used in this application. The specific values are for illustrative purposes, and the aspects described are not limited to these specific values.

[0127] Various embodiments relate to decoding. As used in this application, "decoding" can cover, for example, all or part of the process performed on a received coded sequence in order to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding and inverse quantization. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a more general decoding process, and is considered well understood by those skilled in the art.

[0128] Various embodiments relate to encoding. In a manner similar to the above discussion regarding "decoding", as used in this application, "encoding" can cover, for example, all or part of the process performed on an input video sequence in order to produce a coded bitstream.

[0129] The various embodiments and aspects described herein can be implemented, for example, in a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single embodiment form (e.g., only as a method), the embodiments of the features discussed can also be implemented in other forms (e.g., a device or a program). The device can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0130] References to "one embodiment" or "an embodiment" or "one embodiment" or "an embodiment" and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearance of the phrase "in one embodiment" or "in an embodiment" or "in one embodiment" or "in an embodiment" and any other variations that occur throughout the specification do not necessarily all refer to the same embodiment.

[0131] Additionally, this application or its claims may relate to "determining" various pieces of information. Determining information can include one or more of the following items: for example, estimating information, calculating information, predicting information, or retrieving information from a memory.

[0132] Further, this application or its claims may relate to "accessing" various pieces of information. Accessing information can include one or more of the following items: for example, receiving information, (e.g., from a memory) retrieving information, storing information, moving information, copying information, calculating information, predicting information, or estimating information.

[0133] Additionally, the present application or its claims may relate to "receiving" various pieces of information. Like "access", receiving is intended to be a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., from a memory or an optical media storage device). Further, during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information, "receiving" is typically involved in one way or another.

[0134] It should be appreciated that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", any use of the following " / " and "and / or" and "at least one" is intended to cover only selecting the first-listed option (A), or only selecting the second-listed option (B), or selecting both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover only selecting the first-listed option (A), or only selecting the second-listed option (B), or only selecting the third-listed option (C), or only selecting the first and second-listed options (A and B), or only selecting the first and third-listed options (A and C), or only selecting the second and third-listed options (B and C), or selecting all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended so much for the listed items.

[0135] Furthermore, as used herein, the word "signal" refers to, among other things, indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a quantization matrix for dequantization. Thus, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Accordingly, for example, the encoder may transmit (explicit signaling) specific parameters to the decoder such that the decoder may use the same specific parameters. Conversely, if the decoder already has specific parameters as well as other parameters, signaling without transmission (implicit signaling) may be used to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual function, bit savings are achieved in various embodiments. It should be appreciated that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0136] It will be apparent to those skilled in the art that the embodiments can generate a variety of signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry the bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.

Claims

1. A method for training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, wherein the method includes, during training, quantizing and freezing the learned decoder parameters layer by layer in different epochs.

2. The method according to claim 1, wherein, quantizing and freezing the learned decoder parameters layer by layer in different epochs includes: quantizing and freezing the learned decoder parameters layer by layer in different epochs from the layer closest to the output of the decoder neural network to the layer closest to the input of the decoder neural network.

3. The method according to claim 1, wherein, the same quantizer is used to quantize the learned decoder parameters of all decoding layers.

4. The method according to claim 1, wherein, a specific quantizer is associated with each decoding layer of the decoder neural network and is used to quantize the learned decoder parameters of the decoding layer.

5. The method according to claim 1, wherein, the decoder neural network includes deconvolution layers, and each deconvolution layer is followed by a rectified linear unit.

6. A decoder neural network including deconvolution layers, each deconvolution layer followed by a rectified linear unit, wherein the parameters of the decoder neural network are integer parameters learned by the method according to claim 1.

7. A computer-readable storage medium having stored thereon instructions for implementing the method according to claim 1 when executed by a processor.

8. A method for training an encoder neural network and a decoder neural network to learn encoder parameters and decoder parameters, wherein the method includes, during training, quantizing and freezing the learned encoder parameters layer by layer in different epochs.

9. The method according to claim 8, wherein, quantizing and freezing the learned encoder parameters layer by layer in different epochs includes: quantizing and freezing the learned encoder parameters layer by layer in different epochs from the layer closest to the output of the encoder neural network to the layer closest to the input of the encoder neural network.

10. The method according to claim 8, wherein, the same quantizer is used to quantize the learned encoder parameters of all encoding layers.

11. The method according to claim 8, wherein, a specific quantizer is associated with each encoding layer of the encoder neural network and is used to quantize the learned encoding parameters of the encoding layer.

12. The method according to claim 8, wherein, the encoder neural network includes deconvolution layers, and each deconvolution layer is followed by a rectified linear unit.

13. An encoder neural network including deconvolution layers, each deconvolution layer followed by a rectified linear unit, wherein the parameters of the encoder neural network are integer parameters learned by the method according to claim 8.

14. A computer-readable storage medium having stored thereon instructions for implementing the method according to claim 8 when executed by a processor.