Uniform vector quantization for end-to-end image / video compression

By using neural networks in image and video compression to acquire latent variables and perform uniform vector quantization and lossless encoding, the problem of difficulty in effectively utilizing VQ technology in the prior art is solved, and lower quantization errors and higher compression efficiency are achieved.

CN119999216APending Publication Date: 2025-05-13INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069575.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-09-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing end-to-end compression technologies have not been standardized, especially in the compression process of images and videos, it is difficult to effectively utilize vector quantization (VQ) technology to achieve better rate distortion performance.

Method used

A method of using a neural network to obtain latent variables of image data and performing video encoding and decoding through uniform vector quantization and lossless encoding methods is proposed. This method uses a neural network to obtain latent variables during the encoding process, performs uniform vector quantization to obtain codewords, and performs lossless encoding of codeword indexes based on the probability distribution of latent variables. The decoding process is the opposite, the codeword index is obtained through lossless decoding, uniform vector dequantization is performed to reconstruct the latent variables, and finally decode the image data using a neural network.

Benefits of technology

Through this method, the quantization error of image and video data can be effectively reduced, while maintaining the same rate distortion performance as scalar quantization, thereby improving compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999216A_ABST
    Figure CN119999216A_ABST
Patent Text Reader

Abstract

In end-to-end compression, an image may be encoded using a deep neural network-based encoder. The embedding output from the encoder is quantized and encoded using a lossless encoder. In one embodiment, we suggest to apply uniform VQ (vector quantization) to embedding trained / optimized for uniform SQ (scalar quantization). In addition, we suggest to employ an existing model and to fine-tune a portion of the model or the entire model according to a new fixed optimal VQ graph. In this manner, the decoder becomes more stable for higher individual quantization errors. Because the encoder can be trained based on SQ and deployment is based on VQ, our method is also referred to as a hybrid SQ / VQ method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present embodiments generally relate to a method and apparatus for end-to-end neural compression of images and videos. Background Art

[0002] End-to-end compression is a compression technique where all components of the process are learned from the given data. End-to-end means that everything is learned from one end (the given data) to the other end (the compressed bitstream). Once the architecture is defined for end-to-end compression, no manual engineering effort is required to design the steps. End-to-end neural compression methods are not yet standardized. Currently, MPEG is exploring these techniques. Summary of the invention

[0003] According to one embodiment, a video encoding method is proposed, which includes: using a neural network to obtain latent variables associated with image data; performing uniform vector quantization on the latent variables to obtain corresponding codewords; and performing lossless encoding on the index of the corresponding codeword based on the probability distribution of the latent variables.

[0004] According to another embodiment, a method for video decoding is proposed, the method comprising: performing lossless decoding based on a probability distribution of latent variables representing image data to obtain a codeword index; performing uniform vector dequantization based on the codeword index to reconstruct the latent variables associated with the image data; and decoding the image data from the reconstructed latent variables using a neural network.

[0005] According to another embodiment, a device for video encoding is provided, the device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: obtain latent variables associated with image data using a neural network; perform uniform vector quantization on the latent variables to obtain corresponding codewords; and perform lossless encoding on an index of the corresponding codeword based on a probability distribution of the latent variables.

[0006] According to another embodiment, a device for video decoding is provided, the device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: perform lossless decoding based on a probability distribution of latent variables representing image data to obtain codeword indexes; perform uniform vector dequantization based on the codeword indexes to reconstruct latent variables associated with the image data; and decode the image data from the reconstructed latent variables using a neural network.

[0007] One or more embodiments also provide a computer program, the computer program including instructions, which when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any one of the embodiments described herein. One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for video encoding or decoding according to the methods described herein.

[0008] One or more embodiments also provide a computer-readable storage medium having video data generated according to the method described above stored thereon. One or more embodiments also provide a method and apparatus for transmitting or receiving video data generated according to the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 Shown is a block diagram of a system in which aspects of the present embodiments may be implemented.

[0010] Figure 2 The training loop of an end-to-end image compression model is shown.

[0011] Figure 3 The encoding process of the end-to-end image compression model is shown.

[0012] Figure 4 The decoding process of the end-to-end image compression model is shown.

[0013] Figure 5 It shows how the current end-to-end image compression model applies quantization, dequantization, encoding, and decoding using given latent variables and probabilities during the test phase.

[0014] Figure 6 A sample cell grid in 2D space and some properties are shown.

[0015] Figure 7 Shown are the quantization grids for the pseudo-2D latent variables in the default model (Uniform SQ) and our proposal (Uniform VQ).

[0016] Figure 8 It is shown how our proposal applies quantization, dequantization, encoding and decoding using given latent variables and probabilities at the testing stage according to an embodiment.

[0017] Fig. 9A Quantization error for a 2D hexagonal grid is shown, where the latent variable can be anywhere within the grid. Fig. 9B The continuous relaxation distribution of the latent variables and their hexagonal grid quantization are shown.

[0018] Fig.10A post-training loop of a deep decoder network for a VQ grid is shown, according to an embodiment.

[0019] Fig.11 A post-training loop for a complete model for a VQ grid is shown, under an embodiment.

[0020] Fig.12 The BD rate gain on the Kodak test set for different qualities is shown. DETAILED DESCRIPTION

[0021] Figure 1 A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be embodied as a device including various components described below, and is configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed over multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or other electronic devices via, for example, a communication bus or by a dedicated input and / or output port communication. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0022] The system 100 includes at least one processor 110 configured to execute instructions loaded therein to implement, for example, various aspects described in the present application. The processor 110 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include a non-volatile memory and / or a volatile memory, including but not limited to an EEPROM, a ROM, a PROM, a RAM, a DRAM, an SRAM, a flash memory, a magnetic disk drive, and / or an optical disk drive. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device.

[0023] The system 100 includes an encoder / decoder module 130, which is configured to process data, for example, to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that can be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated into the processor 110 as a combination of hardware and software known to those skilled in the art.

[0024] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0025] In several embodiments, memory inside the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be a memory 120 and / or a storage device 140, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations (such as for MPEG-2, MPEG-4, HEVC, or VVC).

[0026] Input to the elements of system 100 may be provided through various input devices, as indicated in block 105. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, such as by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0027] In various embodiments, the input device of block 105 has corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the desired frequency band again by filtering, down-conversion and perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some and / or add other elements of similar or different functions in these elements. Adding element can include inserting element between existing element, for example, inserting amplifier and analog-to-digital converter. In various embodiments, the RF part includes antenna.

[0028] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 110 as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 110 as desired. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 operating in conjunction with memory and storage elements, to process the data streams as desired for presentation on an output device.

[0029] The various elements of system 100 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data therebetween using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.

[0030] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired and / or wireless medium.

[0031] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 100 using a set-top box that passes data through an HDMI connection of the input block 105. Likewise, other embodiments provide streaming data to the system 100 using an RF connection of the input block 105.

[0032] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of embodiments, other peripheral devices 185 include one or more of independent DVRs, disk players, stereo systems, lighting systems, and other devices that provide functions based on the output of the system 100. In various embodiments, control signals are transmitted between the system 100 and the display 165, the speaker 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output device can be communicatively coupled to the system 100 via a dedicated connection through the corresponding interface 160, 170, and 180. Alternatively, the output device can be connected to the system 100 using a communication channel 190 via the communication interface 150. The display 165 and the speaker 175 can be integrated in a single unit with other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.

[0033] For example, if the RF input section 105 is part of a separate set-top box, the display 165 and speaker 175 may alternatively be separate from one or more of the other components. In various embodiments where the display 165 and speaker 175 are external components, the output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0034] Training phase

[0035] Figure 2 The training phase of a modern end-to-end compression system is shown. The input image to be compressed x∈R n×n×3 First processed by the depth encoder (210), where y = g a (x; φ), where φ is the function g to be optimized during the training phase a The output of the encoder y∈R m×m×o is called the primary embedding (or primary latent variable) of the image. Here, without loss of generality, we assume that the image is square (n×n) and the primary embedding is also square (m×m). However, they do not have to be square and can be of any shape. The input image can be fed to the encoder at the image level, or can be divided into image regions, where separate image regions are taken as input to the encoder.

[0036] y then enters the quantizer (260) as to get the main code of the image, followed by the dequantization block To obtain the reconstructed main latent variable The most advanced neural models use a hyper-prior entropy model. Specifically, the side embedding (or side latent variable) z∈R k×k×f By another deep neural network (220) through z = h a (y; Φ) to learn, where Φ is the function h a are trainable parameters that are optimized during the training phase, and the side embeddings are obtained by To quantize (270) to obtain the side code Then dequantize To obtain the reconstructed lateral latent variables These For learning The probability model of . Usually, It can be modeled by a Gaussian distribution, where the parameters are obtained by another deep network (240), such as Decompressed image Used by the depth decoder (230) Get.

[0037] The values ​​of m and k depend on the defined architecture, and m is typically 1 / 16 of the image spatial resolution, and k is typically 1 / 8 of the size of the embedding. For example, in the most common architectures, the value of o is fixed to 128. In an example, if the image size is 256x256x3, typically y will be 16x16x128, and z will be 2x2x128.

[0038] In this model, the side code The lower bound of the bit length of is calculated by the decomposition entropy model (250). This model accepts As input and learn side code The probability density function (PDF) of the code is given by the quantization method. Using this PDF, the model calculates the probability mass function (PMF) of each code under a certain quantization method. The PMF value is sufficient to calculate a lower bound on the bit length of the side code. On the other hand, the main code The lower bound of the bit length is calculated by the hyper-prior entropy model. This model is usually implemented by a Gaussian distribution. Basically, Each reconstructed main latent variable in should follow a Gaussian distribution, where the parameters (μ i , σ i ) is obtained from the side information in the previous step. Therefore, under a certain quantization method, the PMF of these Gaussians is sufficient to calculate the lower bound of the bit length of the main code.

[0039] In this setting, the depth encoder (g a (.;φ)), deep decoder (g s (.;θ)), deep hyper-prior encoder (h a (.; Φ)), deep super prior decoder (h s (.; θ)) and the decomposition entropy model p f (.; ω) consists of multiple neural layers, such as convolutional layers. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called a bias, and then applies a nonlinear function to the resulting value. The shape of the tensor (and other features) and the type of nonlinear function are called the architecture of the network. We will use the term "weights" to refer to the values ​​of tensors and biases. The parameters of weights and (if applicable) nonlinear functions are called parameters. The architecture and parameters define a "model". As mentioned above, we denote their parameters by φ, θ, Φ, Θ, and ω.

[0040] Many end-to-end architectures have been proposed recently. Figure 2 The ones shown in are more complex, but they all retain deep encoders and decoders. State-of-the-art models can rival traditional codecs in terms of rate-distortion tradeoff.

[0041] The model must be trained on a massive database D of images to learn the weights of the encoder, decoder, and entropy model. Typically, the weights are optimized to minimize the training loss:

[0042]

[0043] where d(.,.)(215) is a measure of the distortion between the original image and the reconstructed image (e.g., mean square error). The rate term (255, 245) is a lower bound on the bit length of the side information. and the lower bound of the bit length of the main information The hyperparameter λ controls the trade-off between the rate (r) and distortion (d) terms. It should be noted that the rate here is based on the lower bound of the bit length of the main information and the side information, but other methods can also be used to estimate the rate.

[0044] Deployment (testing) phase

[0045] When this model is trained and the optimal parameters (φ, θ, ω, Φ, Θ) are obtained, we can deploy the model in the encoding and decoding devices, respectively. Figure 3 and Figure 4 This phase is usually called the testing phase. During the testing phase, both the encoding and decoding devices require a PMF table. In addition to the PMF table, the encoding and decoding devices must also agree on which quantization method to use.

[0046] like Figure 3 As shown, in the encoding device, the main code and the side code (310, 320) are obtained from a given image in the same manner as in the training phase. Afterwards, the encoding device converts the side code into a bitstream through a lossless encoder, such as an arithmetic encoder AE (390) driven by a learned PMF table provided by a decomposition entropy model (350). The side code obtained from the quantization (370) is dequantized (375), and the reconstructed side latent variables are obtained in the next step. The super prior entropy decoder (340) decodes the reconstructed side latent variables to find the distribution of the main code. Finally, the main code obtained from the quantization (360) is converted into a bitstream using these distributions by a lossless encoder (such as an arithmetic encoder AE (380)). Finally, the two bitstreams are cascaded through a certain pointer in the middle.

[0047] like Figure 4 As shown, in the decoding device, the process starts with decoding the side code from the bitstream, which can be done by a lossless decoder (e.g., an arithmetic decoder AD (470)) using the learned PMF table provided by the decomposition entropy model (460). In the next step, the side code is dequantized (450) and the reconstructed side latent variables are obtained. The super prior entropy decoder (440) accepts the reconstructed side latent variables as input in order to find the distribution of the main code. These distributions tell the AD (430) how to read the main code from the bitstream. Dequantization (420) is then applied to the main code to reconstruct the main latent variables. The decoding device can obtain a reconstructed image by feeding the reconstructed main latent variables to the depth decoder (410).

[0048] Quantification

[0049] It is necessary to apply a quantization step denoted as Q(.) and its corresponding dequantization step Q -1(.) to obtain the reconstructed main latent variable and side latent variable. For simplicity, we explain the main latent variable. The same quantization and dequantization can also be applied to the side latent variable.

[0050] In the end-to-end compression literature, quantization is applied in two different ways, either through scalar quantization (SQ) or vector quantization (VQ). SQ is the embedding of each scalar (y) in a tensor (y). i,j,c ) are assigned codes respectively, and VQ firstly starts from y∈R m×m×o Construct v-dimensional vector And for each v-dimensional latent variable Assign codes. Therefore, the number of codes in VQ is v times smaller than the number of codes in SQ. However, the number of different values ​​in the codebook (dictionary length) of VQ should be higher than that of SQ, so the probability of each code in VQ will be much lower than that of the code in SQ. If the quantization boundaries of SQ are used in all dimensions of VQ, the same rate and distortion performance will occur. Quantization boundaries are the boundaries between the quantization grids (partitioned regions) of the quantization map. For example, for scalar quantization, if a number in the range [1.5, 2.5] is quantized to 2, the boundaries of this quantization sub-region (bin) are 1.5 and 2.5. In the VQ case, the data has a higher dimensionality.

[0051] However, VQ can not apply the same quantization bounds to each dimension. It can find a better quantization grid on a higher dimensional space, thus improving the rate-distortion performance.

[0052] Regardless of the quantization technique chosen, the quantization step cannot be directly used for gradient-based training because quantization has non-informative gradients (zero almost everywhere and infinite on the quantization boundary). One approach is to use a continuous relaxation of uniform SQ. In this solution, quantization is simulated by Q(x) = x + ∈, where ∈ is a random sample from uniform noise with a width of 1 sub-region around zero, such as ∈~U(-0.5, 0.5) in the training phase, where U(-0.5, 0.5) represents a uniform distribution that generates random numbers between -0.5 and 0.5. But in the testing phase, deterministic hard quantization is applied. For example, in the case of uniform SQ with a width of 1 sub-region, in the testing phase, it applies rounding to the nearest integer Q(x) = round(x). In this approach, there are no trainable parameters in the quantization step and the dequantization step is the identity function, which is Q -1 (x) = x. Since the encoder network can apply nonlinear transformations, it is expressive enough to shrink and / or expand the embedded values. Therefore, it seems reasonable to optimize the model using quantization with uniform fixed subregion width (width of 1 subregion).

[0053] In the case of VQ, a predefined number of quantization grids (Cj ∈Rv , The codewords (quantization centers) for vj=1..M) are usually learned jointly with the rest of the model parameters. In this document, we use the terms "codeword" and "quantization center" interchangeably, both referring to representative points of the quantization grid in vector quantization. During the test phase, each v-dimensional embedding should be assigned to the index of the closest quantization center as

[0054] And the quantization center coordinates of the dequantization step are returned as Q -1 (j) = C j Since the argmin operator applies hard assignment (assigning the index of the closest quantization center), it is not differentiable. During the training phase, soft assignment is used. Soft assignment assigns all M codes in the dictionary with different probabilities according to their distance Therefore, its solution is the weighted sum of all quantized centers, where the weights are calculated by probability according to their distance.

[0055] Quantitative suboptimality

[0056] Both quantization methods have some advantages and disadvantages. On the one hand, SQ is quite consistent during the training phase and has low computational complexity. But it can be shown that SQ is not the best quantizer, even though the embeddings are considered to be independent and identically distributed values ​​(iid, no correlation between codes). Even among uniform quantization, SQ is not optimal. It is theoretically known that when we use infinitely high-dimensional VQ, we can reach the theoretical limit of rate-distortion performance. On the other hand, learning VQ quantization is unstable in an end-to-end manner, and it is difficult to learn a better quantization map than SQ (a set of quantization grids that fill the space without gaps). In addition, VQ learns non-uniform quantization maps, and it is not clear whether we need non-uniform quantization maps, because the deep encoder can learn any non-linear transformation that converts non-uniform quantization maps to uniform quantization maps. Therefore, uniform quantization grids can be sufficient, and they can be learned more robustly. So far, most current end-to-end image compression methods adopt uniform SQ quantizers and do not use the theoretical advantages of VQ.

[0057] While VQ is theoretically superior to SQ, current modern end-to-end compression methods do not exploit this advantage. Optimizing models using uniform SQ is reported to be the most empirically robust solution in the literature to date. This is likely due to the difficulty of optimizing the quantization center in higher dimensions along with the rest of the model parameters, and because this is the default setting in VQ, requiring a double lookup of the non-uniform quantization grid. In this paper, we apply the theoretically powerful VQ approach to models optimized using uniform SQ.

[0058] As will be described in detail below, in a first embodiment, we exploit both SQ and VQ by proposing to apply uniform VQ to embeddings trained / optimized for uniform SQ. This embodiment does not require any retraining phase, but rather it takes any pre-trained model optimized for uniform SQ on latent variables and changes its quantization map using the theoretically known optimal VQ grid in a higher dimension. One of the drawbacks of this approach is that even if our average quantization error is lower than the default, the individual quantization errors of some latent variables may be higher than the maximum expected quantization error in SQ. Therefore, the decoder may not be robust to the quantization error of this individual latent variable. On average, the proposed method never makes the rate-distortion performance worse than the default model, but the proposal is clearly not optimal. To this end, we propose second and third embodiments that take an existing model and fine-tune part of or the entire model according to a new fixed optimal VQ map. In this way, the decoder becomes more stable to these higher individual quantization errors. Since the encoder can be trained based on SQ and the deployment is based on VQ, we also call our method a hybrid SQ / VQ method. According to our tests, we improve the BD rate of the state-of-the-art end-to-end compression algorithm by more than 2%.

[0059] Most of the latest end-to-end image / video compression models combine two types of information (i.e., primary code and side code ) is encoded into the bitstream. According to our analysis, the vast majority of the bitstream is used by the main code (about 95% to 99%), and a small portion of the bitstream is used by the side code. This is why the improvement of the main code encoding is much more important than the side code. However, our method can be applied to both types of codes.

[0060] definition

[0061] Without loss of generality, we assume that there is a m×m×o is a type of latent variable, and we have sPDFp (c) :R→R,c=1...s, is learned by the entropy model under uniform SQ with fixed subregion width. For simplicity, we can assume that the quantization subregion width is 1. We also know which latent variable in the m×m×o tensor corresponds to which PDF in s. We can use the set s to To reference this information, each set shows the index into a tensor of size m×m×o. For example, It shows that its corresponding PDF is p (1) Therefore, the total number of indices in each set must equal the total number of latent variables in y, i.e. (The total number of elements in the cth set is ). We can Define a new vector that shows the same PDF as (c) The corresponding latent variable. Each y (c) can be regarded as a subset of the embedding y, so that y = y (1) ∪y (2) ∪…∪y (s) For the side information, which latent variable uses which PDF is determined by the last index of the embedding (the feature band refers to the corresponding PDF), which shows a position out of f positions. For example, if z has 4x4x128 dimensions, then z(2, 1, 46) shows a scalar in the 2nd row, 1st column, and 46th feature band, and the last index is 46. For the main code, this information is obtained by the entropy model using the side information.

[0062] State-of-the-art methods

[0063] Currently, if Figure 5 As shown, most end-to-end compression methods use As input, and in the deployment phase, the latent variable is quantized according to the uniform SQ of 1 sub-region width (510) as Also through The PMF table is calculated (550) for the quantized center x (integer center). Later, the PMF table P can be used (c) (.) The code is passed to AE(520) is converted to a bit stream. In this way, the expected bit length will be This process should be repeated for all c=1...s. Therefore, the final expected bit length (rate) will be PMF table P( c)( .) is also used by AD (530). Since the dequantization step (540) of the nearest integer quantization is the identity function, the dequantized latent variable is the same as the code, i.e., Since quantization is not a reversible function, the given latent variable y (c) and the reconstructed latent variables There are some errors between them, which define the quantization error This error can be considered as an approximation of the error in the reconstructed image.

[0064] Proposed method

[0065] Even if the vector y (c)Given that the latent variables in are independent and identically distributed, and there is no known redundancy between the latent variables, VQ can theoretically achieve better rate-distortion performance than SQ. In our approach, we prevent VQ from learning a non-uniform quantization grid and instead stick to a uniform quantization grid. We choose a uniform quantization grid because there is doubt about the necessity of non-uniform quantization maps in end-to-end compression. We believe that the encoder network has sufficient expressiveness to produce the necessary transformations. The encoder network can expand or shrink the necessary latent variable values ​​before quantization. Therefore, quantization can be applied over a uniform quantization grid. Another reason is that in our first embodiment, we do not retrain models pre-trained for uniform quantization grids. To avoid mismatches, we retain the uniform quantization grid. Since the encoder will be able to scale the latent variables (e.g., multiply them with a scalar), it will also be meaningless to learn the quantization subregion width for a specific rate-distortion trade-off (determining λ in the loss function). Therefore, the choice of the quantization subregion width is arbitrary, and we can use choices from the literature, such as 1. Our main motivation is to reduce the quantization error of latent variables by applying VQ with an optimal uniform quantization grid while keeping the rate at the same level (or almost at the same level) as the trained scalar quantization.

[0066] We believe our approach has at least the following novelties.

[0067] Restricting VQ to apply uniform quantization maps is new. To date, in the literature, end-to-end compression models based on VQ jointly learn the quantization grid with the model parameters and all find non-uniform quantization grids which we believe are unnecessary and hurt optimization. Since the optimal uniform VQ quantization grid for a certain subregion width is theoretically known, we do not need to learn these quantization centers.

[0068] It is a new approach to reduce the quantization error of a uniform SQ based trained model by rearranging the quantization centers and PMF tables with the best uniform VQ grid. This is our first implementation and can be viewed as a post-processing of any uniform SQ based trained model.

[0069] Like our second and third embodiments, fine-tuning the SQ-based model to be able to adapt to the optimal uniform VQ quantization grid is also a new approach. These embodiments can be seen as post-fine-tuning of our first embodiment.

[0070] First Example - Optimal Uniform VQ as Post-Processing

[0071] It is known that when VQ is applied to increasing dimensions, we can approach the theoretical rate-distortion limit. Therefore, when VQ is applied to infinite dimensions, we can reach the limit. But in practice, VQ in small dimensions already provides stronger results than SQ. Due to the above advantages, the latent variables to be quantized in the end-to-end compression model are learned for a uniform quantization grid, so we need to determine the optimal center under the constraint that the quantization grid should be uniform. As described below, this constraint helps us get rid of the time-consuming iterative Lloyd algorithm or the unstable quantization center learning process jointly applied with the rest of the model parameters.

[0072] Uniform VQ is optimal among uniform quantizers in v-dimensions. If the v-dimensional space is filled without any empty areas, and there is no overlap, and the repetition of a single v-dimensional shape has the minimum surface / volume ratio (minimum energy), then the optimal VQ solution is equivalent to this uniform partitioning of the v-dimensional shape.

[0073] In 2D space, a square grid can fill the space very well. However, the energy (perimeter / area) of a square grid per unit area is 4. A circle with unit area has the minimum energy, where But it cannot fill a 2D space without gaps and overlaps. However, a hexagonal grid can fill the space very well and has an energy of approximately 3.72, which is better than a square grid. These examples can be found in Figure 6 The hexagonal grid is known as the best uniform quantization grid in 2D space. Therefore, in 2D VQ, we can tile the space with M hexagonal grids, whose centers are initially calculated as C j ∈R 2 , j=1..M.

[0074] In 3D space, truncated octahedra are considered the best candidates for uniform quantization grids. For 3D VQ, the best choice would be M truncated octahedra that tile the 3D space with their centers at C j ∈R 3 , j = 1..M. Our method is not limited to 2D and 3D optimal uniform quantization centers. In addition, it can accept any initially calculated v-dimensional quantization center C j ∈R v , j = 1..M. Since we can view SQ as a special type of VQ where it applies a square grid in all dimensions, any initially computed uniform grid (whose energy is smaller than the square grid) will improve the quantization error of the latent variable.

[0075] Figure 8The method applied in the test phase according to an embodiment is shown. The encoder can decide the value of v. For example, the encoder can take a fixed value, v = 2 or 3. The encoder can also test different values, such as v = 2, 3, 4, 5, and then use the best option and write the value of v as additional information to the bitstream. After determining v, a uniform grid at the zero reference point grid and the center C of all grids j ∈R v , j = 1..M, we need to Create a pseudo v-dimensional vector in . Since y (c) All scalars in are independent and identically distributed, and we can reshape (810)y in any order, even randomly. (c) Since the decoder device should know this order, we choose to reshape the vectors in raster order, for example, as If the length of the vector cannot be divided into v groups, we can use zero padding to fill the missing parts. Now, each column vector of the reshaped latent variable is a point in the v-dimensional space, which is And there are v-dimensional points need to be assigned corresponding codes. As an example, these pseudo 2-dimensional latent variables can be Figure 7 As can be seen in , the quantized unit square grid is the default SQ model, and the hexagonal grid is for the special case where v = 2.

[0076] In the quantization step (820), we find The index of the nearest quantized center. This step can be done by Get. Now we have the code representing the index of the quantization center

[0077] To encode these codes into the bitstream, the AE (830) requires the PMF of each different code, where the default PMF shows the probability of each integer scalar. To obtain the necessary PMF values, we first use the entropy model p (c) : R → R provides the PDF (815) to create a v-dimensional PDF (825). Since both axes follow the same PDFp (c) (.), we pass Get the PDF where u∈R v Later, we can quantize all probabilities within the grid by integrating (870) To discrete quantization center In order to have a grid Here j iterates over all quantization centers j = 1...M, where C j +grid means C jThe jth quantized grid is centered on . In the 2D case, it is a hexagonal grid. This integral can be calculated analytically. However, for more complex grids, Monte Carlo simulations can be useful. We can generate millions of v-dimensional data points, where the dimensions are independent and identically distributed and each follows p (c) (.). We can then count how many points are in each grid and get the frequency of each grid. The normalized frequency will give a fairly good approximation of the PMF for the grid. AE can be done by using To code

[0078] Also like Figure 8 As shown, the decoder device also needs to first calculate the bit stream from the arithmetic decoder AD (840) To get the code This step should be followed by a dequantization step (850), which returns the value of the reconstruction code The corresponding quantization grid shown is the quantization center. Then, we need to reshape (860) by for to cancel the reshaping process at the beginning of encoding.

[0079] Next, we illustrate the step-by-step process in three steps, first the steps that should be performed offline (before deployment), then the steps that should be completed after deployment in the encoding device, and finally the steps that should be completed after deployment in the decoding device.

[0080] Steps to complete before deployment:

[0081] Step 1. Define v, grid and optimal uniform quantization center C j ∈R v , j = 1..M. As mentioned above, in our embodiment, the quantization center for uniform VQ is known (eg, for v = 2 or 3, the quantization center of a grid is the geometric center of the corresponding grid).

[0082] Step 2. For all c=1...s

[0083] By using the default PDF (c) (.) to calculate the quantitative center PMF.

[0084] Step 3. Share v, grid, C with both encoding and decoding devices j ∈R v , j = 1..M and

[0085] Steps to be completed in the encoding device:

[0086] Step 1. Input image x∈R n×n×3 .

[0087] Step 2. Obtain embedding y∈R through a deep encoder network m×m×o .

[0088] Step 3. Get the PMF corresponding set ( The side information is fixed, and the main information can be obtained from the side information by following steps 3.1 to 3.5).

[0089] Step 3.1 Pass z = h through the deep encoder network a (y;Φ) to get the embedding z∈R k×k×f .

[0090] Step 3.2 Obtain the reconstructed lateral latent variable code

[0091] Step 3.3 Obtain the distribution parameters of the main latent variables where μ,σ∈R m×m×o .

[0092] Step 3.4: Each scalar y ijk Its corresponding scalar σ ijk into s groups. If σ (b-1) ≤σ ijk <σ (b) , y ijk Classify into group b. There are s+1 predefined classes [0, σ (1) , ..., σ (s) ], and each y ijk Assigned to one of the s groups.

[0093] Step 3.5 Create a collection This set stores the y assigned to the cth group. ijk The index of .

[0094] Step 4. Divide the latent variables into subsets with respect to their corresponding PMF correspondence

[0095] Step 5. For all c=1...s

[0096] Step 5.1 Reshape

[0097] Step 5.2: pass To apply VQ;

[0098] Step 5.3 Used by AE Will Encoded into the bitstream.

[0099] Step 6. Return the bitstream.

[0100] Steps to be completed in the decoding device:

[0101] Step 1. Input bitstream.

[0102] Step 2. For all c=1...s

[0103] Step 2.1 Used by AD Read code from bitstream

[0104] Step 2.2 Pass To dequantize the code;

[0105] Step 2.3 Reshape

[0106] Step 3. Get the PMF corresponding set ( For side information, it is fixed, and for main information, it can be obtained from the side information).

[0107] Step 4. Incorporate latent variables A subset of

[0108] Step 5. Obtain the reconstructed image through the deep decoder network

[0109] Post-training the new quantized grid

[0110] One of the disadvantages of the first embodiment is that the single maximum quantization error is higher in VQ than in SQ. Figure 6 This problem is shown for the 2D case. Even though the average quantization error in the hexagonal grid (best uniform VQ in 2D) is smaller than SQ, the maximum absolute quantization error in one dimension can be larger than the SQ case. Figure 6 As shown, for a square grid (SQ), the maximum distance from a single latent variable (in one dimension) to the quantization center is 0.5. However, in a hexagonal grid (optimal uniform VQ), this distance can be approximated to 0.62. During the training phase, quantization is simulated as ∈~U(-0.5, 0.5) by adding uniform noise sampled between -0.5 and 0.5. This means that after training, the deep decoder network will be robust enough if any latent variable differs in the range of [-0.5, 0.5] after quantization, which is the expected quantization error in SQ. In the following, we describe how we can simulate the quantization error of a new grid during the training phase.

[0111] Since the v-dimensional grid i Any v-dimensional point within should be quantized at the center of the grid, and we do not have any prior information where the upcoming latent variable point is likely to appear, so on average the best prior is to assume that the latent variable can appear in the grid with equal probability (uniform distribution). i Any location within. Fig. 9A This point is shown for the case of the best 2D uniform VQ (hexagonal grid). Therefore, the quantization error should follow the distribution U grid , where the probability is 1 within the grid centered at the reference point and 0 otherwise.

[0112] Definition: U grid is a uniform distribution in v-dimensional space, defined by:

[0113]

[0114] where ∫U grid (x)dx=1、x∈R v , and grid is a v-dimensional grid centered at zero.

[0115] We can see that for scalar quantization with a subregion width of 1 (e.g., nearest integer rounding), U grid = U(-0.5, 0.5). We can think of the scalar 1 subregion width quantization in v-dimensional space as applying a single v-dimensional cube grid. In this case, U grid =U(-0.5, 0.5) v For the case of 2D VQ, since the optimal grid is a unit volume hexagonal grid, the grid Any random sample of will be located anywhere within the unit volume hexagonal grid with equal probability. This yields a continuous relaxation of VQ that can be written as , where grid is a fixed uniform grid in v-dimensional space.

[0116]

[0117] This quantization relaxation does not assign grid indices to the codes, but simply assigns some uniform noise to the latent variables. Therefore, during training, we do not need to perform hard assignments on the codes, and its dequantization step is the identity function, since it is calculated by Same as SQ. Fig. 9B The continuous relaxation of the hexagonal grid quantization is shown.

[0118] Second Example - Deep Decoder Post-Training for VQ Grid

[0119] In order to adapt the deep decoder network to the new quantization grid, we propose to fine-tune the deep decoder by continuously relaxing these new quantization grids. Since the decoder network is independent of entropy (bit length), we can optimize the decoder network only for the distortion loss. To do this, we first encode the images in the dataset and create a new dataset consisting of images x∈R n×n×3 , their embedding x∈R m×m×o , PMF corresponding set and the quantization error U of the new grid grid This can be done without the PMF corresponding set. But in this way, all scalars in the pseudo v-dimensional latent variable can follow different distributions. We found it better to create a v-dimensional latent variable with latent variables from the same probability. However, it can also be done by skipping the partitioning of the latent variable into subsets and performing the reshape operation on the entire latent variable.

[0120] The forward pass computation steps of the training loop can be shown as follows, also as Fig.10 We can also apply random permutations to the latent variables ( Fig.10 ), so that the position of the latent variable becomes unimportant, thus creating a pseudo v-dimensional vector in step 3.1. If this is the case, we should retrieve this random permutation in step 3.6.

[0121] Forward pass step for deep decoder post-training:

[0122] Step 1. Input image x∈R n×n×3 , embedding y∈R m×m×o , PMF corresponding set and distribution U grid .

[0123] Step 2. Divide the latent variables into subsets with respect to their corresponding PMF correspondence

[0124] Step 3. For all c=1...s

[0125] Step 3.1 (Optional) Randomly permute y (c) The sequential latent variable.

[0126] Step 3.2 Reshape

[0127] Step 3.3 Apply quantification Continuous relaxation (1030).

[0128] Step 3.4 Apply virtual dequantization

[0129] Step 3.5 Reshape

[0130] Step 3.6 (if step 3.1) Retrieve Random arrangement of .

[0131] Step 4. Incorporate latent variables A subset of

[0132] Step 5. Obtain the reconstructed image through the deep decoder network

[0133] Starting from the pre-trained parameters, we can fine-tune the parameters of the deep decoder network by minimizing the following loss:

[0134]

[0135] The fine-tuned deep decoder network will give better performance than the first embodiment, than the first embodiment alone, and than the default model.

[0136] Third Example - Full Model Post-Training of VQ Grid

[0137] Even though the decoder adapts to the new quantization trellis in the second embodiment, the bit length (rate) is still optimized for uniform SQ.In this third embodiment, we propose to jointly optimize everything for the new VQ trellis.

[0138] For side information, which latent variable belongs to which PDF is predefined and fixed. Therefore, we can get the same PDF by using (c) (.) to create a pseudo v-dimensional latent variable. In this way, the v-dimensional PDF can be obtained simply by multiplying the same distribution as the 2D PDF: However, in the main information, each latent variable follows its own distribution during the training phase. They are grouped into s distributions during the testing phase. Therefore, each pseudo v-dimensional latent variable should have its own distribution. However, during the training phase, we do not need to create pseudo v-dimensional latent variables with the same distribution for all dimensions. Basically, latent variables in the same v-dimensional vector can follow different distributions. If we know the distribution of each latent variable, we can create a pseudo v-dimensional PDF by simply multiplying the PDF of each dimension.

[0139] In our design, we transform the image y∈R m×m×o The embedding of is reshaped into a pseudo v-dimensional embedding Start with , and apply the continuous relaxation of the newly designed v-dimensional grid as Then dequantize Dequantization code is treated by the entropy model in the same way as the default model, and its The PDF of each single scalar is found in . These PDFs are Gaussian functions whose parameters are calculated by the entropy model and contain side information about the main code. For the side information, these PDFs are the outputs of the neural network. Without loss of generality, we can get To represent each scalar Later, we can use To create a pseudo v-dimensional vector The PDF of v represents any point in v-dimensional space. We can It is called v-dimensional PDF. In order to obtain We need to calculate the PMF at the point The PDF is integrated over the grid on . This can be obtained by Again, u∈R v represents any point in v-dimensional space. Completing these calculations, we will therefore have Probability value of quantity Since we cannot perform Monte Carlo simulations during the training phase, the above integral should be calculated analytically. Therefore, if there is no analytical solution for the integral within the grid, some approximation can be used (the simpler version of the grid has an analytical solution for the integral). This can be done by taking the sum of the negative log-likelihoods (via ) is used as the rate term in the loss function of the main information to find the expected bit length.

[0140] Fig.11 A post-training loop for a complete model for a VQ grid is shown, under an embodiment.

[0141] Forward pass step for full model post-training:

[0142] Step 1. Input image x∈R n×n×3 , unit volume v-dimensional grid grid and distribution U grid .

[0143] Step 2. Obtain embedding y∈R through a deep encoder network m×m×o (1110).

[0144] Step 3. Reshape the embedding into a pseudo v-dimensional embedding

[0145] Step 4. Apply continuous relaxation of the grid as:

[0146]

[0147] Step 5. Obtain (1165) quantitative latent variables from the entropy model Each scalar PDF(1160).

[0148] Step 6. Pass Compute the pseudo v-dimensional PDF (1175) (1170) where u∈R v .

[0149] Step 7. Pass Calculate the grid PMF(1185) at point.

[0150] Step 8. Use negative log-likelihood The rate term (1180) is calculated by summing

[0151] Step 9. Reshape Back

[0152] Step 10. Use the Deep Decoder Network A reconstructed image is acquired (1150).

[0153] Step 11. Pass The distortion term is calculated (1155).

[0154] The parameters of the full model can be fine-tuned by starting from the pre-trained parameters as the initial parameters, which minimizes the following loss by using the trade-off parameter λ in the default model.

[0155]

[0156] The fine-tuned full model will give better performance than the first embodiment, than the first embodiment alone and than the default model.

[0157] result

[0158] We used a pre-trained model from the CompressAI library, named mbt2018_mean, proposed by D. Minnen et al. in “Joint Autoregressive and Hierarchical Priors for Learning Image Compression”, 2018, pp. 10794–10803. We implemented our first implementation for 2D (hexagonal meshes) and 3D (truncated octahedral meshes) and Fig.12Results on the Kodak test dataset for different models optimized for different rate-distortion tradeoffs are shown in . Our gains in quality are consistent. According to the results, the hexagonal grid can improve the performance of lower quality by more than 1%, and on average by 0.95%. Using 3D VQ, we can improve the lower quality results by more than 1.5%, and in general by more than 1%. These results are consistent with the theoretical background that VQ is better than SQ. The second embodiment improves the average gain by 1.5% over the first embodiment, and the third embodiment obtains a 2% BD rate gain over the first embodiment, which is significant in the compression literature.

[0159] It should be noted that our approach is not limited to a specific neural network architecture, e.g. Figure 5 The architecture shown in is known as a hyper-prior model. In contrast, our approach can be used for other neural network architectures, such as fully factorized neural image / video models, implicit neural image / video compression models, neural image / video compression models based on recurrent networks, or image / video compression methods based on generative models.

[0160] Various numerical values ​​are used in this application. The specific values ​​are only for illustrative purposes, and the described aspects are not limited to these specific values.

[0161] Various methods are described herein, and each of these methods includes one or more steps or actions of implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, in various embodiments, terms such as "first", "second" can be used to modify elements, parts, steps, operations, etc., such as, for example, "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply that the modified operation is sorted. Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can occur in a time period overlapping with the second decoding, such as before, during, or with the second decoding.

[0162] Various embodiments relate to decoding. "Decoding" as used in this application may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, and inverse transform. Based on the context of the specific description, whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear and will be well understood by those skilled in the art.

[0163] Various embodiments relate to encoding.In a similar manner to the above discussion about "decoding", "encoding" as used in this application may encompass all or part of a process performed on an input video sequence, for example, to produce an encoded bitstream.

[0164] The embodiments and aspects described herein may be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if only discussed in the context of a single implementation form (e.g., discussed only as a method), the embodiments of the features discussed may also be implemented in other forms (e.g., a device or program). The device may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in, for example, a device, for example, a processor that generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0165] Reference to "one embodiment" or "an embodiment" or "one implementation" or "implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0166] Additionally, the present application may refer to “determining” various information. Determining information may include one or more of the following: for example, estimating information, calculating information, predicting information, or retrieving information from a memory.

[0167] Furthermore, the present application may refer to "accessing" various information. Accessing information may include one or more of: for example, receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0168] Additionally, the present application may involve "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of: for example, accessing information or retrieving information (e.g., retrieving from a memory). Furthermore, "receiving" is often referred to in one way or another during operation, e.g., storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0169] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any of the following " / ", "and / or", and "at least one" is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many listed items as are apparent to those of ordinary skill in this and related arts.

[0170] As will be apparent to one of ordinary skill in the art, embodiments may produce signals formatted to carry a variety of signals that may be, for example, stored or transmitted. Information may include, for example, instructions for executing a method, or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

Claims

1. A video encoding method, comprising: Using neural networks to obtain latent variables associated with image data; Performing uniform vector quantization on the latent variable to obtain a corresponding codeword; as well as The index of the corresponding codeword is losslessly encoded based on the probability distribution of the latent variable.

2. A device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: Using neural networks to obtain latent variables associated with image data; Performing uniform vector quantization on the latent variable to obtain a corresponding codeword; and The index of the corresponding codeword is losslessly encoded based on the probability distribution of the latent variable.

3. The method of claim 1 or the apparatus of claim 2, wherein a codeword used to represent a partitioned area in the uniform vector quantization is fixed at a center of the partitioned area.

4. A method as claimed in claim 1 or 3 or an apparatus as claimed in claim 2 or 3, wherein the dimensions of the partitioned regions are encoded.

5. The method of any one of claims 1, 3 and 4 or the apparatus of any one of claims 2 to 4, wherein the latent variable corresponds to main information or side information.

6. The method of any one of claims 1 and 3 to 5 or the apparatus of any one of claims 2 to 5, wherein the latent variables are reshaped prior to the uniform vector quantization.

7. The method of any one of claims 1 and 3 to 6 or the apparatus of any one of claims 2 to 6, wherein the parameters of the neural network are trained based on uniform scalar quantization.

8. The method of claim 7 or the apparatus of claim 7, wherein a continuous relaxation of quantization is applied during the uniform scalar quantization.

9. The method of claim 8 or the apparatus of claim 8, wherein a rate term is estimated in the training.

10. A method for video decoding, comprising: performing lossless decoding based on a probability distribution of a latent variable representing the image data to obtain a codeword index; performing uniform vector dequantization based on the codeword index to reconstruct latent variables associated with the image data; as well as The image data is decoded from the reconstructed latent variables using a neural network.

11. A device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: performing lossless decoding based on a probability distribution of a latent variable representing the image data to obtain a codeword index; performing uniform vector dequantization based on the codeword index to reconstruct latent variables associated with the image data; and The image data is decoded from the reconstructed latent variables using a neural network.

12. The method of claim 10 or the apparatus of claim 11, wherein a codeword used to reconstruct a partitioned region in the uniform vector dequantization is fixed at a center of the partitioned region.

13. A method as claimed in claim 10 or 12 or an apparatus as claimed in claim 11 or 12, wherein the dimensions of the partitioned regions are decoded.

14. The method of any one of claims 10, 12 and 13 or the apparatus of any one of claims 11 to 13, wherein the latent variable corresponds to main information or side information.

15. The method of any one of claims 10 and 12 to 14 or the apparatus of any one of claims 11 to 14, wherein the dequantized codewords from the uniform vector dequantization are reshaped to obtain the reconstructed latent variables.

16. A signal comprising video data formed by performing the method as claimed in any one of claims 1 and 3 to 9.

17. A computer-readable storage medium having stored thereon instructions for encoding or decoding a video according to the method of any one of claims 1, 3 to 10, and 12 to 15.