Signal coding using generative models and latent domain quantization

The proposed coding scheme reallocates bits to quantized latent frames in generative models, addressing inefficiencies in audio signal modeling by enhancing coding quality and enabling low-latency transmission.

JP7868040B2Active Publication Date: 2026-06-01DOLBY INTERNATIONAL AB

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY INTERNATIONAL AB
Filing Date
2021-10-11
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

Generative models struggle to efficiently model audio signals due to complexity constraints, limited training data, and algorithm limitations, leading to modeling mismatches and inefficiencies in bit allocation for conditioning information transfer.

Method used

A coding scheme that allocates bits to transfer quantized latent frames instead of entire bit budget for conditioning information, using reversible mappings and generative neural networks to reconstruct audio signals, enabling flexible rate-distortion trade-offs and packet loss concealment.

Benefits of technology

Facilitates efficient coding by allocating bits to quantized latent frames, improving coding quality and enabling low-latency audio transmission by filling in gaps left by generative models, and allowing backward compatibility with legacy codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007868040000007
    Figure 0007868040000007
  • Figure 0007868040000008
    Figure 0007868040000008
  • Figure 0007868040000009
    Figure 0007868040000009
Patent Text Reader

Abstract

The present disclosure provides a decoder configured to receive a finite bitrate stream including quantized latent frames, the quantized latent frames including a quantized representation of a current frame of a signal in a latent domain different from a first domain; generate a reconstructed latent frame from the quantized latent frame; use a generative neural network model to perform a task for which the generative neural network model is trained, the task including generating parameters for a reversible mapping from a latent domain to a first domain; reconstruct the current frame of the signal in the first domain, the reconstructed latent frame including mapping the reconstructed latent frame to the first domain using the reversible mapping; and update a state of the generative neural network model using the reconstructed current frame of the signal in the first domain.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-references to related applications This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 092,642 (reference number: D20048USP1) filed on 16 October 2020 and European Application No. 21151296.7 (reference number: D20048EP) filed on 13 January 2021.

[0002] Technical field This disclosure relates to the field of signal coding. In particular, this disclosure relates to signal coding using generative models and latent domain quantization. [Background technology]

[0003] Generative models implemented using (deep) neural networks have demonstrated their usefulness in signal synthesis tasks. An example of such a task is audio coding, which includes decoding, for example, in which an audio signal is reconstructed based on its finite bitrate representation provided by a corresponding encoder. In such coding tasks, a generative model can perform signal reconstruction by implementing a conditional probability distribution function intended to represent the audio signal and then reconstructing the signal according to this distribution. The probability distribution function may be conditioned on one or more previously reconstructed frames of the audio signal. Furthermore, additional conditioning information is often provided in the finite bitrate representation and is typically updated periodically (e.g., once per frame of the signal) to reflect the variability of the signal.

[0004] However, audio signals can be difficult to model, and in at least some situations, it may not be possible to train an efficient generative model, even when additional conditioning information is provided (e.g., due to complexity constraints, the availability of training data, and / or limitations of the specific algorithms used by the generative model). In some situations, the generative model only approximates the truly unknown model, which can lead to a modeling mismatch.

[0005] Therefore, in light of the above, an improved generative model-based coding method is needed. [Overview of the project] [Problems that the invention aims to solve]

[0006] Therefore, the purpose of this disclosure is to at least partially satisfy the identified needs described above. [Means for solving the problem]

[0007] According to a first aspect of this disclosure, a decoder is provided. The decoder is configured to receive a finite bitrate stream containing quantized latent frames. The quantized latent frames contain a quantized representation of the current frame of the signal in a latent region. The latent region is distinct from a first region. The decoder is configured to generate reconstructed latent frames from the quantized latent frames. The decoder is configured to use a generative neural network model (hereinafter "the Model") to perform a task for which the Model is trained. This task includes generating parameters for a reversible mapping from the latent region to the first region. The decoder is configured to reconstruct the current frame of the signal in the first region, which includes mapping the reconstructed latent frames to the first region using a reversible mapping. The decoder is further configured to use the reconstructed current frame of the signal in the first region to update the state of the generative neural network model (so that the Model is ready to process the next future frame of the signal).

[0008] According to a second aspect of this disclosure, an encoder is provided. The encoder is configured to receive the current frame of a signal in a first region. The encoder is configured to use a (generative neural network) model to perform a task for which the model is trained. This task includes providing parameters for a reversible mapping from the first region to a latent region. The latent region is distinct from the first region. The encoder is configured to generate a latent frame by using the reversible mapping to map at least a portion of the current frame of the signal to the latent region. The latent frame includes a representation of the current frame of the signal in the latent region. The encoder is configured to generate a quantized latent frame based on the generated latent frame. The encoder is further configured to generate a finite bitrate stream containing the quantized latent frame.

[0009] A third aspect of this disclosure provides a method for decoding the current frame of a signal. This method includes steps performed by the decoder described above in accordance with the first aspect.

[0010] A fourth aspect of this disclosure provides a method for encoding the current frame of a signal. This method includes steps performed by the encoder described above in accordance with the second aspect.

[0011] According to the fifth and sixth aspects of this disclosure, each non-temporary computer-readable medium is provided. Each medium stores instructions that, when executed by at least one computer processor belonging to the computer hardware, cause the computer hardware to perform a method of encoding and / or decoding the current frame of a signal as described above according to the third and fourth aspects, respectively.

[0012] According to a seventh aspect of this disclosure, a coding system for transferring the current frame of a signal is provided. This coding system includes at least one encoder as described in accordance with the first aspect and at least one decoder as described in accordance with the second aspect. This coding system further includes means for transferring a finite bitrate stream containing quantized latent frames between the encoder and the decoder.

[0013] Audio signals can be difficult to model, and in at least some situations, it may be impossible to train an efficient generative model (for example, due to complexity constraints, the availability of training data, and / or limitations of the specific algorithms used by the generative model). In some situations, the generative model only approximates the truly unknown model, which can lead to a modeling mismatch.

[0014] In many current technologies that use such generative models for coding, the coding scheme employed further involves spending the entire bit budget on the transfer of conditioning information between the encoder and decoder. This disclosure improves upon the current technology by providing a coding scheme in which, instead, one or more bits in the bitstream transferred between the encoder and decoder are allocated to convey a quantized latent frame. This quantized latent frame may allow coding of one or more aspects of the signal not otherwise described by the generative model. This may help achieve rate-distortion scalability of coding schemes using generative models by facilitating flexible rate-distortion trade-offs and by facilitating coding in the latent region. As will be described in more detail later in this paper, this disclosure may also provide, for example, a generative model trained for 0-bit rate conditioning, resulting in a coding scheme in which the entire bit budget is instead allocated to a quantized latent frame. For example, if the latent variable is transmitted in a packet, such a coding scheme may facilitate packet loss concealment, since the model can be trained to replace such lost packets with synthetically reconstructed latent variables. Furthermore, encoders that implement generative models trained for zero-bitrate conditioning may not require the additional coding delays that might otherwise be needed to estimate the conditioning variables. Such coding schemes can be useful in low-latency coding, such as low-latency transmission of speech or audio.

[0015] This disclosure relates to all possible combinations of the features described in the claims. The purposes and features described according to the first aspect may be combined with, or replaced by, the purposes and features described according to the second aspect, the third aspect, and / or the fourth aspect, and vice versa.

[0016] Further objectives and advantages of various embodiments of this disclosure are described below by exemplary embodiments. [Brief explanation of the drawing]

[0017] Exemplary embodiments are described below with reference to the attached drawings.

[0018] [Figure 1a] This is one of the figures illustrating schematic embodiments of the encoder according to this disclosure. [Figure 1b] This is one of the figures illustrating schematic embodiments of the encoder according to this disclosure. [Figure 1c] This is one of the figures illustrating schematic embodiments of the encoder according to this disclosure. [Figure 1d] This is one of the figures illustrating schematic embodiments of the encoder according to this disclosure. [Figure 1e] This is one of the figures illustrating schematic embodiments of the encoder according to this disclosure.

[0019] [Figure 2a] This is one of the figures illustrating schematic various embodiments of the decoder according to this disclosure. [Figure 2b] This is one of the figures illustrating schematic various embodiments of the decoder according to this disclosure.

[0020] [Figure 3a] This diagram schematically illustrates the flow of various embodiments of the encoding method described herein. [Figure 3b] This diagram schematically illustrates the flow of various embodiments of the decoding method described herein.

[0021] [Figure 4] An embodiment of the coding system described herein is schematically shown.

[0022] In the drawings, similar elements are given the same reference numerals unless otherwise specified. Unless the opposite is explicitly stated, the drawings show only the elements sufficient to illustrate an exemplary embodiment, while other elements may be omitted or merely suggested for clarity. [Modes for carrying out the invention]

[0023] The elements of the coding scheme envisioned in this disclosure are described in more detail here.

[0024] As the first example, x t x is an N-dimensional vector representing the signal sample frame at time t, and such a vector x t It may be assumed that is the realization of some unknown random variable X. Here, N may be equal to, for example, the number of samples / number of channels / number of bandwidths, the length of the signal transformation, or, for example, the number of audio channels sampled in the time domain. The probability model implemented by the neural network is the probability density function

number

[0025] Generally, the generative model may also be conditioned on some conditioning information about the current and / or one or more future frames, in which case the conditional probability density function

Number

[0026] In the following, for the sake of simplicity of notation, such additional conditioning information will be omitted from the description unless the contrary is explicitly stated.

[0027] The above formulation means that consecutive (vector) samples are conditionally independent. However, the components of each N-dimensional vector may be correlated (within that vector). The conditional probability distribution at time t can be, for example,

Number

[0028] A practical challenge may relate to the fact that neural networks must provide a good estimate of the (positive semidefinite) covariance matrix. One conceivable example of how to achieve this is by using the Cholesky decomposition of the covariance matrix (which exists only for positive semidefinite matrices). In such a case,

number

[0029] It should also be noted that other factorization methods may be used instead of Cholesky decomposition. For example, factorization Σ=LDL T (D is a diagonal matrix), or factorization Σ=UDU T It has also been conceivable to use (where U is orthonormal and D is diagonal). The constraint of orthonormality in such factorization may be further enforced by decomposing U into a series of prime Givens rotations or Hausholder transformations.

[0030] Furthermore, it is conceivable that a complete lower triangular matrix L is not actually required. For example, by restricting the number of diagonals of the triangular matrix L, it may be possible to limit the number of parameters output by the model. In some embodiments, it is conceivable that the matrix L may be diagonal (i.e., contain non-finite elements only on the principal diagonal).

[0031] The equation provided by equation (3) represents a Gaussian mixture model, but the concept provided by this disclosure is that the output stage includes multiple reversible mappings, and each such mapping is a parameter w estimated by the generative model. j It should be noted that this can be further generalized to the associated scenarios. This parameter may then be used to calculate the latent vector (frame). For example, in the case of a Gaussian mixture model, w j w may also represent the component probability (weight of the component), and a single model component has probabilities w0, w1, ..., w J They may be randomly selected according to. In other embodiments, w j While it may still represent component probabilities, it is conceivable to use, for example, the maximum likelihood principle instead for component selection. In general, it is conceivable that the generative model is trained to provide the parameters required for either of the above factorizations / decompositions of Σ.

[0032] Here, various examples of embodiments of encoders, for example, in accordance with this disclosure will be described in more detail with reference to Figures 1a to 1e.

[0033] Figure 1a schematically shows the encoder 100, and the generative model 110 (which is implemented using, for example, a deep neural network) performs a reversible mapping 161 (G -1 It is trained to perform a task that provides parameter 120 for ).

[0034] Encoder 100 also displays the current frame of the signal 160(x t ) also receives. As mentioned above in this paper, this may contain single or multiple (vector) samples of the signal. Current frame x tThis is a representation of the signal in a first domain, which may be, for example, a time / signal domain (e.g., relating to pulse code modulation PCM), a filter bank domain (e.g., relating to one or more quadrature mirror filters QMF), a transformation domain (e.g., relating to one or more modified discrete cosine transforms MDCT), or even a processed transformation domain (e.g., relating to one or more flattened MDCTs). Other domain representations are also conceivable.

[0035] Reversible Mapping G -1 161 is that encoder 100 uses this mapping to determine the current frame x of the signal. t At least a portion of it is mapped from the first region to the latent region, thereby creating latent frame 162(Z t ) may be generated. This latent frame Z t This includes the representation of the current frame t of the signal in the latent region. Using the quantization operation 163(Q), encoder 100 calculates the latent frame Z t Quantized latent frame 164

number

number

[0036] Figure 1b schematically shows a slightly different encoder 101. Here, at least the generative model 110, parameter 120, and reversible mapping G are shown. -1 161. The quantization operation Q 163 may also be used similarly for the encoder 100 described with reference to Figure 1a, in all cases for the signal frame x t Mapping from the first domain to the latent domain, latent frame Z t Quantize 162 to obtain the quantized latent frame^Z tThis helps to output 164 in a finite bitrate stream 130. However, in addition to this, the encoder 101 also has a local decoder unit 140 which is a quantized latent frame ^Z t Take 164 (for example, provided in bitstream 130) as input and perform a dequantization operation 166 (Q -1 ) reconstructed latent frame 167(Z t * ) is generated, and then the reconstructed signal frame 169(x in the first region) is generated. t * To generate ), reversible mapping G -1 Apply the inverse 168(G) to this reconstructed signal frame x t * This is then fed back to the generative model 110 and used to update the state of the generative model so that it is ready to process the next frame t+1. As described earlier, in some embodiments, the generative model 110 receives (additional) conditioning information 165(θ) for either or both the current frame t and (one or more) future frames > t. ≧t ) may be received.

[0037] Figure 1c schematically shows encoder 102, including a local decoder 140, similar to encoder 101 described with reference to Figure 1b. Here, reversible mapping G -1 This is the input signal frame x t The signal frame average predicted from 172 (μ) t Subtract ) and then perform the reversible transformation 170(F -1 This is an affine mapping that includes the following: ). A reversible transformation 170 is, for example, a matrix (single or multiple) L j The parameters used to define a Σ, such as the scaling according to at least a subset 273 of the parameters 220 provided by Model 210, may include a reversible mapping G -1 It is not necessarily linear.

[0038] After affine mapping, latent frame Z t It is quantized by operation Q, and the quantized latent frame ^Z t This generates a finite bitrate stream 130 which is then provided to a decoder, for example. As previously mentioned, the encoder 102 includes a local decoder 140 which reconstructs the current signal frame x t * It outputs this, which is fed back and used to update the state of Model 110. In the local decoder 140, reversible mapping G -1 The inverse G is the reversible transform F -1 It is implemented using the inverse 171(F).

[0039] As previously described in this paper, in some embodiments, the encoder described herein may have multiple different reversible mappings available for selection. For example, when using a Gaussian mixture model, the model may have multiple sets of parameters {w j ,μ j ,L j} may be generated, and the exact mapping to use is probabilistic w j They may be selected according to the following. In other embodiments where the mixed model is not necessarily Gaussian, w j It may still be assumed that represents the mapping probability, and the exact mapping to be used may be selected, for example, using the maximum likelihood principle.

[0040] As shown in at least Figures 1a, 1b, and 1c, Model 110 also, in some embodiments, associates additional conditioning information θ with at least one of the current frame t and future frame > t. ≧t The model may accept this, and the model will perform that task (i.e., reversible mapping G based on such conditioning information) -1The encoder may be trained to perform actions including providing parameters for the condition. Using one or more "look-ahead" frames of conditional information can further improve coding quality, for example, at the cost of increased latency introduced in the coding / decoding process. Generally, when additional conditional information θ is used, the encoder described in this paper may also be configured to output a finite bitrate stream containing the conditional information, for example, which may be provided to a decoder. Conditional information θ provided to generative model 110 ≧t The signal may or may not be quantized, depending on the exact content of the conditioning information. Such conditioning information may include, for example, a parametric description of a signal; a waveform approximation to a coded signal encoded at a finite bitrate; and / or an indication of a signal category (e.g., speech, general audio, music, certain types of music, e.g., piano music).

[0041] In some embodiments, the encoders described herein may be further configured to output the same finite bitrate stream containing the quantized latent frames, and conditional information associated with an indication that a) the current frame and at least one of the future frames (as described above), or b) that such conditional information is not contained in the same finite bitrate stream. This indication can help, for example, a decoder to understand whether the current frame contains conditional information or whether only the quantized latent information is provided by the encoder.

[0042] Using such a two-stage signal description, it is conceivable that, for example, legacy codec data may be provided instead of explicitly providing conditioning information in the first stage of the same finite bitrate stream. In that case, the decoder described herein may generate the conditioning information itself, for example, by reconstructing the signal from the legacy codec data, and then use such reconstructed signal as at least part of the conditioning information. This may allow for the construction of a bitstream that can be decoded by a legacy decoder (for example, by ignoring the quantized latent from the bitstream). This may allow, for example, a decoder following this disclosure, i.e., the decoder described herein, to operate in a backward-compatible manner and benefit from legacy data. The encoder may be conditioned with the same legacy codec data to generate the quantized latent frame provided in the second stage of the same finite bitrate stream. Additionally, or instead of such waveform reconstruction, the legacy codec data may include other parameters for conditioning.

[0043] Figure 1d schematically shows another embodiment of encoder 103, where the quantized latent frame ^Z t A perceptual rate allocation of 150 is used to determine the bit allocation for the following. During encoding, the parameters of the lossless mapping (e.g., μ) are used. t For example, F -1 and G -1 The reversible L or L used j ) is available to encoder 103 and can be used to determine rate assignment. For example, μ tIt is conceivable to use L to estimate the variance of the signal after subtracting . These variances may then be quantized in 3dB steps. A set of quantizers is then designed to provide a 1.5dB improvement in the signal-to-noise ratio (SNR), followed by an exchange of the 3dB increment of the variance for a 1.5dB improvement in the SNR. In other embodiments, the box indicated by 150 can instead accommodate different types of rate assignments. An example of a conceivable case is, for example, the received current frame of the signal in the first region, i.e., x t And the reconstructed current frame of the signal in the first region, i.e., x t * This involves assigning a bitrate to a finite bitrate stream 130 based on the difference between the two. This may include, for example, the use of a sample distortion index such as (perceptually) weighted squared error. In other embodiments, the rate assignment may be based on output parameters 120 provided by, for example, a generative model 110. For example, in the case of affine mapping, the scale (defined, for example, by L and Σ) is such that the quantized latent frame ^Z t It can be used to distribute the bitrate for that purpose.

[0044] As another example, in some embodiments, to achieve in-frame bitrate assignment, it is conceivable to use multiple quantizers with different quantization step sizes. It is also conceivable to allow quantization at 0 bits (i.e., "infinite step size"), which may mean that the dimension quantized at 0 bits may be replaced by a (pseudo) random realization drawn from, for example, an N(0,1) distribution. In other words, the quantizer can be selected from multiple quantizers associated with a zero-rate noise filter and different levels of SNR. For each coordinate of the latent frame, a quantizer is selected that is governed by the rate assignment. The rate assignment is backward adaptive if it is derived, for example, from parameters related to the mapping G (or F), and forward adaptive if it is derived on the encoder side and involves additional parameters transmitted to the decoder in the bitstream.

[0045] Figure 1e schematically illustrates an example of the quantization operation described herein. Here, the quantization operation (e.g., one used in any of encoders 100, 101, 102, or 103) involves subtractive dithering followed by the use of gain. For example, dither 174(d) may be uniformly distributed as d~U(-Δ / 2, Δ / 2) and may be added and subtracted before quantization Q. The resulting output is, for example, p = 1.0 / √(1.0+Δ 2 It may receive a gain of 175(p), which can be defined as / 12). Of course, other preferred forms of both dither and gain have also been conceived. An example of such a quantization operation may facilitate the reconstruction of latent frames having a distribution that is approximately the same as the distribution of latent frames assumed during model training. This may facilitate the reconstruction of the signal frame in the first domain using the quantized latent frames and updating the model state based on it.

[0046] Since the quantization structure described above may involve the use of subtractive dithering, pseudorandomness may be required to ensure that the encoder and decoder operate synchronously. This can be achieved, for example, by allowing both the encoder and decoder (where the decoder may be local to the encoder, as mentioned above, or a remote decoder that does not form part of the encoder, as will be discussed later in this paper) access to the same source of randomness (i.e., "common randomness"), or at least the same source of pseudorandomness. Here, "the same source" does not necessarily have to be the same physical source, or even the same logical source. For example, it is conceivable that two separate random number generators using the same seed would suffice. For example, there may be separate random number generators initialized with the same seed (implementing the same random number generation algorithm). Similarly, this common randomness may also be used for zero-rate noise filling, which can occur both on the encoder side (e.g., in the encoder's local decoder) and the remote decoder side.

[0047] In some embodiments of the encoders described herein, reversible mapping is envisioned to be implemented by a neural network having an architecture that facilitates reversibility and enables training to model multivariate distributions in the latent region, etc. An example of such a network is known as invertible flow. For example, mapping G and G -1 This may be implemented using one or more flow models with parameters controlled by the generative model. This can be particularly useful when the mapping is nonlinear and the flow structure allows for reversibility with respect to such mapping.

[0048] Various embodiments of the decoder according to this disclosure will be described in more detail with reference to Figures 2a and 2b.

[0049] Figure 2a shows the quantized latent frame 264(^Z) containing the quantized representation of the current frame t of the signal in a latent region distinct from the first region. t A decoder 200 is schematically shown, configured to receive a finite bitrate stream 230 containing ) and this quantized latent frame ^Z t Dequantization operation 266(Q) acts on it. -1 Using ), decoder 200 quantizes the latent frame ^Z t Reconstructed from latent frame 267(Z t * The decoder 200 includes a generation model 210, which generates a mapping 268(G) from the latent region to the first region (i.e., the reversible mapping G mentioned earlier). -1 It is trained for the task of generating parameter 220 for the inverse of ( ). Mapping G is reconstructed latent frame Z t * The first region is mapped to the reconstructed signal frame 269(x t * ) generates the reconstructed current frame x of the signal. t * This is then used to update the state of Model 210 (the generative neural network) so that Model 210 is ready to process future frames >t.

[0050] As previously described in this paper, in some embodiments, the decoder 200 and other decoders of the present disclosure provide (additional) conditioning information 265(θ ≧t The system may also be trained to perform the task of generating parameter 220 based on the following: Such conditioned data may be associated with at least one of the current frame t and one or more future frames >t, as previously described in this paper.

[0051] As explained earlier in this paper, quantized latent frame^Z t and conditional information θ ≧tThis can optionally be provided as a two-stage description of the signal in separate finite-bitrate streams, or in the same finite-bitrate stream. If the conditioning information is not included, it is conceivable that the finite-bitrate stream may include at least an indication that the conditioning information is not included, so that the decoder knows what to expect and how to process the quantized latent frame using the generative model. For example, if such an indication is present, the generative model can perform the task of generating parameters without using such conditioning information. In some embodiments, i.e., if legacy codec data is provided instead of actual conditioning information, the decoder can generate such conditioning information itself using the legacy codec data (e.g., by waveform reconstruction or from other information provided in the legacy codec data).

[0052] In general, the task of generating / generating parameters for mapping G may include predicting the current frame of the signal in a first domain, and generating the current frame of the signal in the first domain may include modifying the predicted current frame using a reconstructed latent frame mapped to the first domain.

[0053] Figure 2b schematically shows one embodiment of decoder 201. Here, similar to the local decoder 140 of encoder 101 described with reference to Figure 1c, the mapping 268(G) is affine, and the reversible transform 271(F) (the transform F used in local decoder 140) -1 Corresponding to the inverse transform, and defined using scaling according to at least a subset 273 of the parameters 220 provided by model 210, and the subsequent predicted mean 272 (μ t This includes the addition of ). The decoder 201 can be used together with the encoder 102, for example, as described with reference to Figure 1c.

[0054] The present disclosure also provides a method for encoding a current frame of a signal and a method for decoding a current frame of a signal. Embodiments of such methods will be described in more detail herein with reference to FIGS. 3a and 3b.

[0055] FIG. 3a schematically shows an exemplary flow of a method 310 for encoding a current frame of a signal. In step S311, the current frame (x t ) of the signal 360 in the first region is received. In step S312, a generative neural network model is used, which is trained to perform a task including providing parameters for a reversible mapping G -1 from the first region to a latent region different from the first region. In step S312, using the generative model, at least a part of the current frame x t of the signal in the first region is mapped to the latent region using the reversible mapping G -1 to generate a latent frame 362 (Z t ). The latent frame Z t includes a representation of the current frame of the signal in the latent region. In step S313, a quantized latent frame 364 (^Z t ) is generated based on the generated latent frame Z t . In step S314, a finite bitrate stream 330 including the quantized latent frame ^Z t is generated.

[0056] FIG. 3b schematically shows an exemplary flow of a method 320 for decoding a current frame of a signal. In step S321, the finite bitrate stream 330 is received. The finite bitrate stream 330 includes the quantized latent frame ^Z t , which includes a quantized representation of the current frame of the signal in a latent region different from the first region. In step S322, from the quantized latent frame ^Z t , a reconstructed latent frame 367 (Z t *) is generated (by dequantization). In step S323, a generative neural network model is used. The model is trained to perform a task that includes generating parameters for a reversible mapping G from the latent region to the first region. In step S323, the generative model is used to generate the signal 369(x) in the first region. t * The reconstructed current frame of the signal in the first region x is generated using reversible mapping G. In step S324, the reconstructed current frame x of the signal in the first region is generated. t * However, this is then used to update the state of the generative model (indicated by the dashed arrow 340).

[0057] It is envisioned that the encoding and decoding methods 310 and 320, respectively, may be modified in accordance with what has been described and / or discussed for any of the encoders and decoders disclosed herein. For example, the generative models used in methods 310 and 320 may use conditional information, and the encoding method 310 may include a local decoder as described with reference to Figure 1b, and so on. In other words, it is envisioned that the flows of methods 310 and 320, respectively, may be modified to correspond to any embodiment of the respective encoders and decoders described herein.

[0058] The Disclosure also provides a non-temporary computer-readable medium for storing instructions. When these instructions are executed by at least one computer process belonging to computer hardware, they can function to cause the computer hardware to perform Method 310 and Method 320, or any embodiment thereof described herein. Here, “cause the computer hardware to perform” means that a computer processor can, for example, receive or output one or more signals using a suitable interface provided by such computer hardware, and / or, to do so, perform any other method step, including, for example, using the processor, memory or working memory of the computer hardware to use and implement a generative neural network model. Embodiments of such medium are not shown in any figures of this Application.

[0059] Finally, an embodiment of the coding system for transferring the current frame of a signal according to this disclosure will be described with reference to Figure 4.

[0060] Figure 4 schematically shows an example of a coding system 400. The coding system 400 includes at least one encoder 410 and at least one decoder 420. The encoder processes the current frame 460(x) of the signal in the first region. t ) is received, and a quantized latent frame 464(^Z) containing a representation of the current frame of the signal in a latent region different from the first region is received. t The coding system can output at least one finite bitrate stream 430, including ). The coding system further includes means 440 for transferring the finite bitrate stream 430 between the encoder 410 and the decoder 420, thereby the decoder 420 receiving the finite bitrate stream (including the quantized latent frame) and, based thereon, the signal 469 in the first domain (x t *) can generate a reconfigured current frame. Here, the encoder and decoder may, of course, be any encoder and decoder as described in the various embodiments of this disclosure. Means 440 may include, for example, a data link (e.g., an infrared link, a laser link, an optical link, a wireless link, etc.), or simply means (various interfaces, etc.) that allow the encoder 410 and decoder 420 to use it to connect to an existing communication infrastructure, including, for example, the Internet, in order to transfer a finite bitrate stream 430.

[0061] In summary, this disclosure provides a general concept in which, instead of allocating the entire bit budget to transferring conditioning information for a generative model, at least some bits are allocated to transferring quantized latents. This may be achieved, for example, by using a two-stage description as described in this paper, where the rate allocation / distribution between the two components (conditioning information and quantized latents) is arbitrary, even including situations where all bits are used instead to transfer quantized latents (which may be useful, for example, in low-rate transfers of speech and / or audio, in which case only the quantized latents are used to reconstruct the signal on the decoding side). This disclosure facilitates coding configurations in which coding resources can be distributed between two parts / stages of a signal description. In particular, this disclosure provides a way to improve coding quality in situations where, for example, a generative model cannot accurately capture all the features of an audio signal that are needed for sufficient reconstruction on the decoder side, by allocating some bits instead to transfer quantized latents that may help "fill in the gaps" present in the generative model.

[0062] Various encoders and decoders of the present disclosure, such as those described in the exemplary embodiments above, may be implemented using computer hardware including, for example, a (computer) processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), one or more application-specific integrated circuits (ASICs), one or more radio frequency integrated circuits (RFICs), or any combination thereof) and memory coupled to the processor. As described above, the processor may be adapted to perform some or all of the steps of the methods described through the present disclosure.

[0063] Computer hardware may include, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a web appliance, a network router, a switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be performed by such computer hardware. Furthermore, this disclosure relates to any collection of computer hardware that individually or collectively execute instructions to perform one or more of the concepts discussed herein.

[0064] As used in this paper, the term “computer-readable media” includes, but is not limited to, data repositories in the form of solid memory, optical media, and magnetic media.

[0065] Unless otherwise specified, as will be apparent from the following discussion, any discussion using terms such as “processing,” “computing,” “calculating,” “determining,” and “analyzing” throughout this disclosure is understood to refer to the actions and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or transform data, expressed as physical quantities, such as electronic quantities, into other data, also expressed as physical quantities.

[0066] Similarly, the term “computer processor” may refer to any device or part of a device that processes electronic data from, for example, registers and / or memory to convert that electronic data into other electronic data that can be stored, for example, in registers and / or memory. “Computer” or “calculator” or “computing platform” or “computer processor” may include one or more processors.

[0067] The concepts described herein are executable by one or more processors that accept computer-readable (also called machine-readable) code, which includes a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. This includes any processor capable of executing a set of instructions (sequential or otherwise) that specify an action to be performed. Thus, one example is a typical processing system (i.e., computer hardware) comprising one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem, which may include main RAM and / or static RAM and / or ROM. A bus subsystem for communication between components may also be included. The processing system may further be a distributed processing system having processors connected by a network. If the processing system requires a display, such a display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include one or more input devices, such as an alphanumeric input unit, such as a keyboard, and a pointing control device, such as a mouse. The processing system may also include a storage system, such as a disk drive unit. The processing system in some configurations may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software) which, when executed by one or more processors, causes one or more of the methods described herein to be performed. Note that, if a method includes several elements, for example, several steps, the ordering of such elements is not implied unless specifically stated.The software may reside on the hard disk, or, during its execution by the computer system, may reside entirely or at least partially in RAM and / or the processor. Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may form or be contained within a computer program product.

[0068] In some exemplary embodiments, the one or more processors may operate as standalone devices or, in a networked deployment, may be connected, for example, to another processor; the one or more processors may operate as a server or user machine in a server-user network environment; or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a cellular telephone, a web appliance, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying the actions to be taken by that machine.

[0069] It should be noted that the term “machine” is also interpreted to include any set of machines that individually or jointly execute a set (or set) of instructions for executing one or more of the methodologies discussed herein.

[0070] Therefore, an exemplary embodiment of each method described herein may take the form of a computer-readable carrier medium carrying a set of instructions, for example, a computer program for execution on one or more processors, for example, one or more processors that are part of a web server configuration. Thus, as those skilled in the art will understand, exemplary embodiments of the Disclosure may be embodied as a method, an apparatus such as a special-purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, for example, a computer program product. The computer-readable carrier medium carries computer-readable code that, when executed on one or more processors, causes one or more processors to perform the method. Thus, aspects of the Disclosure may take the form of a method, an exemplary embodiment entirely of hardware, an exemplary embodiment entirely of software, or an exemplary embodiment combining software and hardware aspects. Furthermore, the Disclosure may take the form of a carrier medium (for example, a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0071] The software may also be transmitted and received over a network via a network interface device. While the carrier medium is a single medium in exemplary embodiments, the term “carrier medium” should be understood to include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The term “carrier medium” should also be understood to include any medium capable of storing, encoding, or carrying a set of instructions for execution by one or more of the processors, causing one or more of the processors to execute one or more of the methods of this disclosure. The carrier medium can take many forms, but is not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory such as main memory. Transmission media include coaxial cables, copper wires, and optical fibers, including wires that constitute a bus subsystem. Transmission media can also take the form of sound waves or light waves, such as those generated during radio and infrared data communications. For example, the term “carrier medium” should be understood to include, but not be limited to, computer products embodied in solid memory, optical and magnetic media; media carrying a propagation signal detectable by at least one or more processors and representing a set of instructions that implement a method at runtime; and transmission media in a network carrying a propagation signal detectable by at least one of the one or more processors and representing a set of instructions.

[0072] It will be understood that, in some exemplary embodiments, the steps of the method discussed are performed by a suitable processor(s) of a processing (e.g., computer) system / hardware that executes instructions (computer-readable code) stored in a memory device. It will also be understood that this disclosure is not limited to any particular implementation or programming technique, and that this disclosure may be implemented using any suitable technique for implementing the functions described herein. This disclosure is not limited to any particular programming language or operating system.

[0073] Throughout this disclosure, any reference, for example, to “one exemplary embodiment,” “several exemplary embodiments,” or “a particular exemplary embodiment” means that any specific feature, structure, or characteristic described in relation to that exemplary embodiment is included in at least one exemplary embodiment of this disclosure. Thus, phrases such as “in one exemplary embodiment,” “in several exemplary embodiments,” or “in a particular exemplary embodiment” in various parts of this disclosure do not necessarily all refer to the same exemplary embodiment. Furthermore, any specific feature, structure, or characteristic can be combined in any suitable way in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.

[0074] Where used herein, unless otherwise specified, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to describe a common object simply indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects described in this way must be in a given order, temporally, spatially, in rank, or in any other way.

[0075] In the claims and descriptions herein, any of the terms including, including, or having are open terms meaning that include at least the listed elements / features, but do not exclude others. Therefore, as used in the claims, the terms including / having should not be interpreted as limiting to the enumerated means, elements, or steps. For example, the expression "apparatus having A and B" should not be limited to an apparatus consisting only of elements A and B. Any of the terms including, including, including, or encompassing as used herein are open terms meaning that include at least the enumerated elements / features, but do not exclude others. Therefore, including is synonymous with having and means having.

[0076] In the above description of the exemplary embodiments of this disclosure, it should be understood that, for the purpose of improving the flow of the disclosure and aiding in the understanding of one or more of the various inventive aspects, various features of the disclosure may be summarized in a single exemplary embodiment, figure, or description thereof. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly described in each claim. Rather, as reflected in the following claims, the aspects of the invention are fewer than all the features of a single, aforementioned exemplary embodiment. Thus, the claims following this specification are hereby expressly incorporated herein, and each claim stands alone as a separate exemplary embodiment of the disclosure.

[0077] Furthermore, while some exemplary embodiments described herein may include some features included in other exemplary embodiments, they may not include others. However, combinations of features from different exemplary embodiments are intended to be within the scope of the disclosure and constitute different exemplary embodiments, as will be understood by those skilled in the art. For example, any of the exemplary embodiments described in the following claims may be used in any combination.

[0078] Numerous specific details are described in the descriptions provided herein. However, it is understood that exemplary embodiments of this disclosure may be carried out without these specific details. On the other hand, well-known methods, structures, and techniques are not described in detail so as not to obscure the understanding of this paper.

[0079] Therefore, while what is considered to be the best form of disclosure is described, those skilled in the art will recognize that other further modifications may be made without departing from the spirit of the disclosure, and that all such changes and modifications are intended to be requested as being included within the scope of this disclosure. For example, any of the formulas described above merely represent possible procedures. Functions may be added or removed from the block diagrams, and operations may be swapped between function blocks. Steps may be added or removed from the methods described within the scope of this disclosure.

[0080] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] Decoder(200,201): A step of receiving a finite bitrate stream (230) containing a quantized latent frame (264), wherein the quantized latent frame contains a quantized representation of the current frame (t) of the signal in a latent region different from the first region; The steps include generating a reconstructed latent frame (267) from the quantized latent frame; A step in which a generative neural network model (210) is used to perform a task for which the generative neural network model is trained, the task comprising generating parameters (220) for a reversible mapping (268) from a latent region to a first region; A step of reconstructing the current frame (269) of the signal in a first region, comprising mapping the reconstructed latent frame to the first region using the reversible mapping; The system is configured to perform the step of updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region. decoder. [EEE2] The decoder according to EEE1, wherein the generative neural network model is also trained to perform the task based on conditioning information (265) associated with at least one of the current frame (t) and future frames (>t). [EEE3] The decoder according to EEE1 or 2, comprising means configured to receive in the same finite bitrate frame stream either the quantized latent frame and conditioning information associated with at least one of the current frame and / or future frame, or an instruction that such conditioning information is not included in the same finite bitrate frame stream, wherein the decoder is configured to perform the task without using such conditioning information if the same finite bitrate frame stream includes the instruction. [EEE4] The decoder according to EEE2, comprising means configured to receive the quantized latent frame and legacy codec data in the same finite bitrate frame stream, and further configured to reconstruct a signal from the legacy codec data as at least part of the conditioning information. [EEE5] The decoder according to any one of EEE1 to 4, wherein the task includes predicting the current frame of the signal in a first region, and generating the current frame of the signal in the first region includes modifying the predicted current frame using the reconstructed latent frame mapped to the first region. [EEE6] The encoder is (100, 101, 102, 103): The first step involves receiving the current frame (160) of the signal in the first region; A step comprising using a generative neural network model (110) to perform a task for which the generative neural network model is trained, the task comprising providing parameters (120) for a reversible mapping (161) from a first region to a latent region different from the first region; A step of generating a latent frame (162) by mapping at least a portion of the current frame of the signal to a latent region using the reversible mapping, wherein the latent frame includes a representation of the current frame (t) of the signal in the latent region; The steps include: generating a quantized latent frame (164) based on the generated latent frame; The system is configured to perform the steps of generating a finite bitrate stream that includes the quantized latent frame, Encoder. [EEE7] The encoder according to EEE6, wherein the generative neural network is trained to perform the task based on conditioning information (165) associated with at least one of the current frame (t) and future frames (>t), and the encoder is further configured to output a finite bitrate stream containing such conditioning information. [EEE8] The encoder according to EEE6 or 7, further configured to output the same finite bitrate frame stream, which includes the quantized latent frames and conditioning information associated with at least one of the current frame and future frames, or an indication that such conditioning information is not contained in the same finite bitrate stream. [EEE9] The steps include generating a reconstructed latent frame (167) from the quantized latent frame; A step of generating a reconstructed current frame (169) of the signal in the first region, The steps include mapping the reconstructed latent frame to a first region using the inverse (168) of the aforementioned reversible mapping; The system is further configured to perform a step of updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region. An encoder as described in any one of the EEE6 or EEE8 clauses. [EEE10] The encoder according to EEE9, wherein the aforementioned reversible mapping includes an affine transform. [EEE11] An encoder according to EEE9 or 10, configured to generate the reversible mapping using a flow model. [EEE12] The encoder according to any one of EEE6 to 11, wherein the bitrate of the finite bitrate stream including the quantized latent frames is assigned based on a perceptual rate assignment model (150). [EEE13] The encoder according to any one of EEE9 to 12, wherein the bitrate of the finite bitrate stream including the quantized latent frame is assigned based on the difference between the received current frame of the signal in the first domain and the reconstructed current frame of the signal in the first domain. [EEE14] An encoder according to any one of EEE6 to 13, configured to generate the quantized latent frame using subtractive dithering (174) followed by gain (175). [EEE15] An encoder according to any one of EEE6 to 14, configured to generate the quantized latent frame by selecting from a plurality of quantizers having different quantization step sizes, including a zero-rate noise fill. [EEE16] A method (320) for decoding the current frame of a signal: Step (S321) of receiving a finite bitrate stream (330) containing a quantized latent frame (264), wherein the quantized latent frame contains a quantized representation of the current frame (t) of the signal in a latent region different from the first region; The step (S322) is to generate a reconstructed latent frame (367) from the quantized latent frame; A step of using a generative neural network model to perform a task for which the generative neural network model is trained, the task comprising generating parameters for a reversible mapping from a latent region to a first region; A step (S323) of reconstructing the current frame (369) of the signal in a first region, comprising mapping the reconstructed latent frame to the first region using the reversible mapping; The process includes (340) updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region (S324), method. [EEE17] A method (310) for encoding the current frame of a signal: The step (S311) of receiving the current frame (360) of the signal in the first region; A step of using a generative neural network model to perform a task for which the generative neural network model is trained, the task comprising providing parameters for a reversible mapping from a first region to a latent region different from the first region; Step (S312) of generating a latent frame (362) by mapping at least a portion of the current frame of the signal to a latent region using the reversible mapping, wherein the latent frame includes a representation of the current frame of the signal in the latent region; The steps include: (S313) generating a quantized latent frame (364) based on the generated latent frame; The process includes the step of generating a finite bitrate stream containing the quantized latent frame (S314), method. [EEE18] A non-temporary computer-readable medium storing instructions that, when executed by at least one computer processor belonging to computer hardware, can function to cause the computer hardware to perform the method of decoding the current frame of a signal as described in EEE16. [EEE19] A non-temporary computer-readable medium storing instructions that, when executed by at least one computer processor belonging to computer hardware, can function to cause the computer hardware to perform the method of encoding the current frame of a signal as described in EEE17. [EEE20] A coding system (400) for transferring the current frame of a signal, comprising at least one decoder (420) according to any one of EEE1 to 5, at least one encoder (410) according to any one of EEE6 to 15, and means (440) for transferring the finite bitrate stream (430) containing the quantized latent frame between the encoder and the decoder.

Claims

1. Decoder (200,201): A step of receiving a finite bitrate stream (230) containing quantized latent frames (264), wherein the quantized latent frames contain a quantized representation of the current frame (t) of the signal in a latent region different from the first region; The steps include: generating a reconstructed latent frame (267) from the quantized latent frame; A step of using a generative neural network model (210) to perform a task for which the generative neural network model is trained, the task comprising generating parameters (220) for a reversible mapping (268) from a latent region to a first region; A step of reconstructing the current frame (269) of the signal in a first region, comprising mapping the reconstructed latent frame to the first region using the reversible mapping; The system is configured to perform the steps of updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region, decoder.

2. The decoder according to claim 1, wherein the generative neural network model is trained to perform the task based on conditioning information (265) associated with at least one of the current frame (t) and future frames (>t).

3. The decoder according to claim 1 or 2, comprising means configured to receive in the same finite bitrate frame stream either the quantized latent frame and conditioning information associated with at least one of the current frame and / or future frame, or an instruction that such conditioning information is not included in the same finite bitrate frame stream, wherein the decoder is configured to perform the task without using such conditioning information if the same finite bitrate frame stream includes the instruction.

4. The decoder according to claim 2, further comprising means configured to receive the quantized latent frame and legacy codec data in the same finite bitrate frame stream, and further configured to reconstruct a signal from the legacy codec data as at least part of the conditioning information.

5. The decoder according to any one of claims 1 to 4, wherein the task comprises predicting the current frame of the signal in a first region, and generating the current frame of the signal in the first region comprises modifying the predicted current frame using the reconstructed latent frame mapped to the first region.

6. Encoders (100, 101, 102, 103): The first step involves receiving the current frame (160) of the signal in the first region; A step of using a generative neural network model (110) to perform a task for which the generative neural network model has been trained based on conditioning information (165) associated with at least one of the current frame (t) and future frames (>t), the task comprising providing parameters (120) for a reversible mapping (161) from a first region to a latent region different from the first region; A step of generating a latent frame (162) by mapping at least a portion of the current frame of the signal to a latent region using the reversible mapping, wherein the latent frame includes a representation of the current frame (t) of the signal in the latent region; The steps include: generating a quantized latent frame (164) based on the generated latent frame; The system is configured to perform the steps of generating a finite bitrate stream that includes the quantized latent frame, Encoder.

7. The encoder according to claim 6, further configured to output a further finite bitrate stream containing such conditioning information.

8. The encoder according to claim 6 or 7, further configured to output the same finite bitrate frame stream, which includes the quantized latent frames and conditioning information associated with at least one of the current frame and future frames, or an indication that such conditioning information is not contained in the same finite bitrate stream.

9. The steps include: generating a reconstructed latent frame (167) from the quantized latent frame; A step of generating a reconstructed current frame (169) of the signal in the first region, A step comprising mapping the reconstructed latent frame to a first region using the inverse (168) of the aforementioned reversible mapping; The system is further configured to perform a step of updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region. The encoder according to any one of claims 6 to 8.

10. The encoder according to claim 9, wherein the reversible mapping includes an affine transform.

11. The encoder according to claim 9 or 10, configured to generate the reversible mapping using a flow model.

12. The encoder according to any one of claims 6 to 11, wherein the bitrate of the finite bitrate stream including the quantized latent frames is assigned based on a perceptual rate assignment model (150).

13. The encoder according to any one of claims 9 to 11, wherein the bitrate of the finite bitrate stream including the quantized latent frames is assigned based on the difference between the current received frame of the signal in the first domain and the reconstructed current frame of the signal in the first domain.

14. The encoder according to any one of claims 6 to 13, configured to generate the quantized latent frame using subtractive dithering (174) followed by gain (175).

15. The encoder according to any one of claims 6 to 14, configured to generate the quantized latent frame by selecting from a plurality of quantizers having different quantization step sizes, including a zero-rate noise fill.

16. A method (320) for decoding the current frame of a signal, performed by a decoder, the following: Step (S321) of receiving a finite bitrate stream (330) containing quantized latent frames (264), wherein the quantized latent frames contain a quantized representation of the current frame (t) of the signal in a latent region different from the first region; The step (S322) is to generate a reconstructed latent frame (367) from the quantized latent frame; A step of using a generative neural network model to perform a task for which the generative neural network model is trained, the task comprising generating parameters for a reversible mapping from a latent region to a first region; A step (S323) of reconstructing the current frame (369) of the signal in a first region, comprising mapping the reconstructed latent frame to the first region using the reversible mapping; The process includes the step of updating the state of the generative neural network model using the reconstructed current frame of the signal in the first region (340) (S324), method.

17. A method (310) for encoding the current frame of a signal, performed by an encoder, the following: The step (S311) of receiving the current frame (360) of the signal in the first region; A step of using a generative neural network model to perform a task for which the generative neural network model is trained based on conditioning information (165) associated with at least one of the current frame (t) and future frames (>t), the task comprising providing parameters for a reversible mapping from a first region to a latent region different from the first region; Step (S312) of generating a latent frame (362) by mapping at least a portion of the current frame of the signal to a latent region using the reversible mapping, wherein the latent frame includes a representation of the current frame of the signal in the latent region; The steps include: (S313) generating a quantized latent frame (364) based on the generated latent frame; The process includes the step (S314) of generating a finite bitrate stream containing the quantized latent frame and such conditioning information (165), method.

18. A non-temporary computer-readable medium storing instructions that, when executed by at least one computer processor belonging to computer hardware, can function to cause the computer hardware to perform the method of decoding the current frame of a signal according to claim 16.

19. A non-temporary computer-readable medium storing instructions that, when executed by at least one computer processor belonging to computer hardware, can function to cause the computer hardware to perform the method of encoding the current frame of a signal according to claim 17.

20. A coding system (400) for transferring the current frame of a signal, comprising at least one decoder (420) according to any one of claims 1 to 5, at least one encoder (410) according to any one of claims 6 to 15, and means (440) for transferring the finite bitrate stream (430) containing the quantized latent frame between the encoder and the decoder.