Signal Encoding Using Latent Feature Prediction

By incorporating context encoding and learnable amplitude compression into the VQ-VAE framework, the solution addresses inefficiencies in neural audio codecs, improving real-time audio communication quality and reducing latency.

JP2025522259APending Publication Date: 2025-07-15MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024563546
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing neural audio codecs fail to fully utilize temporal correlation in latent features, leading to redundancies and inefficiencies in data transmission, which can cause delays and quality issues in real-time audio communication.

Method used

Introduce context encoding using temporal prediction in the latent representation of a VQ-VAE framework, employing a learnable extractor and synthesizer to fuse latent features and quantization outputs, and utilize time/frequency bins, learnable amplitude compression, and vector quantization for efficient encoding and decoding.

Benefits of technology

The solution reduces latency and enhances coding efficiency by removing redundancies, achieving high-quality audio transmission at low bitrates with resilience against packet loss and background noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522259000001_ABST
    Figure 2025522259000001_ABST
Patent Text Reader

Abstract

Describe techniques and solutions for encoding and decoding signals such as audio data. The disclosed innovations can find specific uses in applications of voice encoding, such as in real-time communication. Using a neural network, context encoding can be used to encode the latent features in the current frame using predictions from the reconstructed latent features of past frames as context. Based on such predictions and latent features of the current frame obtained using an encoder, an extractor learns features such as residuals. These features such as residuals are then quantized. In the decoder part of the encoding framework, the quantized features are dequantized and then combined with predictions from the latent features reconstructed so far to generate the reconstructed features of the current frame, which can then be processed by the decoder to generate the reconstructed signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Field

[0001] This disclosure generally relates to signal encoding. Specific implementations provide neural encoding of audio data using latent feature prediction.

Background Art

[0002] Background

[0002] Since at least the early 1970s, digital technologies have been used to record, store, and transmit audio information. With the advent of the Internet, the use of digital audio transmission has exploded, including the use of real-time streaming in voice-over-IP applications and services, such as Microsoft Teams (Microsoft Corp., Redmond, Washington). The computing power of personal computing devices, like the networking infrastructure, continues to improve, but it remains interesting to improve audio quality while reducing the amount of data required to transmit audio information. Specifically, real-time audio can be relatively sensitive to transmission and processing delays because there may be only limited buffering available for audio signals. For example, delays in audio processing can prevent participants in a call from effectively communicating with each other. Thus, there is room for improvement.

Summary of the Invention

Means for Solving the Problems

[0003] Summary

[0003] This summary of the invention is presented to introduce selected concepts in a simplified form that will be further described below in the mode for carrying out the invention. This summary of the invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004]

[0004] Techniques and solutions for encoding and decoding signals such as audio data are described. The disclosed innovation can find specific uses in applications such as real-time communication. Using a neural network, context encoding can be used to encode the latent features in the current frame using predictions from the reconstructed latent features of past frames as context. Based on such predictions and latent features of the current frame obtained using an encoder, an extractor learns features such as residuals. The features such as residuals are then quantized. In the decoder part of the encoding framework, the quantized features such as residuals are dequantized and then combined with predictions from the latent features reconstructed so far to generate the reconstructed features of the current frame, which can then be processed by the decoder to generate the reconstructed signal.

[0005]

[0005] In one aspect, a method for encoding a signal such as digital audio data is provided. Using an encoder, one or more latent features are extracted from a frame of the input signal. The prediction of the one or more latent features is determined using the reconstructed latent features in a plurality of previous frames. From the one or more latent features and the prediction, features such as residuals are extracted. The features such as residuals, or sufficient data to reconstruct the features such as residuals, are transmitted to the client.

[0006]

[0006] The present disclosure also includes a computing system and a tangible, non-transitory, computer-readable storage medium configured to perform the foregoing method or including instructions for performing the foregoing method. As described herein, various other features and advantages can be incorporated into each technique as desired.

Brief Description of the Drawings

[0007] Brief Description of the Drawings

Figure 1

[0007] Diagram of a conventional vector quantization variational autoencoder.

Figure 2

[0008] Diagram of a vector quantization variational autoencoder according to an embodiment of the present disclosure.

Figure 3

[0009] Present a diagram of a filtering technique that can be used in the vector quantization variational autoencoder of FIG. 2.

Figure 4

[0010] Show a diagram showing a modified form of the vector quantization variational encoder of FIG. 2.

Figure 5

[0011] Graph of the audio quality of various audio coding techniques at various bitrates.

Figure 6

[0012] Show a table showing the comparison results of various implementation forms of the vector quantization variational autoencoder of FIG. 2.

Figure 7

[0013] Diagram of an integrated system according to the present disclosure for using latent feature prediction as context in a vector quantization variational autoencoder.

Figure 8A

[0014] Show the details of the encoder part of the system of FIG. 7.

Figure 8B

[0014] Show the details of the decoder part of the system of FIG. 7.

Figure 9

[0015] Diagram showing a technique for vector quantization for each group.

Figure 10

[0016] Flow diagram of an exemplary signal coding technique according to the present disclosure.

Figure 11

[0017] Diagram of an exemplary computing system that can implement some of the described embodiments.

Figure 12

[0018] Exemplary cloud computing environment that can be used with the technology described herein.

Best Mode for Carrying Out the Invention

[0008] Detailed Description Example 1 - Overview

[0019] Since at least the early 1970s, digital technology has been used to record, store, and transmit audio information. With the advent of the Internet, the use of digital audio transmission, including the use of real-time streaming in voice-over-IP applications and services, such as Microsoft Teams (Microsoft Corp., Redmond, Washington), has exploded. The computing power of personal computing devices continues to improve, as does the networking infrastructure, but there is still a need to improve audio quality while reducing the amount of data required to transmit audio information. Specifically, real-time audio can be relatively sensitive to transmission and processing delays because there may only be limited buffering available for audio signals. For example, delays in audio processing can prevent participants in a call from effectively communicating with each other. Therefore, there is room for improvement.

[0009]

[0020] Artificial intelligence / machine learning techniques, such as neural networks, have been applied to audio data, including for real-time communication. Existing neural audio codecs can be classified into two types. One type of neural audio codec is based on a generative decoder model. At least some generative decoder models extract acoustic features from audio data, encode them after quantization and entropy coding, and use a powerful decoder to recover the waveform based on the generative model.

[0010]

[0021] Another type of audio codec that has been studied is based on end-to-end neural audio coding. An end-to-end neural network typically utilizes, as an example 100 shown in FIG. 1, a VQ-VAE (vector quantization variational autoencoder) framework to learn, in an end-to-end manner, an encoder 110, a vector quantizer 120, and a decoder 130 as shown in FIG. 1. The latent features to be quantized, generated from the encoder, are mainly blindly learned using a convolutional neural network (CNN) without any prior knowledge of their semantics. These methods can enhance the coding efficiency by achieving high quality at low bitrates. However, in these algorithms, temporal correlation is not fully utilized. In the encoded features, there still exist a lot of redundancies between adjacent frames. In contrast, the disclosed technological innovation incorporates context coding into a VQ-VAE-based neural codec framework to remove such redundancies in the latent domain and thus further enhance the coding efficiency.

[0011]

[0022] In the encoding of images, videos, and audio, such as JPEG, HEVC, H.264 / AVC, DPCM / ADPCM, prediction has been used to remove redundancy. In the encoding of video images and within frames, in the area of pixels or frequencies, the reconstructed adjacent blocks are used to predict the current block, and the predicted residual is quantized and encoded into the bitstream. In the inter-frame encoding of video codes, the reconstructed reference frame is used to predict the current frame using motion compensation. The residual after prediction is much sparser, and the entropy is significantly reduced. In neural video coders, such temporal correlation can be exploited by using the motion-aligned reference frame as a prediction or context for encoding the current frame. In audio encoding, DPCM / ADPCM has been used to encode audio samples or acoustic parameters. However, such techniques have not yet been studied for use in neural audio coders.

[0012]

[0023] The present disclosure realizes the introduction of context encoding using temporal prediction into the VQ-VAE framework for neural audio encoding. To reduce latency, this prediction is performed in the latent representation. Different from the conventional video / audio encoding that determines the residual by subtracting the sample from the prediction, a learnable extractor and synthesizer are used to fuse the latent features and quantization output with the prediction.

[0013]

[0024] The disclosed innovation has specific application examples for low-latency audio coding, but can be incorporated into other coding techniques, can be used with other types of signals besides audio voice data, and includes data other than audio data. The present disclosure implements several innovations that, while not essential, can be used with each other. These innovations use time / frequency bins as input for a neural encoder, learnable amplitude compression, context encoding of the latent space for an end-to-end neural audio codec, an improved vector quantization technique with rate control, and the possibility of a relatively high transmission bit rate to achieve scalable quality using the same coding framework, including using an extensible coding framework.

[0014] Example 2 - Exemplary variational autoencoder using time filtering

[0025] In one aspect, the present disclosure provides a codec that includes a neural network, sometimes referred to as "TFNet", that uses a time / frequency input. A specific implementation form 200 of TFNet is shown in FIG. 2. This implementation form includes a causal 2D encoder 204, the output of which is processed using a time filter 208 that includes a time convolution module (TCM) and a gated recurrent unit for each group (G-GRU) in an interleaved manner. The output of filter 208 is quantized by a vector quantizer 212 using a codebook 216 to supply a quantized input 220. The quantized input 220 is then supplied to a time filter 224 that is configured similarly to filter 208 and includes a time convolution module interleaved with a gated recurrent unit for each group, (such as after being transmitted through the network). The output of filter 224 is supplied to a causal 2D decoder 228. Next, the operation of the TFNet implementation form 200 will be further described.

[0015]

[0026] The TFNet-based codec takes a time / frequency spectrum input. This time / frequency spectrum input can be obtained by splitting the audio samples into overlapping windows and applying the Short-Time Fourier Transform (STFT) to each window input to obtain frames, where the hop size determines how often this input is processed. These parameters can be selected as desired, but when used for audio processing, a window size of 20 ms with a hop length of 5 ms can yield good results.

[0016]

[0027] Optionally, the input can be further processed before being fed into the encoder neural network. Specifically, power-law compression to the amplitude can be applied to the input. The dynamic range of audio can be high due to harmonics. This compression acts to normalize the input so that the importance of various frequencies is balanced and training is more stable. Optionally, other compression techniques can be used to compress the amplitude of the input to the encoder 204.

[0017]

[0028] Encoder 204 utilizes local two-dimensional (2D) correlation relationships. Temporal filters 208, 224 utilize relatively long-term temporal dependencies with past frames for feature extraction. This two-level feature extraction helps to learn to extract features with good representational ability, realizes error tolerance against packet loss, and in some cases, removes unnecessary information such as background noise when necessary. The learned features are then quantized via a learned vector quantizer and encoded in fixed-length coding or Huffman coding. In decoding, there are several temporal filtering blocks followed by a decoder for reconstruction. When power-law compression is applied to the amplitude in encoding, inverse power-law compression can be applied to the amplitude of the decoded spectrum. Considering packet loss in real-time communication, it is preferable that the decoding has resilience against these losses in case of recovery ability and minimum error propagation. Thus, a heterogeneous structure is brought about with relatively more temporal filtering blocks in decoding than in encoding.

[0018]

[0029] The entire network is trained end-to-end to optimize the reconstruction quality under rate constraint conditions. The convolution is causal in the temporal dimension, and thus the system can maintain a short latency, such as 20 ms latency in some cases.

[0019] Encoder and Decoder of Example 3 - TFNet

[0030] Referring to Figure 2, encoder 204 includes several causal 2D convolutional layers, each followed by batch normalization (BN) and parametric ReLU (PReLU) for non-linearity. After each convolutional layer, the features are downsampled by a factor of 2 or 4 in the frequency dimension, and finally, all frequency information is folded into each channel.

[0020]

[0031] Let the input feature be X I ∈R T×F×2It is represented by. After being processed by the encoder 204, this feature is X at the input to the temporal filter 208 E ∈R T×1×C This is the case. T, F, and C are the number of frames, frequency bins, and channels, respectively. The convolution is causal along the time dimension, and thus T is maintained without downsampling. The decoder 228 is symmetric to the encoder 204 having a causal 2D inverse convolution layer. The output of the decoder is the reconstructed spectrum X R ∈R T×F×2 which is processed using the inverse short-time Fourier transform to generate the output waveform.

[0021] Example 4 - Exemplary Temporal Filtering

[0032] As shown in Example 2 and as shown in FIG. 3, the filters 208, 224 of the implementation form of TFNet include an extended temporal convolution module (TCM) 300 and a gated recurrent unit per group (G-GRU) 350. Both of these filter elements are causal and have low complexity. The TCM module can be implemented in the same way as the module described in "TCNN: Temporal convolutional neural network for real-time speech enhancement in the time domain" by Pandey et al., IEEE International Conference on Acoustics, Speech and Signal Processing, 6875 - 6879 (2019). According to Pandey's reference, the following is shown. The residual block consists of three convolutions: an input 1×1 convolution, a depthwise convolution, and an output 1×1 convolution. The depthwise convolution is used to further reduce the number of parameters. In the depthwise convolution, the number of channels is maintained the same, and only one filter per input channel is used for the output calculation.

[0022]

[0033] The TCM module includes two 1×1 kernel size convolutional layers 304, 308 for changing the channel dimension, and an extended depthwise convolutional layer 312 for exploiting low-complexity temporal correlations. Several TCM blocks with different expansion rates are grouped as large blocks to increase the receptive field and diversity.

[0023]

[0034] For each group of filters 208, 224, the GRU part divides the channels into N groups and independently exploits the temporal dependencies within each group. The operation of the gated recurrent unit is described in Cho et al., "On the Properties of Neural Machine Translation: Encoder-Decoder Approaches", arXiv:1409.1259 (2014). Specifically, Cho's reference describes that gating can be achieved using activation functions. This extends the normal logistic sigmoid activation function using two gating units called the reset gate r and the update gate z. Each gate depends on the previous hidden state h (t-1) and the current input x t to control the flow of information.

[0024]

[0035] Further details of the gated recurrent unit are presented in Cho et al., "Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation", arXiv 1406.1078v3 (2014). The following is described in this Cho reference. When the reset gate is close to 0, the hidden state has to ignore the previous hidden state and is reset only by the current input. This allows the hidden state to effectively discard any information that has been determined to be irrelevant later, thus enabling a more compact representation. On the other hand, the update gate controls how much information from the previous hidden state is carried over to the current hidden state. This operates similarly to the memory cells within the LSTM network and helps the RNN remember long-term information.

[0025]

[0036] Each group's GRU variant not only reduces complexity but also improves flexibility and representational ability for realizing frequency recognition time filtering when channels are learned from frequency. TCM can help explore short-term and medium-term temporal evolution, and GRU can help capture long-term dependencies. Therefore, interleaving these two techniques helps capture short-term and long-term temporal correlations at various depths. The experimental results presented in Example 8 verify that the interleaved structure is more efficient than the single structure.

[0026] Example 5 - Exemplary Vector Quantization

[0037] The vector quantizer discretizes the learned features when encoding using a set of learnable codebooks according to the target bitrate. Before quantization, the features of X Q ∈R T×1×C after being encoded are X by 1×1 convolution (C’ < C) Q ∈R T×1×C’It is reduced to. Group quantization is obtained by dividing channel C’ into N groups and encoding each group with an independent codebook. Let S represent the number of codewords in each codebook and K = C’ / N represent the dimension of each codeword. In a specific example of implementation form 200, an STFT window length of 20 ms and a hop length of 5 ms are adopted. Therefore, when using fixed-length encoding, the bit rate is given by N×log2S / 5 kbps. At 6 kbps, C’, N, S, and K can be set to 120, 3, 1024, and 40 respectively, although other parameter values can be used as needed. The codebook is learned using exponential moving average according to the technique described in "Neural discrete representation learning" by van den Oord et al., arXiv:1711.00937 (2017). According to that technique, the encoder network outputs a discrete code instead of a continuous code and uses the learned ones so far instead of static ones. The discrete code can be determined using a nearest neighbor search procedure using a shared embedding space. Since the encoder and decoder share the same dimensional space, learning is performed by passing the gradient from the decoder input to the encoder. The shared embedding space, that is, the codebook, is updated as a function of the moving average of the encoder output z e (x).

[0027]

[0038] Specifically, the input x is passed through the encoder to generate an output z e (x), where a shared embedding space e (with embedding vector e j ) for nearest neighbor lookup can be used to determine a discrete latent variable z. Then, the encoder output can pass through a discretization bottleneck and then be mapped to the nearest embedding e. The following formula can be used, where q(z = k|x) is the posterior categorical distribution probability and z q (x) is the nearest embedding.

Equation

[0028]

[0039] Quantized features

Number

[0029] Example 6 - Exemplary loss function

[0040] An exemplary loss function that can be used in the system 200 is a combination of two terms L = L recon + αL VQ where L recon is the reconstruction loss, and L VQ imposes a constraint condition on the vector quantization. In the reconstruction loss, the mean squared error can be used for the power-law compression spectrum between the original signal and the decoded signal. To help achieve the consistency of the STFT, the decoded spectrum can be first converted to the waveform domain via the inverse STFT and then converted back to the time / frequency domain via the STFT to calculate the loss. The second term L VQ is the commitment loss used in the VQ-VAE that causes the encoder 204 to generate a representation close to its codeword, and α is a weight coefficient for balancing the two terms.

[0030] Example 7 - Exemplary TFNet implementation form

[0041] In real-time communication, in addition to quality degradation due to audio encoding, such as background noise and packet loss, there are several types of degradation. With the disclosed end-to-end learnable codec, it is possible to jointly optimize audio encoding using speech enhancement (SE) and packet loss concealment (PLC) when used for audio applications. Two ways of joint optimization are realized, namely, (1) a cascade network (network 400 in FIG. 4) having an enhancer before the codec and a PLC network thereafter, and (2) an all-in-one network (network 450 in FIG. 4) having a network structure similar to the codec but optimized for noisy inputs with packet loss.

[0031]

[0042] The cascade network 400 in FIG. 4 includes three modules: a preprocessing enhancer 410, an audio codec (encoder 420 and decoder 424), and a postprocessing PLC network 440. Since speech is relatively more efficient in compression than noisy audio, the enhancer 410 is placed before the codec 420. The enhancer 410, encoder 420 and decoder 424, and PLC network 440 can all be based on a structure such as TFNet (in FIG. 2 etc.) and are jointly trained in an end-to-end manner. That is, for example, the encoder 420 can include the functions of the encoder 208 and filter 208, and the decoder 424 can include the functions of the filter 224 and decoder 228.

[0032]

[0043] The pre - processing enhancer 410 takes in noisy audio as input, outputs enhanced audio, and supplies it to the codec. Different from the TFNet - based codec implementation form 200, in order to remove information loss, there is a skip connection between the encoder and the decoder in the enhancer 410. In the decoder, a causal - gated block can be used to output the amplitude gain and phase for reconstruction, and this can be implemented in the same way as described in Zheng et al.'s "Interactive speech and noise modeling for speech enhancement" in AAAI (2021). In Zheng, this gated block aims to "learn a multiplication mask for the corresponding features from the encoder and suppress its unnecessary parts."

[0033]

[0044] Under packet loss, the neural codec is adjusted in that when decoding, it treats the quantized features with lost packets as zero and takes a mask indicating where this loss occurs as input. This mask is also injected into each time - filtering block when decoding. The post - processing PLC module 440 operates in the waveform domain and takes as input a TFNet - based structure having both the decoded audio and the mask. Similar to the enhancer 410, there is also a skip connection in the PLC network 440. As a restoration task, the PLC network 440 outputs complex residuals in the time / frequency domain, and this is added to the spectrum of the audio decoded for reconstruction.

[0034]

[0045] For training, the three networks can be connected and jointly trained end - to - end. For better quality, two - stage training can be used. First, the enhancer 410 and the codec 420 can be trained separately using noisy data and noise - free data respectively. Then, the cascade network 440 is... reconUsing the same reconstruction loss, two additional monitors can be used respectively in the outputs of the enhancer and the codec, and then fine-tuned.

[0035]

[0046] The all-in-one network 450 has resilience to both background noise and packet loss using only a single codec network having the same general structure as the TFNet implementation form 200, including an encoder 460 (including the functions of both the encoder 204 and the filter 208) and a decoder 470 (including the functions of both the filter 224 and the decoder 228). To adapt to packet loss, the decoding part in the codec is adjusted in the same way as that in the cascade network 400. This is trained from the beginning using the auxiliary monitoring added in the encoding part to remove noise for efficient encoding. This is achieved by adding a decoder after the time filtering block of the encoder, which will necessarily output noise-free audio during training. During inference, this decoder is not required.

[0036] Example 8 - Exemplary comparison results

[0047] From the "Deep Noise Suppression Challenge" at ICASSP (2021), 16 kHz noisy audio with 890 hours was synthesized, having noise-free speech, noise, and indoor impulses. The noise-free audio included multi-language speech, emotions, and singing clips. The signal-to-noise ratio was randomly selected to be between -5 dB and 20 dB, and the speech level was within the range of -40 to -10 dB. Each audio was cut into 3-second segments for training. Voice enhancement performed both noise removal and reverberation removal. Packet loss was simulated according to the three-state model described in Milner et al.'s "An analysis of packet loss models for distributed speech recognition" Proceedings INTERSPEECH, 8th International Conference on Spoken Language Processing (2004). In this three-state model, one state corresponds to a "good" state where no packet loss occurs, another state corresponds to a "bad" state where there is a possibility of packet loss, and the final state can represent a transition from the "good" state to a new state that is also not related to packet loss. For testing, 1400 audios were used, each with a length of 10 seconds and no overlap with the training data.

[0037]

[0048] During training, the Adam optimizer (see Kingma et al.'s "Adam: A Method for Stochastic Optimization" arXiv:1412:6980 (2014)) was used at a learning pace of 0.0004. The network was trained in 100 epochs with a batch size of 200. The "Adam" algorithm is "first-order gradient-based optimization of probabilistic objective functions based on adaptive estimates of lower-order moments".

[0038]

[0049] In the evaluation, except for the subjective listening test, three metrics, namely, PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and DNSMOS (Deep Noise Suppression Mean Opinion Score), were used for ablation studies to evaluate simultaneous optimization and for evaluation of the time filter type. These metrics were not designed and optimized for exactly the same task, but they were found to be in good agreement with perceptual quality for the same type of distortion in all the compared methods.

[0039]

[0050] The codec network was trained and measured on noise-free data from the "Deep Noise Suppression Challenge". The subjective listening test was conducted using a crowdsourcing method inspired by MUSHRA (Multiple Stimuli with Hidden Reference and Anchor). There were 10 participants. Each participant evaluated 12 samples. The TFNet-based neural codec was compared with two codecs used for real-time communication, Lyra (a neural audio codec from Google LLC) and Opus (Xiph.Org Foundation). As shown in Figure 5, the disclosed TFNet technique at about 3 kbps clearly outperforms Lyra at 3 kbps, and TFNet at 6 kbps is much better than Opus at 6 kbps, which proves the superiority of the disclosed TFNet technique.

[0040]

[0051] The simultaneous optimization of codec, voice enhancement, and PLC (packet loss concealment) was evaluated using noisy / noisy paired data with simulated packet loss traces. The following three methods were compared: a baseline using separately trained enhancement, encoding, and PLC models, a cascade network, and an all-in-one network. In the baseline, the encoding and PLC networks were trained using only the raw noise-free data. The enhancer and PLC networks had 470K parameters and 1.2M MACs per 20 ms, which was much less than the codec network with 5M parameters.

[0041]

[0052] Tables 610 and 620 in FIG. 6 present the comparison results for two- and three-task simultaneous optimization, respectively. It is observed that the two simultaneous optimization methods clearly outperform the baseline in all metrics. Although no preprocessing or postprocessing networks are used, the all-in-one network functions competitively with the cascade network and demonstrates the strong discriminative and representational capabilities of TFNet. Another observation is that the PLC network trained on raw noise-free data in the baseline method becomes sensitive to input mismatches.

[0042]

[0053] The interleaved structure in the TFNet neural codec was compared with the separate use of two modules, TCM and GRU, commonly used in the regression task of voice enhancement. In encoding and decoding, all schemes were compared under the same computational complexity using 1.4M parameters and 3.3M MACs per 20 ms window. All time filtering modules were used only for decoding to evaluate their recovery capabilities.

[0043]

[0054] Table 630 in FIG. 6 shows the comparison results. It can be seen that the interleaved structure is optimal for capturing both short-term and long-term temporal correlations.

[0044] Example 9 - Overview of Vector Quantized Variational Autoencoder Using Latent Feature Prediction

[0055] Examples 9 to 13 describe a low bitrate scalable context neural audio codec for real-time communication based on the VQ-VAE framework. This codec incorporates the features of the codec described in Examples 1 to 8. The codecs of Examples 9 to 13 learn encoding, vector quantization codebook, and decoding in an end-to-end manner. Different from existing neural audio codecs that utilize either acoustic features or learned blind features using a convolutional neural network for encoding where temporal redundancy still exists within the quantized features, context encoding using latent feature prediction is introduced into the VQ-VAE framework to further remove such redundancy. Using group vector quantization for each channel with random dropout helps to provide bitrate scalability in a single model and a single bitstream. Subjective evaluation shows that the disclosed technique can achieve an acceptable audio quality at 1 kbps and a nearly transparent quality at 6 kbps.

[0045]

[0056] The disclosed techniques realize several features and advantages that can be used in real-time communication applications as well as other applications, including compressing other types of audio information. One feature is that time / frequency bins are used as network inputs for end-to-end neural audio encoding. Another feature is the use of learnable amplitude compression for low bitrate encoding. Context encoding in the latent domain is used for end-to-end neural audio encoding. The disclosed techniques also realize the feature of vector quantization that supports rate control. A further feature is per-channel bitrate scalability, where the audio quality can be adjusted to a relatively high level as the bitrate increases.

[0046] Example 10 - Exemplary Vector Quantized Variational Autoencoder Using Latent Feature Prediction

[0057] FIG. 7 shows an exemplary neural codec 700 in specific Examples 9 - 13 according to the present disclosure that performs context encoding in the latent representation to reduce latency. As shown in FIGS. 8A and 8B, the codec 700 is divided into an encoding part 800 and a decoding part 850. This technique is described with a specific application to low-latency audio encoding, but can be used for other applications, such as as a codec for other types of audio. The basic encoder and decoder networks 800, 850 are the same as those described for Examples 1 - 8.

[0047]

[0058] The encoder 704 is applied to extract a latent representation r from the input audio x (FIG. 8A). For each frame r in r, the encoder 704 utilizes a prediction learned from the reconstructed past latent code t through a predictor 708 having a receptive field of N past frames. Then, the extractor 712 extracts r

Number

[0048]

[0059] In the decoding part 850 (FIG. 8B), features like dequantized residuals are merged with a prediction p from the past latent features reconstructed via a synthesizer 730 t to obtain a reconstructed current latent code

Number

[0049] Example 11 - Exemplary Amplitude Compression of Input Data

[0060] A typical neural network takes either time-domain samples in end-to-end neural coding or mel-scale features in generative neural coding. The disclosed technique uses the short-time Fourier transform (STFT) domain for feature extraction. As an encoder input, the time / frequency spectrum X by STFT t,f is used. Due to the harmonics of the audio, X t,fThe dynamic range in is large, which may cause the training to become unstable. To balance the importance of various frequencies and bitrates,

Number

[0050] Example 12 - Exemplary Latent Region Context Encoding

[0061] Since context encoding is autoregressive, reducing the delay (at 750) is studied in the latent region. As shown in Figure 7 and as the encoding / decoding split is shown in Figures 8A and 8B, the reconstructed past latent features

Number

Number

Number

[0051]

[0062] The predictor 708 uses a window of N frames to

Number

[0052]

[0063] To induce the predictor 708 with good prediction accuracy, L p =E(D(p t ,sg(r t ))) is used as the prediction loss incorporated into the training, where D(·) is the distance metric given by L1. sg(·) is the gradient stop operator used for relatively stable training.

[0053]

[0064] Both the extractor 712 and the synthesizer 730 include one convolutional layer with a kernel size of 1, followed by a parametric ReLU as a non-linear activation function.

[0054]

[0065] Since quantization is not differentiable, techniques are used to learn the codebook and perform backpropagation through the vector quantization process. Suitable methods include VQ-VAE with a commitment loss, exponential moving average (EMA), Gumbel-Softmax (see Jang et al.'s "Categorical Reparameterization with Gumbel-Softmax" ICLR (2017)), and soft-to-hard (see Augustsson et al.'s "Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations" arXiv 1704.00648v2 (2017)). According to Jang, Gumbel-Softmax includes "a continuous distribution on the simplex that can approximate categorical samples and whose parameter gradients can be easily computed using the reparameterization trick." Further, "the Gumbel-Softmax distribution interpolates between the discrete one-hot encoded categorical distribution and the continuous categorical density." According to Augustsson, "soft-to-hard" uses a soft assignment of a given scalar or vector that is quantized to quantization levels. The parameter controls the "hardness" of the assignment and can gradually transition from soft assignment to hard assignment during training. In contrast to rounding-based or probabilistic quantization schemes, our encoding scheme is directly differentiable and thus end-to-end trainable. Among these methods, Gumbel-Softmax and soft-to-hard take into account the probability of selecting the codeword and thus enable rate control.

[0055]

[0066] However, Gumbel-Softmax selects a codeword using a linear projection without clearly correlating with quantization error. The soft-to-hard technique gives a soft assignment based on the distances to various codewords, but due to quantization in training, a weighted average of codewords is used instead of a single codeword, which leads to a gap between training and inference.

[0056]

[0067] In light of this, in a specific implementation form, a modified mechanism combines the distance-to-soft mapping with Gumbel-Softmax to achieve a non-linear projection as opposed to the linear projection of Gumbel-Softmax. Let K represent the number of codewords in the codebook C. The probability of selecting the k-th codeword c k and quantizing n t is given by the following formula. d t,k =D(n t ,c k )

Equation

[0057]

[0068] Here, τ is the temperature of Gumbel Softmax, and v k ∈Gumbel(0,1) is a sample drawn from the Gumbel distribution. D(·) is a distance metric, and L2 is used in a specific implementation form. α is a scalar for controlling the mapping from the distance D(n t ,c k ) to logits. A hard assignment can be used during the forward pass, and the gradient with respect to the logits is used during the backward pass.

[0058]

[0069] For q t,k , rate control is performed over each mini-batch according to the following formula.

Equation

[0059]

[0070] Here, E b.t (·) is the expected value over all T frames of all B audio in each training mini-batch. R target is the average number of target bits per frame. This loss L r not only constrains the average rate, but also [Number] also performs optimization of rate distortion by []. The current entropy, when higher than R target will push similar features quantized to the same codeword through the trade-off between rate and distortion, while on the other hand, when this is lower than R target similar features are quantized to various codewords and may be of higher quality but also retain a higher rate.

[0060]

[0071] Group vector quantization is adopted to reduce the size of the codebook for easy training. Specifically, each frame n t is divided into G groups along the channel dimension (Figure 9), and each group n t,c is quantized using a separate codebook with K codewords. In a specific implementation form, using a large codebook size can help capture the actual distribution of latent features through rate distortion optimization. For example, if it is desired to achieve 6 kbps in 16 kHz audio, each new 20 ms data is expected to consume 120 bits. Then, the codebook is set by G·log2(K)>120.

[0061] Example 13 - Exemplary Techniques for Bitrate Scalability

[0072] Bitrate scalability is a desirable feature for streaming and real-time communication to support various receivers with different network conditions without any transcoding. Bitrate scalability can support multiple bitrates in a single bitstream. Specifically, the bitstream can be divided into S layers {B i |i = 0, 1, ..., S - 1}, where B0 is the base layer and B1, B2, ..., B S-1 are enhancement layers. A receiver with only B0 has the lowest quality, while a receiver with B0, B1, B2, ... B i-1 (i < S) has higher quality. The best quality is achieved when i = S.

[0062]

[0073] Existing scalable neural audio coders generally utilize residual vector quantization to achieve bitrate scalability, where each channel is trained using a single codebook at the lowest bitrate and relatively more codebooks are used at relatively high bitrates to encode the residual between the encoder features and their previous reconstruction. Instead, dropout during training can be used to utilize group VQ for each channel as described above to achieve bitrate scalability for each channel. As shown in Figure 9, the i-th group feature n t,i uses a separate codebook for quantization. n t,0 can be regarded as the base layer, and n t,1 , n t,2 , ..., n t,G-1 can be regarded as enhancement layers, thus supporting a G bitrate. During training, for each mini-batch, a bitrate b s is randomly selected, which uses only {n t,i |i = 0, 1, ..., s, s < S} and sets the features for achieving a bitrate exceeding b s to zero. That is, {n t,i|i = s + 1, s + 2,..., S - 1} is set to zero. In this way, the decoder is induced to learn restoration from multiple bitrates, and thus scalability can be achieved in a single model.

[0063] Example 14 - Exemplary Encoder Using Latent Feature Prediction

[0074] FIG. 10 is a flowchart of an exemplary signal encoding method 1000 according to the present disclosure. At 1010, using an encoder, one or more latent features are extracted from a frame of an input signal. The prediction of one or more latent features is determined at 1020 using the reconstructed latent features in a plurality of previous frames. At 1030, from the one or more extracted latent features and the prediction, features such as residuals are extracted. For decoding, at 1040, features such as residuals, or data sufficient to reconstruct features such as residuals, are transmitted to the client.

[0064] Example 15 - Additional Examples

[0075] Example 1 is a computing system comprising at least one memory and at least one hardware processor coupled to the at least one memory. The computing system further includes one or more computer-readable storage media that, when executed, store computer-executable instructions that cause the computing system to perform various operations. The operations include using an encoder to extract one or more latent features from a frame of an input signal, resulting in one or more extracted latent features. The prediction of one or more latent features is determined using the reconstructed latent features in a plurality of previous frames. From the one or more extracted latent features and the prediction, features such as residuals are extracted. Features such as residuals, or data sufficient to reconstruct features such as residuals, are transmitted to the client.

[0065]

[0076] Example 2 includes the subject matter of Example 1 and further specifies that the input signal includes audio data such as voice data.

[0066]

[0077] Example 3 includes the subject matter of Example 1 or Example 2, and further specifies that extraction includes the use of at least one convolutional layer.

[0067]

[0078] Example 4 includes the subject matter of any one of Examples 1 to 3, and further specifies that the input signal includes time / frequency spectrum data.

[0068]

[0079] Example 5 includes the subject matter of Example 4, and further specifies that the time / frequency spectrum data is obtained using the short-time Fourier transform of a time window of the input signal.

[0069]

[0080] Example 6 includes the subject matter of Example 4 or Example 5, and further specifies that amplitude compression is applied to the time / frequency spectrum data.

[0070]

[0081] Example 7 includes the subject matter of Example 6, and further specifies that amplitude compression is applied using values determined during the training of the encoder.

[0071]

[0082] Example 8 includes the subject matter of Example 7, and further specifies that the values are different at various encoding bitrates.

[0072]

[0083] Example 9 includes the subject matter of any one of Examples 1 to 8, and further specifies that the encoder includes a plurality of convolutional layers.

[0073]

[0084] Example 10 includes the subject matter of any one of Examples 1 to 9, and further specifies that determining the prediction includes processing the reconstructed latent features for a plurality of previous frames using a plurality of convolutional layers.

[0074]

[0085] Example 11 includes any one of the themes of Examples 1 to 10, and further specifies that quantizing a feature such as a residual includes dividing the feature such as a residual into a plurality of groups along the channel dimension and separately quantizing each of the plurality of groups.

[0075]

[0086] Example 12 includes the theme of Example 11 and further specifies that a given group among the plurality of groups includes a plurality of frequencies.

[0076]

[0087] Example 13 includes the theme of Example 12 and further specifies that the channel is quantized using various codebooks. For a set of input training data used during the training of the encoder, one of the plurality of groups is randomly selected, and each group is associated with a set of bitrates that gradually increase. While training the encoder using the set of input training data, only the selected group among the plurality of groups and a certain group among the plurality of groups associated with a lower bitrate than the selected group are used.

[0077]

[0088] Example 14 includes any one of the themes of Examples 1 to 13, and further specifies that quantizing a feature such as a residual includes determining a distance between the feature such as a residual and the codeword of the codebook used for vector quantization of the feature such as a residual in a frame, and determining a probability of selecting at least partially the codeword using this distance.

[0078]

[0089] Example 15 includes the theme of Example 14 and further specifies that this probability is determined as a non - linear projection.

[0079]

[0090] Example 16 includes the theme of Example 14 or Example 15 and further specifies that determining the probability includes selecting an element of a Gumbel distribution.

[0080]

[0091] Example 17 includes any one of the themes of Examples 1 to 16, and further specifies that features such as residuals, or data sufficient to reconstruct features such as residuals, are transmitted as part of a bitstream having a rate. During the training of the encoder, a bitrate is determined for the training input data, where determining the bitrate includes determining the difference between a target bitrate and the entropy of the probability of selecting a specific codeword of the codebook in the frame of the training input data.

[0081]

[0092] Example 18 includes the theme of Example 17 and further specifies optimizing a rate-distortion coefficient determined as a trade-off between the determined distortion and the bitrate in the training input data.

[0082]

[0093] Example 19, when executed, is one or more computer-readable media storing computer-executable instructions that cause a computing system to perform various operations. These operations include using an encoder to extract one or more latent features from a frame of an input signal, resulting in one or more extracted latent features. The prediction of one or more latent features is determined using the reconstructed latent features in a plurality of previous frames. From the one or more extracted latent features and the prediction, features such as residuals are extracted. Features such as residuals, or data sufficient to reconstruct features such as residuals, are transmitted to a client. Additional examples include the theme of Example 19 in the form of computer-executable instructions, as well as any one of the themes of Examples 2 to 18 and Examples 27 to 31.

[0083]

[0094] Example 20 is a method that can be implemented in hardware, software, or a combination thereof. An encoder is used to extract one or more latent features from a frame of an input signal, resulting in one or more extracted latent features. The prediction of the one or more latent features is determined using the reconstructed latent features in multiple previous frames. From the one or more extracted latent features and the prediction, features such as residuals are extracted. Features such as residuals, or information sufficient to reconstruct features such as residuals, are transmitted to a client. Additional examples include the subject matter of Example 20 in the form of additional elements of this method, as well as the subject matter of any of Examples 2-18 and Examples 27-31.

[0084]

[0095] Example 21 is a computing system comprising at least one memory and at least one hardware processor coupled to the at least one memory. The computing system further includes one or more computer-readable storage media that store computer-executable instructions that, when executed, cause the computing system to perform various operations. Each operation includes receiving features such as residuals, or data sufficient to reconstruct features such as residuals. The prediction of one or more latent values is determined using the reconstructed latent features in multiple previous frames. The prediction and features such as residuals are combined to result in one or more reconstructed latent features in a frame of the input signal. The one or more reconstructed latent features are provided to a decoder to result in a decoded output signal.

[0085]

[0096] Example 22 includes the subject matter of Example 21 and further specifies that the output signal includes audio data such as voice data.

[0086]

[0097] Example 23 includes the subject matter of Example 21 or Example 22 and further specifies that the decoder includes a plurality of convolutional layers.

[0087]

[0098] Example 24 includes any of the themes of Examples 21-23, and further specifies that determining the prediction includes processing the reconstructed latent features for a plurality of previous frames using a plurality of convolutional layers.

[0088]

[0099] Example 25 includes any of the themes of Examples 21-24, and further specifies that data sufficient to reconstruct features such as residuals includes quantization indices to a codebook used for dequantization.

[0089]

[0100] Example 26 includes the theme of Example 25, and further specifies that the quantization index is received in a bitstream.

[0090]

[0101] Example 27 includes any of the themes of Examples 1-18, and further includes quantizing features such as residuals.

[0091]

[0102] Example 28 includes the theme of Example 27, and quantizing features such as residuals provides quantization indices to a codebook used in quantization.

[0092]

[0103] Example 29 includes the theme of Example 28, and further includes encoding the quantization index into a bitstream.

[0093]

[0104] Example 30 includes the theme of Example 29, and further specifies that the encoding is entropy encoding.

[0094]

[0105] Example 31 includes the theme of Example 30, and further specifies that this entropy encoding is Huffman encoding.

[0095] Example 16 - Computing System

[0106] FIG. 11 shows a generalized example of a suitable computing system 1100 that can implement the described technological innovation. Since the technological innovation may be implemented in various general-purpose or special-purpose computing systems, computing system 1100 does not imply any limitation with respect to the scope of use or functionality of the present disclosure.

[0096]

[0107] Referring to FIG. 11, computing system 1100 includes one or more processing devices 1110, 1115, and memories 1120, 1125. In FIG. 11, this basic configuration 1130 is included within the dashed lines. Processing devices 1110, 1115 execute computer-executable instructions, such as to implement the features described in Examples 1-15. The processing device can be a general-purpose central processing unit (CPU), a processor within an application-specific integrated circuit (ASIC), or any other type of processor. In a multiprocessing system, multiple processing devices execute computer-executable instructions to increase processing power. For example, FIG. 11 shows a central processing unit 1110, as well as a graphics processing unit or coprocessing unit 1115. The tangible memories 1120, 1125 can be volatile memories (e.g., registers, caches, RAM), non-volatile memories (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two, accessible by processing devices 1110, 1115. Memories 1120, 1125 store software 1180 that implements one or more of the technological innovations described herein in the form of computer-executable instructions suitable for execution by processing devices 1110, 1115.

[0097]

[0108] Computing system 1100 may have additional features. For example, computing system 1100 includes a storage device 1140, one or more input devices 1150, one or more output devices 1160, and one or more communication connections 1170, which include input devices, output devices, and communication connections for interacting with a user. An interconnect mechanism (not shown), such as a bus, a control device, or a network, interconnects the components of computing system 1100. Typically, an operating system software (not shown) provides an operating environment for other software to execute in computing system 1100 and coordinates the operation of the components of computing system 1100.

[0098]

[0109] The tangible storage device 1140 may be removable or non-removable and may include a magnetic disk, magnetic tape or cassette, CD-ROM, DVD, or any other medium that can be used to store information in a persistent manner and that can be accessed within computing system 1100. The storage device 1140 stores instructions for software 1180 that implements one or more of the technological innovations described herein.

[0099]

[0110] The input device 1150 may be a keyboard, mouse, pen, or touch input device such as a trackball, voice input device, scanning device, or any other device that supplies input to computing system 1100. The output device 1160 may be a display device, printer, speaker, CD writer, or any other device that supplies output from computing system 1100.

[0100]

[0111] The communication connection unit 1170 enables communication with another computing entity via a communication medium. The communication medium transmits information such as computer-executable instructions, audio or video input or output, or other data within a modulated data signal. A modulated data signal is a signal having one or more of the signal's characteristics set or changed to encode information within the signal. By way of example and without limitation, the communication medium can use electricity, light, RF, or other carrier waves.

[0101]

[0112] In the general context of computer-executable instructions, such as instructions included in program modules that are executed on a target physical or virtual processor in a computing system, the technological innovation can be described. Generally, a program module or program component includes routines, programs, libraries, classes, objects, components, data structures, etc., which perform specific tasks or implement specific abstract data types. As desired in various embodiments, the functions of program modules may be combined or divided among program modules. The computer-executable instructions for program modules may be executed within a local computing system or a distributed computing system.

[0102]

[0113] The terms "system" and "apparatus" are used interchangeably herein. Unless the context clearly indicates otherwise, neither term implies any limitation to a type of computing system or computing apparatus. Generally, a computing system or computing apparatus can be local or distributed and can include any combination of dedicated hardware and / or general-purpose hardware and software implementing the functions described herein.

[0103]

[0114] In various examples described herein, a module (e.g., a component or an engine) can be “encoded” to perform a particular operation or realize a particular function, which indicates that computer-executable instructions for the module can be executed to perform such an operation, cause such an operation to be performed, or otherwise realize such a function. The functions described for software components, modules, or engines can be executed as individual software units (e.g., programs, functions, class methods), but need not be implemented as individual units. That is, the functions can be incorporated into a larger program, or a more general-purpose program, such as one line or multiple lines of code in a larger program or a more general-purpose program.

[0104]

[0115] For purposes of explanation, in the detailed description, terms such as “determine” and “use” are used to describe computer operations in a computing system. These terms are high-level abstractions of operations performed by a computer and should not be confused with acts performed by a human. The actual computer operations corresponding to these terms vary depending on the implementation.

[0105] Example 17 - Cloud Computing Environment

[0116] FIG. 12 illustrates an exemplary cloud computing environment 1200 that can implement the technology described. The cloud computing environment 1200 includes cloud computing services 1210. The cloud computing services 1210 can include various types of cloud computing resources such as computer servers, repositories of data storage devices, networking resources, etc. The cloud computing services 1210 can be centrally located (e.g., provided by a corporate or organizational data center) or distributed (e.g., provided by various computing resources located in different locations such as different data centers and / or located in different cities or countries).

[0106]

[0117] The cloud computing services 1210 are utilized by various types of computing devices (e.g., client computing devices) such as computing devices 1220, 1222, 1224. For example, the computing devices (e.g., 1220, 1222, and 1224) can be a computer (e.g., a desktop computer or a laptop computer), a mobile device (e.g., a tablet computer or a smartphone), or other types of computing devices. For example, the computing devices (e.g., 1220, 1222, and 1224) can utilize the cloud computing services 1210 to perform computing operations (e.g., data processing, data storage, etc.).

[0107] Example 18 - Implementation Form

[0118] Some operations of the disclosed methods are described in a specific order for convenience of explanation, but it should be understood that this manner of explanation encompasses rearrangement unless a particular ordering is required by the specific language described herein. For example, operations described sequentially may, in some cases, be rearranged or performed simultaneously. Further, for the sake of brevity, the accompanying figures may not show various ways in which the disclosed methods can be used in conjunction with other methods.

[0108]

[0119] Any of the disclosed methods can be implemented as computer-executable instructions or a computer program product stored on one or more computer-readable storage media and executed by a computing device (any available computing device, including, for example, a smartphone or other mobile device equipped with computing hardware). A tangible computer-readable storage media is any available tangible media that can be accessed within a computing environment (e.g., one or more optical media disks such as a DVD or CD, components of volatile memory (such as DRAM or SRAM), or components of non-volatile memory (such as flash memory or a hard drive)). By way of example, referring to FIG. 11, computer-readable storage media includes memories 1120 and 1125, as well as storage device 1140. The term computer-readable storage media does not include signals and carrier waves. Further, the term computer-readable storage media does not include a communication connection (e.g., 1170).

[0109]

[0120] Any of the computer-executable instructions for implementing the disclosed techniques, as well as any data created and used during the implementation of the disclosed embodiments, can be stored on one or more computer-readable storage media. The computer-executable instructions can be, for example, a dedicated software application, or a portion of a software application accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer (such as any suitable commercially available computer), or using one or more network computers, in a network environment (such as via the Internet, a wide area network, a local area network, a client / server network such as a cloud computing network or other such network).

[0110]

[0121] For clarity, only certain selected aspects of the software-based implementation forms are described. It should be understood that the disclosed technology is not limited to any particular computer language or computer program. For example, the disclosed technology can be implemented by software written in C++, Java, Perl, JavaScript, Python, Ruby, ABAP, SQL, Adobe Flash, or any other suitable programming language, or, by way of example, a markup language such as html or XML, or software written in a combination of a suitable programming language and a markup language. Similarly, the disclosed technology is not limited to any particular type of computer or hardware.

[0111]

[0122] Furthermore, any of the software-based embodiments (including computer-executable instructions for causing a computer to execute any of the disclosed methods) can be uploaded, downloaded, or remotely accessed via appropriate communication means. Such appropriate communication means include, for example, the Internet, the World Wide Web, an intranet, a software application, a cable (including an optical fiber cable), magnetic communication, electromagnetic communication (including RF, microwave, and infrared communication), electronic communication, or other such communication means.

[0112]

[0123] The disclosed methods, apparatuses, and systems should in no way be construed as limiting. Instead, the present disclosure is directed to any novel and non-obvious features and aspects of the various disclosed embodiments, alone and in various combinations and sub-combinations with each other. The disclosed methods, apparatuses, and systems are not limited to any particular aspect or feature, or combination thereof, and the disclosed embodiments do not require the presence of one or more particular advantages or the solving of any particular problems.

[0113]

[0124] The techniques from any example can be combined with the techniques described in any one or more of the other examples. In view of the many possible embodiments to which the principles of the disclosed techniques may be applied, it should be recognized that the illustrated embodiments are examples of the disclosed techniques and should not be construed as limiting the scope of the disclosed techniques. Rather, the scope of the disclosed techniques includes what is encompassed by the scope and spirit of the appended claims.

Claims

1. At least one hardware processor, At least one memory coupled to the at least one hardware processor, When executed, Using an encoder, extracting one or more latent features from a frame of an input signal to yield the one or more extracted latent features, Determining a prediction of the one or more latent features using the reconstructed latent features in a plurality of previous frames, Extracting features such as residuals from the one or more extracted latent features and the prediction, and Transmitting to a client the features such as residuals, or data sufficient to reconstruct the features such as residuals One or more computer-readable storage media including computer-executable instructions that cause a computing system to perform operations including A computing system comprising.

2. The computing system according to claim 1, wherein the input signal includes audio data.

3. The computing system according to claim 1, wherein the extracting includes using at least one convolutional layer.

4. The computing system according to claim 1, wherein the input signal includes time / frequency spectrum data.

5. The computing system according to claim 4, wherein the time / frequency spectrum data is obtained using a short-time Fourier transform of a time window of the input signal.

6. The computing system according to claim 4, further comprising applying amplitude compression to the time / frequency spectrum data.

7. The computing system according to claim 6, wherein the amplitude compression is applied using a value determined during training of the encoder.

8. The computing system according to claim 7, wherein the value varies at various encoding bitrates.

9. The computing system according to claim 1, wherein the encoder comprises a plurality of convolutional layers.

10. The computing system according to claim 1, wherein determining the prediction includes processing the reconstructed latent features in the plurality of previous frames using a plurality of convolutional layers.

11. The computing system according to claim 1, wherein the operation further includes dividing features such as the residual along a channel dimension into a plurality of groups, and separately quantizing the groups of the plurality of groups.

12. The computing system according to claim 11, wherein a given group of the plurality of groups includes a plurality of frequencies.

13. The channel is quantized using various codebooks, and the operation is during the training of the encoder, randomly selecting one of the plurality of groups in a set of input training data used during the training of the encoder, wherein the group is associated with a set with gradually increasing bitrate; while training the encoder using the set of input training data, using only the selected group of the plurality of groups and the groups of the plurality of groups associated with a lower bitrate than the selected group The computing system according to claim 12, further including.

14. The operation is quantizing features such as the residual, for the frame, determining a distance between features such as the residual and a codeword of a codebook used for vector quantization of the features such as the residual; determining a probability of selecting the codeword at least partially using the distance, for quantizing the features such as the residual The computing system according to claim 1, further including.

15. The computing system according to claim 14, wherein determining the probability is determined as a non-linear projection.

16. The computing system according to claim 14, wherein determining the probability includes selecting an element of a Gumbel distribution.

17. Features such as the residual, or the data sufficient to reconstruct the features such as the residual, are transmitted as part of a bitstream having a rate, and the operation is During the training of the encoder, determining the bit rate in the training input data, including determining the difference between the target bit rate and the entropy of the probability of selecting a specific codeword of the codebook in the frame of the training input data. Determining the bit rate in the training input data The computing system according to claim 1, further comprising

18. The computing system according to claim 17, wherein the operation further includes optimizing a rate-distortion coefficient determined as a trade-off between the determined distortion and the bit rate in the training input data.

19. A method implemented in a computing system comprising at least one hardware processor and at least one memory coupled to the at least one hardware processor, comprising: Using an encoder to extract one or more latent features from a frame of an input signal, resulting in the extracted one or more latent features; Using the reconstructed latent features in a plurality of previous frames to determine a prediction of the one or more latent features; Extracting features such as residuals from the extracted one or more latent features and the prediction; Transmitting to the client the features such as the residuals or data sufficient to reconstruct the features such as the residuals. A method comprising

20. Computer-executable instructions that, when executed by a computing system comprising at least one hardware processor and at least one memory coupled to the at least one hardware processor, cause the encoder to extract one or more latent features from a frame of an input signal, resulting in the extracted one or more latent features; Computer-executable instructions that, when executed by the computing system, cause the computing system to determine a prediction of the one or more latent features using the reconstructed latent features in a plurality of previous frames; Computer-executable instructions that, when executed by the computing system, cause the computing system to extract features such as residuals from the extracted one or more latent features and the prediction; When executed by the computing system, computing-executable instructions that cause the computing system to send to the client features such as the residual, or data sufficient to reconstruct features such as the residual One or more computer-readable storage media including