Compressing Audio Waveforms Using Neural Networks and Vector Quantizers
The neural network-based audio compression/decompression system with a sequence of vector quantizers efficiently compresses and decompresses audio waveforms, addressing inefficiencies in existing systems by optimizing parameters and codebooks for flexible bit rates and reduced latency.
Patent Information
- Application Number
- JP2023579832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2022-07-05
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-07-05
AI Technical Summary
Existing audio compression systems are inefficient in terms of computational resources and latency, particularly when performing noise reduction and compression/decompression tasks, and often require separate modules for these functions, leading to increased latency.
A neural network-based compression/decompression system using a sequence of vector quantizers and a decoder neural network, trained end-to-end, to efficiently compress and decompress audio waveforms while performing noise reduction, reducing the need for separate modules and minimizing latency.
The system achieves efficient audio data compression and transmission by optimizing neural network parameters and codebooks, allowing for flexible bit rates and adaptable performance, reducing computational resources and latency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Patent Application No. 17 / 856,856, filed Jul. 1, 2022, which claims the benefit of priority of U.S. Provisional Application No. 63 / 218,139, filed Jul. 2, 2021. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated herein by reference.
[0002] This specification relates to processing data using a machine learning model.
Background Art
[0003] A machine learning model receives an input and generates an output, such as a predicted output, based on the received input. Some machine learning models are parametric models and generate an output based on the received input and the values of the model's parameters.
[0004] Some machine learning models are deep models that employ multiple layers of the model to generate an output for the received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non - linear transformation to the received input to generate the output.
Prior Art Documents
Non - Patent Documents
[0005]
Non - Patent Document 1
Non-Patent Document 2
Summary of the Invention
Means for Solving the Problems
[0006] This specification generally describes a compression system implemented as a computer program on one or more computers in one or more locations that can compress an audio waveform. This specification further describes a restoration system implemented as a computer program on one or more computers in one or more locations that can restore the audio waveform.
[0007] Generally, the compression system and the restoration system can be located at any suitable location. In particular, the compression system can, in some cases, be located remotely from the restoration system. For example, the compression system can be implemented by one or more first computers at a first location, while the restoration system can be implemented by one or more second (different) computers at a second (different) location.
[0008] In some implementations, the compression system can generate a compressed representation of the input audio waveform and store the compressed representation in a data store, such as a logical data storage area or a physical data storage device. The restoration system can later access the compressed representation from the data store and process the compressed representation to generate the corresponding output audio waveform. The output audio waveform can be, for example, a reconstruction of the input audio waveform or an enhanced (e.g., noise-reduced) version of the input audio waveform.
[0009] In some implementations, the compression system can generate a compressed representation of the input audio waveform and transmit the compressed representation to a destination via a data communication network, such as a local area network, a wide area network, or the Internet. The restoration system can access the compressed representation at the destination and process the compressed representation to generate the corresponding output waveform.
[0010] According to a first aspect, a method is provided that is executed by one or more computers, the method comprising receiving an audio waveform comprising respective audio samples for each of a plurality of time steps; processing the audio waveform using an encoder neural network to generate a plurality of feature vectors representing the audio waveform; using a plurality of vector quantizers each associated with a respective codebook of code vectors to generate a respective coded representation of each of the plurality of feature vectors, wherein the respective coded representation of each feature vector comprises respective code vectors from the codebook of the vector quantizer that defines the quantized representation of the feature vector, identifying a plurality of code vectors; and generating a compressed representation of the audio waveform by compressing the respective coded representation of each of the plurality of feature vectors.
[0011] In some implementations, for each of a plurality of vector quantizers ordered in a sequence, the step of generating a coded representation of a feature vector includes, for a first vector quantizer in the sequence of vector quantizers, receiving the feature vector; identifying, based on the feature vector, each code vector from a codebook of the vector quantizer for representing the feature vector; and determining a current residual vector based on an error between (i) the feature vector and (ii) the code vector representing the feature vector, wherein the coded representation of the feature vector identifies the code vector representing the feature vector.
[0012] In some implementations, for each of a plurality of feature vectors, the step of generating a coded representation of the feature vector includes, for each vector quantizer after the first vector quantizer in the sequence of vector quantizers, receiving the current residual vector generated by the previous vector quantizer in the sequence of vector quantizers; identifying, based on the current residual vector, each code vector from a codebook of the vector quantizer for representing the current residual vector; and, if the vector quantizer is not the last vector quantizer in the sequence of vector quantizers, further including updating the current residual vector based on an error between (i) the current residual vector and (ii) the code vector representing the current residual vector, wherein the coded representation of the feature vector identifies the code vector representing the current residual vector.
[0013] In some implementations, the step of generating a compressed representation of an audio waveform includes entropy encoding each respective coded representation of each of the plurality of feature vectors.
[0014] In some implementations, each quantized representation of a feature vector is defined by a sum of a plurality of code vectors identified by the coded representation of the feature vector.
[0015] In some implementations, all of the codebooks of the plurality of vector quantizers contain the same number of code vectors.
[0016] In some implementations, an encoder neural network and the codebooks of a plurality of vector quantizers are trained together with a decoder neural network, and the decoder neural network receives, for each of a plurality of feature vectors representing an input audio waveform generated using the encoder neural network and the plurality of vector quantizers, a respective quantized representation of the feature vector, and processes the quantized representation of the feature vector representing the input audio waveform to generate an output audio waveform.
[0017] In some implementations, training includes: obtaining a plurality of training examples each including (i) a respective input audio waveform and (ii) a corresponding target audio waveform; using an encoder neural network, a plurality of vector quantizers from a sequence of vector quantizers, and a decoder neural network to process each input audio waveform from each training example to generate an output audio waveform that is an estimate of the corresponding target audio waveform; determining a gradient of an objective function that depends on each output and target waveform for each training example; and updating one or more of a set of encoder neural network parameters, a set of decoder neural network parameters, or the codebooks of the plurality of vector quantizers using the gradient of the objective function.
[0018] In some implementations, for one or more of the training examples, the target audio waveform is a Emphasis modified version of the input audio waveform.
[0019] In some implementations, for one or more of the training examples, the target audio waveform is a noise-removed version of the input audio waveform.
[0020] In some implementations, for one or more of the training examples, the target audio waveform is the same as the input audio waveform.
[0021] In some implementations, the step of processing each input audio waveform to generate a corresponding output audio waveform includes conditioning an encoder neural network, a decoder neural network, or both on data that defines whether the corresponding target audio waveform is (i) the input audio waveform or (ii) a processed version of the input audio waveform. Emphasis is a processed version.
[0022] In some implementations, the method further includes, for each training example, selecting the respective number of vector quantizers to be used when quantizing the feature vector representing the input audio waveform, and generating the corresponding output audio waveform using only the selected number of vector quantizers from the sequence of vector quantizers.
[0023] In some implementations, the selected number of vector quantizers to be used when quantizing the feature vector representing the input audio waveform varies among the training examples.
[0024] In some implementations, for each training example, the step of selecting the respective number of vector quantizers to be used when quantizing the feature vector representing the input audio waveform includes randomly sampling the number of vector quantizers to be used when quantizing the feature vector representing the input audio waveform.
[0025] In some implementations, the objective function comprises a reconstruction loss that measures, for each training example, the error between (i) the output audio waveform and (ii) the corresponding target audio waveform.
[0026] In some implementations, for each training example, the reconstruction loss measures the multi-scale spectral error between (i) the output audio waveform and (ii) the corresponding target audio waveform.
[0027] In some implementations, training comprises, for each training example, using a discriminator neural network to process data derived from the output audio waveform to generate a set of one or more discriminator scores, wherein each discriminator score characterizes an estimated likelihood that the output audio waveform is an audio waveform generated using an encoder neural network, a plurality of vector quantizers, and a decoder neural network, and the objective function comprises an adversarial loss that depends on the discriminator scores generated by the discriminator neural network.
[0028] In some implementations, the data derived from the output audio waveform comprises the output audio waveform, a downsampled version of the output audio waveform, or a Fourier-transformed version of the output audio waveform.
[0029] In some implementations, for each training example, the reconstruction loss measures the error between (i) one or more intermediate outputs generated by a discriminator neural network by processing the output audio waveform and (ii) one or more intermediate outputs generated by the discriminator neural network by processing the corresponding target audio waveform.
[0030] In some implementations, the codebooks of multiple vector quantizers are repeatedly updated during training using the exponential moving average of the feature vectors generated by the encoder neural network.
[0031] In some implementations, the encoder neural network comprises a sequence of encoder blocks, each configured to process each set of input feature vectors according to a set of encoder block parameters to generate a set of output feature vectors having a lower temporal resolution than the set of input feature vectors.
[0032] In some implementations, the decoder neural network comprises a sequence of decoder blocks, each configured to process each set of input feature vectors according to a set of decoder block parameters to generate a set of output feature vectors having a higher temporal resolution than the set of input feature vectors.
[0033] In some implementations, the audio waveform is a speech waveform or a music waveform.
[0034] In some implementations, the method further includes transmitting a compressed representation of the audio waveform over a network.
[0035] According to another aspect, a method is provided that is executed by one or more computers. The method includes receiving a compressed representation of an input audio waveform; restoring the compressed representation of the input audio waveform to obtain a respective coded representation of each of a plurality of feature vectors representing the input audio waveform, wherein the coded representation of each feature vector includes respective code vectors from respective codebooks of a plurality of vector quantizers that define a quantized representation of the feature vector, and identifying a plurality of code vectors; generating a respective quantized representation of each feature vector from the coded representation of the feature vector; and using a decoder neural network to process the quantized representation of the feature vector to generate an output audio waveform.
[0036] According to another aspect, a system is provided that includes one or more computers and one or more storage devices communicatively coupled to the one or more computers. The one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of each method described herein.
[0037] According to another aspect, one or more non-transitory computer storage media storing instructions are provided that, when executed by one or more computers, cause the one or more computers to perform the operations of each method described herein.
[0038] The subject matter described herein may be implemented in certain embodiments so as to realize one or more of the following advantages.
[0039] The compression / decompression system described herein can enable audio data to be compressed more efficiently than some conventional systems. By enabling more efficient audio data compression, the system enables more efficient audio data transmission (e.g., by reducing the communication network bandwidth required to transmit the audio data) and more efficient audio data storage (e.g., by reducing the amount of memory required to store the audio data).
[0040] The compression / decompression system includes an encoder neural network, a set of vector quantizers, and a decoder neural network that are trained together (i.e., “end-to-end”). By training the neural network parameters of each of the encoder and decoder neural networks together with the codebook of the vector quantizer, the parameters of the compression / decompression system can be harmonized and adapted to achieve more efficient audio compression than would otherwise be possible. For example, as the neural network parameters of the encoder neural network are adjusted iteratively, the codebook of the vector quantizer is simultaneously optimized to enable more accurate quantization of the feature vectors generated by the encoder neural network. The neural network parameters of the decoder neural network are also simultaneously optimized to enable more accurate reconstruction of the audio waveform from the quantized feature vectors generated using the updated codebook of the vector quantizer.
[0041] Performing vector quantization of the feature vectors representing the audio waveform using a single vector quantizer, where each feature vector is represented using r bits, is of size 2 rA codebook may be required. That is, the size of the codebook of the vector quantizer can increase exponentially along with the number of bits allocated to represent each feature vector. As the number of bits allocated to represent each feature vector increases, learning and storing the codebook becomes computationally infeasible. To address this problem, the compression / decompression system performs vector quantization using a sequence of multiple vector quantizers, each maintaining its own codebook. The first vector quantizer can directly quantize the feature vectors generated by the encoder neural network, while each subsequent vector quantizer can quantize the residual vectors that define the quantization error generated by the previous vector quantizer.
[0042] The sequence of vector quantizers can iteratively improve the quantization of the feature vectors while each maintaining a dramatically smaller codebook than would be required by a single vector quantizer. For example, each vector quantizer can maintain a codebook of size
[0043]
Number
[0044] where r is the number of bits allocated to represent each feature vector and N q is the number of vector quantizers. Thus, by performing vector quantization using a sequence of multiple vector quantizers, it becomes possible for the compression / decompression system to reduce the memory required to store the quantizer codebooks, and for vector quantization to be performed in situations where it would otherwise be computationally infeasible.
[0045] By performing vector quantization using a set of multiple vector quantizers (i.e., rather than a single vector quantizer), the compression / decompression system can also control the compression bit rate, e.g., the number of bits used to represent each second of audio data. To reduce the bit rate, the compression / decompression system can perform vector quantization using fewer vector quantizers. Conversely, to increase the bit rate, the compression / decompression system can perform vector quantization using more vector quantizers. During training, the number of vector quantizers used for compression / decompression of each audio waveform is varied (e.g., randomly) over the training examples, allowing the compression / decompression system to learn a single set of parameter values that enable effective compression / decompression over a range of possible bit rates. Thus, the compression / decompression system can reduce the consumption of computational resources by eliminating any requirement to train and maintain multiple respective encoders, decoders, and vector quantizers that are each optimized for each respective bit rate.
[0046] The compression / decompression system can be trained to perform both audio data compression and audio data Emphasis , e.g., noise removal, together. That is, the compression and decompression system can be trained to simultaneously Emphasis process (e.g., remove noise from) the audio waveform as part of compressing and decompressing the waveform without increasing the overall latency. In contrast, some conventional systems apply separate audio Emphasis algorithms to the audio waveform either on the transmitter side (i.e., before compression) or on the receiver side (i.e., after decompression), which can result in an increase in latency.
[0047] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0048]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Best Mode for Carrying Out the Invention
[0049] Like reference numerals and designations in the various drawings indicate like elements.
[0050] FIG. 1 shows an exemplary audio compression system 100 that can compress an audio waveform using an encoder neural network 102 and a residual vector quantizer 106. Similarly, FIG. 2 shows an exemplary audio restoration system 200 that can restore a compressed audio waveform using a decoder neural network 104 and a residual vector quantizer 106. For clarity, both FIGS. 1 and 2 are referred to when describing the various components involved in compression and restoration. The audio compression / restoration system 100 / 200 is an example of a system implemented as a computer program on one or more computers at one or more locations where the systems, components, and techniques described hereinafter are implemented.
[0051] The compression / decompression system 100 / 200 utilizes a neural network architecture (neural codec) that can be superior to legacy codecs, such as waveform and parametric codecs, with respect to the operating bitrate and general audio compression. For comparison, waveform codecs typically use time / frequency domain conversion for compression / decompression and make few or no assumptions about the source audio. As a result, waveform codecs produce high-quality audio at medium to high bitrates but tend to introduce coding artifacts at low bitrates. Parametric codecs can overcome this problem by making specific assumptions about the source audio (e.g., speech), but are not effective for general audio compression. In contrast, the audio compression and decompression system 100 / 200 can compress and decompress speech, music, and general audio at the target bitrate typically targeted by speech-adapted codecs (e.g., parametric codecs). Thus, the audio compression / decompression system 100 / 200 can operate in modalities that are not possible with conventional codecs. In some implementations, the compression and decompression system 100 / 200 can be configured for a specific type of audio content (e.g., speech) due to the flexibility afforded by the neural network architecture.
[0052] Referring to FIG. 1, compression system 100 receives an audio waveform 112 to be compressed. Waveform 112 can include audio samples at each time step, where the time step typically corresponds to a particular sampling rate. A higher sampling rate captures higher frequency components of audio waveform 112. For example, a standard audio sampling rate used by professional digital equipment is 48 kHz because professional digital equipment can reconstruct sounds at frequencies up to 24 kHz (e.g., the human upper audible limit). Such a sampling rate is ideal for comfortable listening, but compression system 100 can generally be configured to process waveform 112 at any sampling rate, even with waveforms involving non-uniform sampling.
[0053] Audio waveform 112 can be generated from any suitable audio source. For example, waveform 112 can be a recording from an external audio device (e.g., voice from a microphone), a purely digital work (e.g., electronic music), or common audio such as sound effects and background noise (e.g., white noise, room tone). In some implementations, audio compression system 100 can, while compressing waveform 112, simultaneously Emphasis perform operations such as suppressing unwanted background noise.
[0054] In a conventional audio processing pipeline, compression and Emphasis are typically performed by separate modules. For example, audio EmphasisThe algorithm can be applied at the input of the encoder 102 before the waveform is compressed or at the output of the decoder 104 after the waveform is restored. In this setup, for example, each processing step contributes to the end-to-end latency due to buffering the waveform to the expected frame length of the algorithm. Conversely, by leveraging the thoughtful training of various neural network components (see Figure 3), the audio compression system 100 can achieve joint compression and Emphasis without relying on separate modules and without incurring additional latency. In some implementations, the audio restoration system 200 implements joint restoration and Emphasis . Generally, the compression system 100, the restoration system 200, or both can be designed to perform audio Emphasis by adequately training neural network components.
[0055] The audio waveform 112 is processed (e.g., encoded) by the encoder 102 to generate a sequence of feature vectors 208 representing the waveform 112. The feature vectors 208 (e.g., embeddings, latent representations) are a compressed representation of the waveform that extracts the most relevant information about the audio content of the waveform. The encoder 102 can downsample the input waveform 202 to generate the compressed feature vectors 208 such that the feature vectors 208 have a lower sampling rate than the original audio waveform 112. For example, the encoder neural network 102 can use multiple convolutional layers with increased stride to generate the feature vectors 208 at a lower sampling rate (e.g., lower temporal resolution).
[0056] Next, the feature vector 208 is processed by a residual (e.g., multi-stage) vector quantizer RVQ106 to generate a coded representation (CFV) 210 of the feature vector and a corresponding quantized representation (QFV) 212 of the feature vector. RVQ106 can generate QFV212 at a specific bit rate by leveraging one or more vector quantizers 108. RVQ106 achieves (irreversible) compression by mapping the higher-dimensional space of the feature vector 208 to the individual subspaces of the code vectors. As detailed below, CFV210 specifies codewords (e.g., indices) from the respective codebooks 110 of each vector quantizer 108, where each codeword identifies the code vector stored in the associated codebook 110. Thus, QFV212 is an approximation of the feature vector 208 defined by the combination of code vectors specified by the corresponding CFV210. Generally, QFV212 is the sum (e.g., linear combination) of the code vectors specified by CFV210.
[0057] In some cases, RVQ106 uses a single vector quantizer 108 with a single codebook 110. The quantizer 108 can compress the feature vector 208 into QFV212 by selecting a code vector in its codebook 110 to represent the feature vector 208. The quantizer 106 can select a code vector based on any suitable distance metric (e.g., error) between two vectors, e.g., the L n norm, cosine distance, etc. For example, RVQ106 can select a code vector with the smallest Euclidean norm (e.g., the L 2 norm) for the feature vector 208. Then, the quantizer 106 can store the corresponding codeword in CFV210. Since codewords generally require fewer bits than code vectors, CFV210 consumes less space in memory and can achieve greater compression than QFV212 without additional loss.
[0058] Nevertheless, the approach of a single vector quantizer 108 can be extremely expensive because as the bit rate is increased, the size of the codebook 110 increases exponentially. To overcome this problem, RVQ106 can utilize a sequence of vector quantizers 108. In this case, each vector quantizer 108 in the sequence includes its own codebook 110 of code vectors. Then, RVQ106 can use an iterative method to generate CFV210 and the corresponding QFV212 such that each vector quantizer 108 in the sequence further improves the quantization.
[0059] For example, in the first vector quantizer 108, the quantizer 106 can receive the feature vector 208 and select a code vector from its codebook 110 to represent the feature vector 208 based on the smallest distance metric. A residual vector can be calculated as the difference between the feature vector 208 and the code vector representing the feature vector 208. The residual vector can be received by the next vector quantizer 108 in the sequence to select a code vector from its codebook 110 to represent the residual vector based on the smallest distance metric. The difference between these two vectors can be used as the residual vector for the next iteration. This iterative method can continue for each vector quantizer 108 in the sequence. Any code vectors identified in this way can be summed into QFV212 and the codeword of each code vector can be stored in the respective CFV210.
[0060] Generally, RVQ106 can utilize any suitable number of vector quantizers 108. The number of quantizers N q and the size of each codebook N icontrols the trade-off between computational complexity and coding efficiency. Thus, the sequence of quantizers 108 provides a flexible means of balancing these two opposing factors. In some cases, the size N of each codebook is such that the entire bit budget is uniformly allocated across each vector quantizer 108 i = N is equal. The uniform allocation provides practical modularity to RVQ106 since each codebook 110 consumes the same amount of space in memory.
[0061] Moreover, for a fixed-size codebook N i the number of vector quantizers N in the sequence q determines the resulting bitrate of QFV212, where a higher bitrate corresponds to a larger number of quantizers 108. Thus, RVQ106 provides a convenient framework for variable (e.g., scalable) bitrates by employing structured dropout of quantizers 108. That is, the audio compression and restoration system 100 / 200 can vary the number of quantizers 108 in the sequence to target any desired bitrate, facilitating adjustable performance and reducing the overall memory footprint compared to multiple fixed-bitrate codecs. These capabilities enable the compression / restoration system 100 / 200 to be adaptable to low-latency implementations suitable, in particular, for devices with limited computing resources (e.g., smartphones, tablets, smartwatches, etc.).
[0062] CFV210 can then further compress the compressed representation 114 of the audio waveform, for example, using an entropy codec 302. The entropy codec 302 can implement any suitable reversible entropy coding, such as arithmetic coding, Huffman coding, etc.
[0063] Referring to FIG. 2, the restoration system 200 receives a compressed representation 114 of an audio waveform. Generally, the compressed audio waveform 114 can represent any type of audio content, such as voice, music, general audio, etc. Indeed, as described above, the audio compression and restoration systems 100 / 200 can be implemented for a specific task (e.g., voice-adapted compression / restoration) in order to optimize around a specific type of audio content.
[0064] The compressed audio waveform 114 can be restored to the CFV210, for example, using an entropy codec 302. Then, the CFV210 is processed by the RVQ106 into the QFV212. As described above, each CFV210 includes a codeword (e.g., an index) that identifies a code vector in the respective codebook 110 of each vector quantizer 108. The combination of code vectors specified by each CFV210 identifies the corresponding QFV212. Generally, the code vectors identified by each CFV210 are summed into the corresponding QFV212.
[0065] Then, the QFV212 can be processed (e.g., decoded) by the decoder 104 to generate an audio waveform 112. The decoder 104 generally reproduces the process of the encoder 102 by starting from the (quantized) feature vector and outputting a waveform. The decoder 104 can upsample the QFV212 to generate an output waveform 206 at a higher sampling rate than the input QFV212. For example, the decoder 104 can use a plurality of convolutional layers with reduced stride to generate the output waveform 206 at a higher sampling rate (e.g., higher temporal resolution).
[0066] Note that the compression / decompression system 100 / 200 can be implemented in various different implementation forms, such as being integrated as a single system or as separate systems. Moreover, each component of the compression / decompression system 100 / 200 does not need to be restricted to a single client device. For example, in some implementation forms, the compression system 100 stores the compressed audio waveform 114 in local storage, and then the compressed audio waveform 114 is retrieved from local storage by the decompression system 200. In other implementation forms, the compression system 100 on the transmitter client transmits the compressed audio waveform 114 over a network (such as the Internet, 5G cellular network, Bluetooth, Wi-Fi, etc.), and the compressed audio waveform 114 can be received by the decompression system 200 on the receiver client.
[0067] As will be described in more detail below, the neural network architecture can be trained using the training system 300. The training system 300 can enable efficient general-purpose compression or adapted compression (such as voice-adapted) by utilizing a suitable set of training examples 116 and various training procedures. Specifically, the training system 300 can train the encoder neural network 102 and the decoder neural network 104 together to efficiently encode and decode the feature vectors 208 of various waveforms included in the training examples 116. Furthermore, the training system 300 can train the RVQ 106 to efficiently quantize the feature vectors 208. In particular, each codebook 110 of each cascading vector quantizer 108 can be trained to minimize the quantization error. To facilitate the trainable codebook 110, each vector quantizer 108 can be implemented, for example, as a vector quantized variational autoencoder (VQ-VAE).
[0068] When implementing this data-driven training solution, the audio compression / restore systems 100 / 200 can be a completely "end-to-end" machine learning approach. In the end-to-end implementation form, the compression / restore systems 100 / 200 utilize neural networks for all tasks involved in training and for inference after training. Processes such as feature extraction are not executed by external systems. Generally, the training system 300 can utilize unsupervised learning algorithms, semi-supervised learning algorithms, supervised learning algorithms, or more complex combinations of these. For example, the training system 300 can balance reconstruction loss and adversarial loss to enable audio compression that is both faithful to the original audio in playback and perceptually similar to the original audio.
[0069] Generally, the neural networks included in the audio compression / restore systems 100 / 200 can have any suitable neural network architecture that enables the audio compression / restore systems 100 / 200 to perform the functions they describe. Specifically, the neural network can each include any suitable number (e.g., 5, 10, or 100 layers) of, and be arranged in any suitable configuration (e.g., as a linear sequence of layers), any suitable neural network layers (e.g., fully connected layers, convolutional layers, attention layers, etc.).
[0070] In some implementations, the compression / decompression system 100 / 200 utilizes a fully convolutional neural network architecture. FIGS. 8A and 8B illustrate exemplary implementations of such an architecture for the neural networks of the encoder 102 and decoder 104. The fully convolutional architecture can be particularly advantageous for low-latency compression because it has a lower scale of connectivity compared to fully connected networks (e.g., multi-layer perceptrons) and has filters (e.g., kernels) that can be optimized to limit coding artifacts. Additionally, convolutional neural networks provide an effective means of resampling the waveform 112, i.e., changing the temporal resolution of the waveform 112, by using different strides for different convolutional layers.
[0071] In a further implementation, when implementing the fully convolutional architecture, the compression / decompression system 100 / 200 uses strict causal convolution such that padding is applied only to the past and not to the future in both training and offline inference. Padding is not necessary for streaming inference. In this case, the overall latency of the compression and decompression systems 100 / 200 is completely determined by the temporal resampling ratio between the waveforms and their corresponding feature vectors.
[0072] FIG. 3 shows the operations performed by an exemplary training system 300 to train the encoder neural network 102, the decoder neural network 104, and the residual vector quantizer 106 together. The neural networks are trained end-to-end in an objective function 214 that can include a number of reconstruction losses. In some implementations, a discriminator neural network 216 is also trained to facilitate an adversarial loss 218 and, optionally, additional reconstruction losses.
[0073] The training system 300 receives a set of training examples 116. Each training example 116 includes a respective input audio waveform 202 and a corresponding target audio waveform 204 that the neural network is trained to reconstruct. That is, the target waveform 204 can be compared with the obtained output audio waveform 206 to evaluate the performance of the neural network using the objective function 214. Specifically, the objective function 214 can include a reconstruction loss that measures the error between the target waveform 204 and the output waveform 206. In some cases, the point - wise reconstruction loss in the raw waveform is implemented using, for example, the mean squared error between the waveforms.
[0074] However, this type of reconstruction loss can, in some cases, have limitations. For example, the reason is that two different waveforms can sound perceptually equal, but point - wise similar waveforms can sound very different. To mitigate this problem, the objective function 214 can utilize a multi - scale spectral reconstruction loss that measures the error between the mel - spectrograms of the target waveform 204 and the output waveform 206. A spectrogram characterizes the frequency spectrum of an audio waveform over time using, for example, the short - time Fourier transform (STFT). A mel - spectrogram is a spectrogram converted to the mel scale. Since humans generally do not perceive audio frequencies on a linear scale, the mel scale can appropriately evaluate frequency components to facilitate fidelity. For example, the target waveform
[0075]
Number
[0076] and the output waveform
[0077]
Number
[0078] The reconstruction loss L between... rec can include terms that measure the absolute error and log error of the mel spectrogram,
[0079]
Number
[0080] However, ||...|| n denotes the L n norm. Other reconstruction losses are possible, but this form of L rec satisfies a strictly proper scoring rule that may be desirable for training purposes. Here,
[0081]
Number
[0082] denotes the t-th frame (e.g., time slice) of a 64-bin mel spectrogram calculated using a window length equal to s and a hop length equal to s / 4. The α s coefficient can be set to
[0083]
Number
[0084] as may be set.
[0085] As described above, the set of training examples 116 can be selected to enable various modalities of the compression / decompression systems 100 / 200, such as general audio compression, audio-adapted compression, etc. For example, for training for general audio compression, the training examples 116 can include speech, music, and general audio waveforms. In other implementations, the training examples 116 can include only music waveforms to facilitate optimal music compression and playback.
[0086] In some cases, the target waveform 204 is equal to the input waveform 202, thereby enabling the neural network to be trained towards a faithful and perceptually similar reconstruction. However, the target waveform 204 can also be modified with respect to the input waveform 202 to facilitate more advanced functionality such as joint compression and Emphasis The nature of Emphasis can be determined by designing training examples 116 with certain qualities. For example, the target waveform 204 can be a voiced version of the input waveform 202, so that the neural network improves the audio dialog during waveform reconstruction. Alternatively or additionally, the target waveform 204 can be a noise-removed version of the input waveform 202 that trains the network to suppress background noise. Generally, any desired audio Emphasis can be enabled using this technique. Emphasis In a further implementation, the encoder 102 and / or the decoder 104 can be conditioned on data typically included in the training examples 116 that define whether the target waveform 204 is equal to the input waveform 202 or a modified version of the waveform 202. For example, the training examples 116 can include a conditioning signal that represents two modes ( Emphasis is enabled or disabled) so that the neural network is trained to enable
[0087] only when a signal is present. To implement this, the encoder 102 and / or the decoder 104 can have a dedicated layer such as a Feature-wise Linear Modulation (FilM) layer to process the conditioning signal. After training, this technique allows the audio compression / decompression system 100 / 200 to Emphasis by supplying the conditioning signal through the network. Emphasis In some cases, the target waveform 204 is equal to the input waveform 202, thereby enabling the neural network to be trained towards a faithful and perceptually similar reconstruction. However, the target waveform 204 can also be modified with respect to the input waveform 202 to facilitate more advanced functionality such as joint compression and Emphasis The nature of Emphasis can be determined by designing training examples 116 with certain qualities. For example, the target waveform 204 can be a voiced version of the input waveform 202, so that the neural network improves the audio dialog during waveform reconstruction. Alternatively or additionally, the target waveform 204 can be a noise-removed version of the input waveform 202 that trains the network to suppress background noise. Generally, any desired audio EmphasisIt can be made possible to flexibly control in real time. Therefore, the compression system 100 implements this controllable Emphasis to enable the compression of acoustic scenes and natural sounds that would otherwise be removed Emphasis (e.g., noise removal).
[0088] Next, we return to the encoder neural network 102. The input waveform 202 is processed by the encoder 102 and encoded into a sequence of feature vectors 208. This process, which may involve multiple encoder neural network layers, can be collectively represented by an encoder function ε θ that maps the input waveform x to the feature vector y, such that y(x) = ε θ (x). The encoder function ε θ is parameterized by the encoder network parameters θ and can be updated using the objective function 214 to minimize the loss during encoding.
[0089] Next, the feature vector 208 is compressed by the RVQ 106 to generate the coded CFV 210 and the corresponding QFV 212. The quantization process of the RVQ 106, which may involve multiple vector quantizers 108, can be collectively represented by an RVQ function Q
[0090]
Number
[0091] that maps the feature vector y to QFV ψ and can be written as
[0092]
Number
[0093] Note that it is of the form of. The RVQ function Q ψis parameterized by codebook parameter ψ that can be updated using objective function 214 to minimize loss during quantization.
[0094] Training system 300 can minimize the quantization loss associated with RVQ106 by appropriately aligning the code vectors with the vector space of feature vector 208. That is, codebook parameter ψ can be updated by the training system 300 by backpropagating the gradient of the objective function 214. For example, codebook 110 can be repeatedly updated during training using the exponential moving average of feature vector 208. Training system 300 can also improve the use of codebook 110 by running the k - means algorithm on the first set of training examples 116 and using the learned centroids as initialization for subsequent training examples 116. Alternatively or additionally, if a code vector is not assigned to the feature vectors 208 for a number of training examples 116, training system 300 can replace it with a random feature vector 208 sampled during the current training example 116. For example, training system 300 can track the exponential moving average of the assignment to each code vector (along with a decay function of 0.99) and replace code vectors whose statistical value of this is less than 2.
[0095] For variable (e.g., scalable) bitrates, to fully train the neural network, training system 300 can select a specific number n of vector quantizers 108 to be used for each training example 116 such that the number of quantizers 108 varies during training examples 116. q For example, training system 300 can uniformly randomly sample n in [1; N q for each training example 116, and for the first i = 1... n in the sequence q qquantizers 108 can be used. Thus, the network can only use quantizers in the range n q =1...N q , and no architecture changes are required for the encoder 102 or the decoder 104. After training, the audio compression and decompression system 100 / 200 can select a particular number of quantizers n during compression and decompression that accommodates the desired bitrate. q can be selected.
[0096] Reference is now made to the decoder neural network 104. The QFV 212 is processed by the decoder 104 and decoded into an output audio waveform 206. Similar to the encoder 102, the decoding is performed using a decoder function D φ The decoder function D may involve multiple decoder network layers, which can be collectively represented by φ is the input QFV
[0097]
number
[0098] The output waveform
[0099]
number
[0100] Mapping to
[0101]
number
[0102] In some implementations, the input waveform x and the output waveform
[0103]
number
[0104] have the same sampling rate, although this need not be the case. The decoder function D φ is parameterized by decoder network parameters φ that can be updated using the objective function 214 to minimize loss during decoding.
[0105] Output waveform
[0106]
Number
[0107] Generally depends on the encoder network parameters θ, the codebook parameters ψ, and the decoder network parameters φ, so an objective function 214 that includes the reconstruction loss between the output waveform 206 and the target waveform 204 of each training example 116 can be used to update these network parameters. Specifically, the gradient of the objective function 214 can be calculated, for example, using backpropagation with gradient descent to iteratively update the network parameters. Generally, the network parameters are updated for the purpose of optimizing the objective function 214.
[0108] In some implementations, the training system 300 utilizes a discriminator neural network 216 to incorporate an adversarial loss 218 into the objective function 214 and potentially additional reconstruction losses. The adversarial loss 218 can promote the perceptual quality of the waveform reconstructed by the neural network. In this case, the discriminator 216 is trained together with the training system 300 to compete with the encoder 102, the decoder 104, and the RVQ 106. That is, the discriminator 216 is trained to distinguish the target waveform 204 from the output waveform 206, while the encoder 102, the decoder 104, and the RVQ 106 are trained to fool the discriminator 216.
[0109] The discriminator 216 calculates a discriminator score L for k={1, 2, ..., K}. k This can be implemented in the adversarial loss 218 by using a set of scores, where each score characterizes the estimated likelihood that the output waveform 206 is not produced as output from the decoder 104. For example, the discriminator 216 may discriminate between the output waveform
[0110]
number
[0111] , and processes the waveform using one or more neural network layers to obtain logits.
[0112]
number
[0113] In this case, g k,t is a discriminator function that maps the input waveform to output logits. k indexes a particular discriminator output and t indexes a particular logit of the discriminator output. In some implementations, the discriminator 216 utilizes a fully convolutional neural network such that the number of logits is proportional to the length of the input waveform.
[0114] The discriminator 216 uses the logit to calculate a respective discriminator score L for each discriminator output k. k For example, each score L k can be determined from the average over the logits as follows:
[0115]
number
[0116] Here, T k is the total number of logits for output k, and E xis the expected value over x. In some implementations, the adversarial loss L adv is the discriminator score L k averaged over.
[0117]
Number
[0118] The adversarial loss L adv may be included in the objective function 214 to promote the perceptual quality of the reconstructed waveform. Additionally, the discriminator 216 minimizes the discriminator loss function L dis to distinguish the target waveform
[0119]
Number
[0120] from the output waveform
[0121]
Number
[0122] and can be trained by the training system 300. In some implementations, L dis has the following form.
[0123]
Number
[0124] The target waveform
[0125]
Number
[0126] is generally equal to the input waveform
[0127] [Number]
[0128] or Emphasis Note that it depends on the input waveform x in that it can be the version that has been generated. L dis Regarding L, the encoder 102, decoder 104, and RVQ 106 train the discriminator 216 to efficiently classify the target waveform 204 from the output waveform 206. adv By minimizing L, they learn to fool the discriminator.
[0129] In some implementations, the discriminator 216 utilizes different versions of the waveform to determine the discriminator score L. k For example, in addition to the original waveform, the discriminator 216 can also use a downsampled version of the waveform (e.g., downsampled by 2, downsampled by 4, etc.) or a Fourier-transformed version of the waveform (e.g., STFT, Hartley transform, etc.), thereby adding diversity to the adversarial loss 218. As a specific implementation using four discriminator scores, L k=1 The score can correspond to the STFT waveform, while the discriminator score L k=2,3,4 can correspond to the original waveform, the waveform downsampled by 2, and the waveform downsampled by 4.
[0130] In a further implementation, the discriminator 216 introduces a reconstruction loss in the form of a "feature loss". Specifically, the feature loss L feat measures the error between the output of the internal layer of the discriminator for the target audio waveform 204 and the output of the internal layer of the discriminator for the output audio waveform 206. For example, the feature loss L feat for each layer l ∈ {1, 2,..., L} measures the difference between the target waveform
[0131] [Number]
[0132] and the output waveform
[0133] [Number]
[0134] the discriminator output for
[0135] [Number]
[0136] can be expressed as the absolute difference between, and is as follows.
[0137] [Number]
[0138] The feature loss can be a useful tool for promoting an increase in the fidelity between the output waveform 206 and the target waveform 204. Considering all the above loss terms, the objective function L can control the trade-off between the reconstruction loss, the adversarial loss, and the feature loss. L = λ rec L rec + λ adv L adv + λ feat L feat
[0139] The weight coefficients λ rec , λ adv , and λ feat are used to weight the appropriate loss terms, by which the objective function 214 can emphasize several properties such as faithful reconstruction, fidelity, perceptual quality, etc. In some implementations, the weight coefficients are set to λ rec = λ adv = 1 and λ feat = 100.
[0140] FIG. 4 is a flowchart of an exemplary process 400 for compressing an audio waveform. For convenience, process 400 is described as being performed by a system of one or more computers located in one or more locations. For example, an audio compression system appropriately programmed in accordance with this specification, such as audio compression system 100 of FIG. 1, can execute process 400.
[0141] The system receives an audio waveform (402). The audio waveform includes respective audio samples at each of a plurality of time steps. In some cases, the time steps may correspond to a particular sampling rate.
[0142] The system processes the audio waveform using an encoder neural network to generate a feature vector representing the audio waveform (404).
[0143] The system processes each feature vector using a plurality of vector quantizers to generate a respective coded representation of the feature vector (406), where each vector quantizer is associated with a respective codebook of code vectors. Each coded representation of the feature vector identifies a plurality of code vectors, including code vectors from the codebook of each vector quantizer, that define a respective quantized representation of the feature vector. In some implementations, each quantized representation of the feature vector is defined by the sum of a plurality of code vectors.
[0144] The system compresses the coded representation of the feature vector to generate a compressed representation of the audio waveform (408). In some implementations, the system compresses the coded representation of the feature vector using entropy coding.
[0145] FIG. 5 is a flowchart of an exemplary process 500 for restoring a compressed audio waveform. For convenience, process 500 is described as being performed by a system of one or more computers located at one or more locations. For example, an audio restoration system suitably programmed in accordance with this specification, such as the audio restoration system 200 of FIG. 2, can perform process 500.
[0146] The system receives a compressed representation of the input audio waveform (502).
[0147] The system restores the compressed representation of the audio waveform to obtain a coded representation of a feature vector representing the input audio waveform (504). In some implementations, the system uses entropy decoding to restore the compressed representation of the input audio waveform.
[0148] For each coded representation of the feature vector, the system identifies a plurality of code vectors, including code vectors from the codebook of each vector quantizer, that define a quantized representation of the respective feature vector (506). In some implementations, each quantized representation of the feature vector is defined by the sum of a plurality of code vectors.
[0149] The system uses a decoder neural network to process the quantized representation of the feature vector to generate an output audio waveform (510). The output audio waveform can include respective audio samples at each of a plurality of time steps. Optionally, the time steps can correspond to a particular sampling rate.
[0150] FIG. 6 is a flowchart of an exemplary process 600 for generating a quantized representation of a feature vector using a residual vector quantizer. For convenience, process 600 is described as being performed by a system of one or more computers located at one or more locations.
[0151] The system receives a feature vector (602) in the first vector quantizer in a sequence of vector quantizers.
[0152] The system identifies a code vector from the codebook of the first vector quantizer in the sequence for representing the feature vector, based on the feature vector (604). For example, a distance metric (e.g., error) can be calculated between the feature vector and each code vector in the codebook. The code vector with the smallest distance metric can be selected to represent the feature vector.
[0153] The system determines a current residual vector based on the error between the feature vector and the code vector representing the feature vector (606). For example, the residual vector can be the difference between the feature vector and the code vector representing the feature vector. The codeword corresponding to the code vector representing the feature vector can be stored in the coded representation of the feature vector.
[0154] The system receives the current residual vector generated by the previous vector quantizer in the sequence in the next vector quantizer in the sequence (608).
[0155] The system identifies a code vector from the codebook of the next vector quantizer in the sequence for representing the current residual vector, based on the current residual vector (610). For example, a distance metric (e.g., error) can be calculated between the current residual vector and each code vector in the codebook. The code vector with the smallest distance metric can be selected to represent the current residual vector. The codeword corresponding to the code vector representing the current residual vector can be stored in the coded representation of the feature vector.
[0156] The system updates the current residual vector (612) based on the error between the current residual vector and the code vector representing the current residual vector. For example, the current residual vector can be updated by subtracting the code vector representing the current residual vector from the current residual vector.
[0157] Steps 608 - 612 can be repeated for each remaining vector quantizer in the sequence. The final coded representation of the feature vector contains the codewords of each code vector selected from its respective codebook during process 600. The quantized representation of the feature vector corresponds to the sum of all the code vectors specified by the codewords of the coded representation of the feature vector. In some implementations, the codebooks of the vector quantizers in the sequence contain an equal number of code vectors such that each codebook is allocated the same space in memory.
[0158] FIG. 7 is a flowchart of an exemplary process 700 for training an encoder neural network, a decoder neural network, and a residual vector quantizer together. For convenience, process 700 is described as being executed by a system of one or more computers located at one or more locations. For example, a training system appropriately programmed in accordance with this specification, such as training system 300 of FIG. 3, can execute process 700.
[0159] The system obtains training examples (702) that include respective input audio waveforms and corresponding target audio waveforms. In some implementations, one or more of the target audio waveforms among the training examples are, for example, a noise - removed version of the input audio waveform, etc., of the input audio waveform EmphasisIt can be the version that has been made. One or more of the target audio waveforms in the training examples can also be the same as the input audio waveform. Alternatively or in addition, the input audio waveform can be a voice or music waveform.
[0160] The system uses an encoder neural network, a plurality of vector quantizers, and a decoder neural network to process the input audio waveform for each training example to generate respective output audio waveforms (704), where each vector quantizer is associated with a respective codebook. In some implementations, the encoder and / or decoder neural network is conditioned on data defining whether the corresponding target audio waveform is the same as the input audio waveform or the version that has been made of the input audio waveform. Emphasis is the version that has been made.
[0161] The system determines, for example, using backpropagation, the gradient of the objective function that depends on the respective output and target audio waveforms for each training example (706).
[0162] The system uses the gradient of the objective function to update one or more of a set of encoder network parameters, a set of decoder network parameters, or the codebooks of the plurality of vector quantizers (708). For example, the parameters can be updated using any suitable gradient descent optimization technique, such as update rules like RMSprop, Adam, etc.
[0163] Figures 8A and 8B show an example of a fully convolutional neural network architecture for the neural networks of the encoder 102 and the decoder 104. C represents the number of channels, and D is the dimension of the feature vector 208. The architecture in Figures 8A and 8B is based on the SoundStream model developed by N. Zeghidour, A. Luebs, A. Omran, J. Skoglund and M. Tagliasacchi, "SoundStream: An End-to-End Neural Audio Codec", in IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 495-507, 2022. This model is an adaptation of the SEANet encoder-decoder network without skip connections, designed by Y. Li, M. Tagliasacchi, O. Rybakov, V. Ungureanu and D. Roblek, "Real-Time Speech Frequency Bandwidth Extension", ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 691-695.
[0164] The encoder 102 includes a Conv1D layer 802 and four subsequent EncoderBlocks 804. Each of the blocks includes three ResidualUnits 812 that each contain dilated convolutions with dilation rates of 1, 3, and 9 respectively, and a downsampling layer in the form of a strided convolution that follows. The convolutional layers inside the EncoderBlock 804 and ResidualUnit 812 are shown in FIG. 8B. The number of channels doubles whenever there is downsampling. The last Conv1D layer 802 with a kernel of length 3 and a stride of 1 is used to set the dimensionality of the feature vector 208 to D. A FiLM conditioning layer 806 can also be implemented to process the conditioning signal for use in joint compression and Emphasis joint restoration. The FiLM layer 806 performs a feature-wise affine transformation on the feature vector 208 of the neural network, conditioned on the conditioning signal.
[0165] In this case, the decoder 104 effectively replicates the encoder 102. The DecoderBlock 810 includes a transposed Conv1D layer 814 for upsampling and three subsequent ResidualUnits 812. The convolutional layers inside the DecoderBlock 810 and ResidualUnit 812 are shown in FIG. 8B. The decoder 104 uses the same strides as the encoder 102, in reverse order, to reconstruct the waveform at the same resolution as the input waveform. The number of channels is halved whenever there is upsampling. The last Conv1D layer 802 with one filter, a kernel of size 7, and a stride of 1 projects the feature vector 208 back to the waveform 112. A FiLM conditioning layer 806 can also be implemented to process the conditioning signal for use in joint restoration and Emphasis joint restoration. In some implementations, both the encoder 102 and the decoder 104 are audio EmphasisAlthough it is executed, in other implementation forms, only one of the encoder 102 or the decoder 104 is responsible.
[0166] This specification uses the term "configured" in relation to system and computer program components. For a system of one or more computers to be configured to perform a particular operation or action means that the system has installed software, firmware, hardware, or a combination thereof that causes the operation or action to be performed on the system during operation. For one or more computer programs to be configured to perform a particular operation or action means that the one or more programs include instructions that cause the operation or action to be performed on the device when executed by a data processing device.
[0167] Embodiments of the subject matter and the functional operations described in this specification may be implemented in digital electronic circuits, in tangibly embodied computer software or firmware, in computer hardware, or in one or more combinations thereof, including the structures disclosed in this specification and their structural equivalents. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or sequential access memory device, or one or more combinations thereof. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, e.g., a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver device for execution by a data processing apparatus.
[0168] The term "data processing apparatus" refers to data processing hardware and encompasses, by way of example, all kinds of apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can also be, or further include, dedicated logic circuits, such as, for example, FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). The apparatus may, in some cases, include, in addition to the hardware, code for creating an execution environment for a computer program, such as, for example, processor firmware, protocol stacks, database management systems, operating systems, or code constituting one or more combinations thereof.
[0169] A computer program, which may also be called or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may correspond to a file in a file system, but need not. The program may be stored in a file that holds other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple cooperating files, such as files that hold one or more modules, subprograms, or portions of code. The computer program may be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.
[0170] In this specification, the term "engine" is widely used to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers at one or more locations. In some cases, one or more computers may be dedicated to a particular engine, and in other cases, multiple engines may be installed and running on the same one or more computers.
[0171] The processes and logical flows described in this specification may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows may also be performed by, for example, a special purpose logic circuit such as an FPGA or ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0172] A computer suitable for the execution of a computer program can be based on a general-purpose or special-purpose microprocessor, or both, or any other kind of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a central processing unit for executing or running instructions, and one or more memory devices for storing the instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, dedicated logic circuitry. Generally, a computer will also include, or be operably coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, for receiving data therefrom, or transferring data thereto, or both. However, a computer need not have such devices. Additionally, a computer can be embedded in another device, such as, to name just a few examples, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.
[0173] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0174] To provide interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may be used as well to provide interaction with the user. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user may be in any form, including acoustic input, voice input, or tactile input. Additionally, the computer can interact with the user by sending documents to the devices used by the user and receiving documents from those devices, for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device, or by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user in return.
[0175] A data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive portions of machine learning training or production, i.e., inference, workloads.
[0176] The machine learning model may be implemented and deployed using a machine learning framework, such as the TensorFlow framework.
[0177] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes back-end components, such as a data server, or includes middleware components, such as an application server, or includes front-end components, such as a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0178] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server may send data, such as an HTML page, to a user device to display data to a user who interacts with a device acting as a client and to receive user input from the user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.
[0179] This specification includes many specific implementation details, but these should not be construed as limiting the scope of any invention or the scope of what may be claimed. Rather, they should be construed as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, features may be described above as acting in some combinations and even being claimed as such initially, but one or more features from the claimed combination may in some cases be removable from that combination, and the claimed combination may be directed to a sub-combination or a variant of a sub-combination.
[0180] Similarly, operations are illustrated in the drawings and described in the claims in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all illustrated operations be performed to achieve a desired result. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.
[0181] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order or sequence shown to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Description of the Signs
[0182] 100 Audio compression system, compression system 102 Encoder neural network, encoder 104 Decoder neural network, decoder 106 Residual vector quantizer, residual (e.g., multi-stage) vector quantizer RVQ, RVQ, quantizer 108 Vector quantizer, quantizer 110 Codebook, trainable codebook 112 Audio waveform, waveform 114 Compressed representation of the audio waveform, compressed audio waveform 116 Training example 200 Audio restoration system, restoration system 202 Input waveform, respective input audio waveform, waveform 204 Corresponding target audio waveform, target waveform, target audio waveform 206 Output waveform, obtained output audio waveform, output audio waveform 208 Feature vector, compressed feature vector 210 Coded representation (CFV) of the feature vector, CFV 212 Quantized representation (QFV) of the feature vector, QFV, input QFV 214 Objective function 216 Discriminator neural network, discriminator 218 Adversarial loss 300 Training System 302 Entropy Coder 802 Conv1D Layer 804 Encoder Block 806 FiLM Conditioning Layer, FiLM Layer 810 Decoder Block 812 Residual Unit 814 Transposed Conv1D Layer
Claims
1. A method performed by one or more computers, comprising: receiving an audio waveform comprising respective audio samples for each of a plurality of time steps; processing the audio waveform using an encoder neural network to generate a plurality of feature vectors representing the audio waveform; generating a respective coded representation of each of the plurality of feature vectors using a plurality of vector quantizers each associated with a respective codebook of code vectors, wherein the respective coded representation of each feature vector identifies a plurality of code vectors including the respective code vector from the codebook of the respective vector quantizer that defines the quantized representation of the feature vector; generating a compressed representation of the audio waveform by compressing the respective coded representation of each of the plurality of feature vectors; and the encoder neural network and the codebooks of the plurality of vector quantizers are trained together with a decoder neural network, the decoder neural network being configured to receive a respective quantized representation of each of a plurality of feature vectors representing an input audio waveform, generated using the encoder neural network and the plurality of vector quantizers; process the quantized representation of the feature vectors representing the input audio waveform to generate an output audio waveform; wherein the training comprises: (i) obtaining a plurality of training examples each including a respective input audio waveform and (ii) a corresponding target audio waveform; processing each respective input audio waveform from each training example using the encoder neural network, a plurality of vector quantizers from a sequence of vector quantizers, and the decoder neural network to generate an output audio waveform that is an estimate of the corresponding target audio waveform; determining a gradient of an objective function that depends on the respective output audio waveform and the target audio waveform for each training example. Updating one or more of a set of encoder neural network parameters, a set of decoder neural network parameters, or the codebook of the plurality of vector quantizers using the gradient of the objective function comprising for each training example selecting a respective number of vector quantizers to be used when quantizing the feature vector representing the input audio waveform, and generating the corresponding output audio waveform using only the selected number of vector quantizers from the sequence of vector quantizers further comprising a method **Claim 2** wherein the plurality of vector quantizers are ordered in a sequence, and for each of the plurality of feature vectors, the step of generating the coded representation of the feature vector for the first vector quantizer in the sequence of vector quantizers receiving the feature vector, and identifying, based on the feature vector, each code vector from the codebook of the vector quantizer for representing the feature vector, and determining a current residual vector based on the error between (i) the feature vector and (ii) the code vector representing the feature vector comprising The method according to claim 1, wherein the coded representation of the feature vector identifies the code vector representing the feature vector **Claim 3** for each of the plurality of feature vectors, the step of generating the coded representation of the feature vector for each vector quantizer after the first vector quantizer in the sequence of vector quantizers receiving the current residual vector generated by the previous vector quantizer in the sequence of vector quantizers, and identifying, based on the current residual vector, each code vector from the codebook of the vector quantizer for representing the current residual vector, and if the vector quantizer is not the last vector quantizer in the sequence of vector quantizers updating the current residual vector based on the error between (i) the current residual vector and (ii) the code vector representing the current residual vector further comprising The method according to claim 2, wherein the coded representation of the feature vector identifies the code vector representing the current residual vector. **Claim 4** The step of generating the compressed representation of the audio waveform comprises entropy encoding the respective coded representations of each of the plurality of feature vectors The method according to claim 1. **Claim 5** The method according to claim 1, wherein the quantized representation of each feature vector is defined by the sum of the plurality of code vectors identified by the coded representation of the feature vector. **Claim 6** The method according to claim 1, wherein all of the codebooks of the plurality of vector quantizers contain an equal number of code vectors. **Claim 7** The method according to claim 1, wherein for one or more of the training examples, the target audio waveform is an enhanced version of the input audio waveform. **Claim 8** The method according to claim 7, wherein for one or more of the training examples, the target audio waveform is a noise-removed version of the input audio waveform. **Claim 9** The method according to claim 7, wherein for one or more of the training examples, the target audio waveform is the same as the input audio waveform. **Claim 10** The step of processing each input audio waveform to generate the corresponding output audio waveform comprises conditioning the encoder neural network, the decoder neural network, or both, on data defining whether the corresponding target audio waveform is (i) the input audio waveform or (ii) an enhanced version of the input audio waveform The method according to claim 9. **Claim 11** The method according to claim 1, wherein the selected number of vector quantizers to be used in quantizing the feature vectors representing the input audio waveform varies among the training examples. **Claim 12** For each training example, the step of selecting the respective number of vector quantizers to be used in quantizing the feature vectors representing the respective input audio waveforms The step of randomly sampling each of the numbers of the vector quantizers that will be used when quantizing the feature vectors representing the respective input audio waveforms The method according to claim 1, comprising: **Claim 13** The method according to claim 1, wherein the objective function comprises a reconstruction loss that measures an error between (i) each of the output audio waveforms and (ii) the corresponding target audio waveform for each training example. **Claim 14** The method according to claim 13, wherein for each training example, the reconstruction loss measures a multi-scale spectral error between (i) each of the output audio waveforms and (ii) the corresponding target audio waveform. **Claim 15** The training is, for each training example, Using a discriminator neural network to process data derived from each of the output audio waveforms to generate a set of one or more discriminator scores, each discriminator score characterizing an estimated likelihood that each of the output audio waveforms is the audio waveform generated using the encoder neural network, the plurality of vector quantizers, and the decoder neural network. Further comprising The method according to claim 1, wherein the objective function comprises an adversarial loss that depends on the discriminator scores generated by the discriminator neural network. **Claim 16** The method according to claim 15, wherein the data derived from the output audio waveforms comprises the output audio waveforms, a downsampled version of the output audio waveforms, or a Fourier-transformed version of the output audio waveforms. **Claim 17** The method according to claim 13, wherein for each training example, the reconstruction loss measures an error between (i) one or more intermediate outputs generated by the discriminator neural network by processing each of the output audio waveforms and (ii) one or more intermediate outputs generated by the discriminator neural network by processing the corresponding target audio waveform. **Claim 18** The method according to claim 1, wherein the codebook of the plurality of vector quantizers is repeatedly updated during the training using the exponentially weighted average of the feature vectors generated by the encoder neural network.
19. The method according to claim 1, comprising a sequence of encoder blocks, wherein each of the encoder neural networks is configured to process each set of input feature vectors according to a set of encoder block parameters to generate a set of output feature vectors having a lower temporal resolution than the set of input feature vectors.
20. The method according to claim 1, comprising a sequence of decoder blocks, wherein each of the decoder neural networks is configured to process each set of input feature vectors according to a set of decoder block parameters to generate a set of output feature vectors having a higher temporal resolution than the set of input feature vectors.
21. The method according to claim 1, wherein the audio waveform is a speech waveform or a music waveform.
22. The method according to claim 1, further comprising transmitting the compressed representation of the audio waveform over a network.
23. A method performed by one or more computers, comprising: receiving a compressed representation of an input audio waveform generated by the method according to claim 1; restoring the compressed representation of the input audio waveform to obtain a respective coded representation of each of a plurality of feature vectors representing the input audio waveform, wherein the coded representation of each feature vector comprises a respective code vector from each of the respective codebooks of the plurality of vector quantizers that define a quantized representation of the feature vector, identifying a plurality of code vectors; generating a respective quantized representation of each feature vector from the coded representation of the feature vector; using a decoder neural network to process the quantized representations of the feature vectors to generate an output audio waveform; and including.
24. A method performed by one or more computers, comprising: receiving a compressed representation of an audio waveform; Restoring the compressed representation of the audio waveform to obtain a coded representation of each of the plurality of feature vectors representing the audio waveform, where identifying a plurality of code vectors, each coded representation of a respective feature vector including a respective code vector from a respective codebook of a plurality of vector quantizers that define a quantized representation of the feature vector; generating a quantized representation of each respective feature vector from the coded representation of the feature vector; processing the quantized representation of the feature vector using a decoder neural network to generate an output audio waveform; including the decoder neural network and the codebooks of the plurality of vector quantizers are trained together with an encoder neural network, the encoder neural network is configured to receive an input audio waveform comprising respective audio samples for each of a plurality of time steps, process the input audio waveform to generate a plurality of feature vectors representing the input audio waveform; and the training includes (i) obtaining a plurality of training examples each including a respective input audio waveform and (ii) a corresponding target audio waveform; using the encoder neural network, a plurality of vector quantizers from a sequence of vector quantizers, and the decoder neural network to process the respective input audio waveform from each training example to generate an output audio waveform for the training example that is an estimate of the corresponding target audio waveform; determining a gradient of an objective function that depends on the respective output audio waveform and the target audio waveform for each training example; updating one or more of a set of encoder neural network parameters, a set of decoder neural network parameters, or the codebooks of the plurality of vector quantizers using the gradient of the objective function; including for each training example Selecting the respective number of vector quantizers to be used when quantizing the feature vectors representing the input audio waveform; Generating the corresponding output audio waveform using only the selected number of vector quantizers from the sequence of vector quantizers; A method further comprising. **Claim 25** The plurality of vector quantizers are ordered in a sequence; For each of the plurality of feature vectors, the sequence of vector quantizers; For the first vector quantizer in the sequence of vector quantizers; Receiving the feature vector; Identifying, based on the feature vector, each code vector from the codebook of the vector quantizer for representing the feature vector; Determining a current residual vector based on the error between (i) the feature vector and (ii) the code vector representing the feature vector; Generating a coded representation of the feature vector by performing an operation including; The method according to claim 24, wherein the coded representation of the feature vector identifies the code vector representing the feature vector. **Claim 26** For each of the plurality of feature vectors, the step of generating the coded representation of the feature vector; For each vector quantizer after the first vector quantizer in the sequence of vector quantizers; Receiving the current residual vector generated by the previous vector quantizer in the sequence of vector quantizers; Identifying, based on the current residual vector, each code vector from the codebook of the vector quantizer for representing the current residual vector; If the vector quantizer is not the last vector quantizer in the sequence of vector quantizers; Updating the current residual vector based on the error between (i) the current residual vector and (ii) the code vector representing the current residual vector; Further comprising; The method according to claim 25, wherein the coded representation of the feature vector identifies the code vector representing the current residual vector. **Claim 27** Restoring the compressed representation of the audio waveform to obtain the respective coded representations of each of the plurality of feature vectors, entropy decoding the respective entropy-coded representations of each of the coded representations of the feature vectors The method according to claim 26, comprising: **Claim 28** The method according to claim 24, wherein the respective quantized representations of each feature vector are defined by the sum of the plurality of code vectors identified by the coded representation of the feature vector. **Claim 29** The method according to claim 24, wherein all of the codebooks of the plurality of vector quantizers contain an equal number of code vectors. **Claim 30** The method according to claim 24, wherein for one or more of the training examples, the target audio waveform is an enhanced version of the input audio waveform. **Claim 31** The method according to claim 30, wherein for one or more of the training examples, the target audio waveform is a noise-removed version of the input audio waveform. **Claim 32** The method according to claim 30, wherein for one or more of the training examples, the target audio waveform is the same as the input audio waveform. **Claim 33** Processing each input audio waveform to generate the corresponding output audio waveform, conditioning the encoder neural network, the decoder neural network, or both on data that defines whether the corresponding target audio waveform is (i) the input audio waveform or (ii) an enhanced version of the input audio waveform The method according to claim 32, comprising: **Claim 34** The method according to claim 24, wherein the selected number of vector quantizers to be used when quantizing the feature vectors representing the input audio waveform varies among the training examples. **Claim 35** For each training example, selecting the respective number of vector quantizers to be used when quantizing the feature vectors representing each of the input audio waveforms The step of randomly sampling each of the numbers of the vector quantizers to be used when quantizing the feature vectors representing the respective input audio waveforms The method according to claim 24, comprising: **Claim 36** The method according to claim 24, wherein the objective function comprises a reconstruction loss that measures an error between (i) each of the output audio waveforms and (ii) the corresponding target audio waveform for each training example. **Claim 37** The method according to claim 36, wherein for each training example, the reconstruction loss measures a multi-scale spectral error between (i) each of the output audio waveforms and (ii) the corresponding target audio waveform. **Claim 38** For each training example, the training Using a discriminator neural network to process data derived from each of the output audio waveforms to generate a set of one or more discriminator scores, wherein each discriminator score characterizes an estimated likelihood that each of the output audio waveforms is an audio waveform generated using the encoder neural network, the plurality of vector quantizers, and the decoder neural network further comprising: The method according to claim 24, wherein the objective function comprises an adversarial loss that depends on the discriminator scores generated by the discriminator neural network. **Claim 39** A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 to 23. **Claim 40** A system comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 24 to 38. **Claim 41** One or more non-transitory computer-readable recording media storing instructions, which, when executed by one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 1 to 23.
42. One or more non-transitory computer-readable recording media storing instructions, which, when executed by one or more computers, cause the one or more computers to perform the operations of the method according to any one of claims 24 to 38.