APPARATUS AND METHOD FOR END-TO-END ADVERSARY BLIND BANDWIDTH EXTENSION USING ONE OR MORE CONVOLUTIONAL AND / OR RECURRENT NETWORKS
By employing adversarial blind bandwidth expansion techniques with convolutional and recurrent networks, the method effectively addresses bandwidth limitations in voice communication technologies, enhancing voice quality and intelligibility while being suitable for real-time implementation on mobile devices.
Patent Information
- Application Number
- JP2021113056
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-07
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2041-07-07
AI Technical Summary
Existing voice communication technologies face challenges in maintaining high voice quality and intelligibility due to bandwidth limitations, especially in mobile phones and public switched telephone networks, where traditional voice codecs take years to be adopted.
The development of an apparatus and method for end-to-end adversarial blind bandwidth expansion using convolutional networks and/or recurrent networks, which processes narrowband speech inputs to generate wideband speech outputs by extrapolating signal envelopes and excitation signals.
This approach significantly improves the perceptual quality and intelligibility of voice communications, reduces the word error rate of automatic speech recognition systems, and can be implemented in real-time on embedded systems like mobile phones.
Smart Images

Figure 0007679244000022 
Figure 0007679244000023 
Figure 0007679244000024
Abstract
Description
Technical Field
[0001] Specification The present invention relates to an apparatus and method for end-to-end adversarial blind bandwidth expansion using one or more convolutional networks and / or recurrent networks.
Background Art
[0002] Voice communication is a technology that most people use every day, creating a vast amount of data that needs to be transmitted over networks such as VoIP (Voice over Internet Protocol), mobile phones, and the public switched telephone network. According to an OFCOM survey in 2017, an average of 156.75 minutes of outgoing calls per contract per month are made to mobile phones (see https: / / www.ofcom.org.uk / research-and-data / multi-sector-research / cmr / cmr-2018 / interactive).
[0003] Even if the amount of data to be transferred is small, it is desirable that the voice quality be high. To achieve this goal, voice compression technology has evolved over the past few decades from the compression of band-limited voice by simple pulse code modulation [1] to voice generation that can code full-band voice and coding schemes [2], [3] according to human perception models. Despite the existence of such standardized voice codecs, their adoption in mobile phones or public switched telephone networks takes years, if not decades. For these reasons, AMR-NB [4] is most frequently used as a codec for voice communication in mobile phones that only encodes frequencies from 200 Hz to 3400 Hz (usually called narrowband, NB). However, transmitting band-limited voice degrades not only the acoustic quality but also the intelligibility [5], [6], [7]. Blind bandwidth expansion (BBWE), also called artificial bandwidth expansion or audio super-resolution, artificially regenerates missing frequency components without transmitting additional information from the encoder. Since BBWE can be added to the decoder's toolchain without changing the transmission network, it functions as an intermediate solution to improve the perceived audio quality and intelligibility until a better codec is introduced into the network [5], [6], [8]. A robust BBWE may be a viable solution for modern voice transmission to save transmission bandwidth or improve quality. Furthermore, in other types of applications such as audio restoration where band-limited voice is stored or archived, BBWE is the only possible solution for expanding the audio bandwidth.
[0004] BBWE has a long tradition in the field of audio signal processing [9],
[10] , but the solutions based on deep neural networks (DNNs) have been considered mostly by researchers with backgrounds in artificial intelligence (AI) and image processing, rather than in audio signal processing. Such DNN-based systems are generally called speech super resolution (SSR). In image processing, the task of estimating a high-resolution image from one or more low-resolution observations is called super resolution and has received significant attention within the computer vision community. Recently, deep convolutional neural networks have produced better results than traditional methods
[11] , and super resolution generative adversarial networks are at the forefront
[12] .
[0005] Excellent BBWE can not only improve the perceptual quality of speech, but also improve the word error rate of automatic speech recognition systems
[13] .
[0006] Generative Adversarial Networks (GANs) can more appropriately restore finer structures for more realistic reproduction. However, some of these systems cannot be directly applied to the scenarios of voice communication. In the design of BBWE, in addition to the different natures of the underlying signals (such as different dimensions), the following points need to be considered. First, it is necessary that the delay of the algorithm (the time delay of the decoded speech from the original speech) does not become too large. Furthermore, the computational complexity and memory consumption must be able to meet the requirements of real-time processing in embedded systems such as mobile phones.
[0007] Recurrent neural networks are suitable for the analysis and prediction of time series such as speech. In fact, speech is considered to be stationary or quasi-periodic in a broad sense with a duration of about 20 to 25 ms, and its temporal correlation can be utilized by an RNN with a relatively small model. On the other hand, CNNs demonstrate performance in tasks such as pattern recognition and upscaling such as super-resolution of images. They also have the advantage of being able to highly parallelize processing. Therefore, in speech processing, especially in BBWE, it is necessary to consider both architectures.
[0008] As described above, in the state-of-the-art technology, the principle of BBWE was first presented by Karl-Otto Schmidt in 1933 [9], and an analog non-linear device was used to expand the bandwidth of the transmitted speech. The idea of performing (non-blind) bandwidth expansion on the excitation signal of a speech codec dates back at least to 1959
[10] . Subsequently, several so-called parametric BWEs were published that utilized the separation of speech signals into excitation and spectral envelopes based on the source-filter model of human speech generation. These systems apply statistical models to extrapolate the spectral envelope while generating the excitation signal by spectral folding
[14] , spectral conversion [8], or non-linearity
[15] . Statistical models for envelope extrapolation include simple codebook mapping
[16] , hidden Markov models
[14] , (shallow) neural networks
[17] , or more recently DNNs
[18] , etc.
[0009] Before using DNNs, the inputs to statistical models were often manually engineered features
[14] ,
[17] ,
[19] ,
[20] . With the introduction of DNNs, this approach can be simplified to directly use the log short-time Fourier transform (STFT) energy
[18] ,
[21] ,
[22] or the time-domain speech signal
[23] ,
[24] ,
[25] . The same holds for the output of statistical models. Instead of modeling sub-band energy [8] or other envelope representations
[21] , DNNs have enough power to model the magnitude of the spectrum for each bin
[15] , without going as far as the entire time-domain speech signal or a combination of the time and frequency domains
[26] . However, when the magnitude of the spectrum is modeled, it is necessary to reconstruct the phase due to spectral folding or transformation
[18] ,
[21] ,
[15] ,
[27] .
[0010] Regarding learning objectives, to design an efficient DNN-based solution, it is necessary to select an appropriate architecture, mainly by carefully choosing the learning loss function and the network type. Representative loss functions include mean squared error
[21] , categorical cross-entropy (CE) loss
[28] , adversarial loss
[29] ,
[30] ,
[25] , or a mixture of losses
[31] . The loss function can also determine the data representation.
[0011] Regarding mean squared error and cross-entropy, an acoustically motivated loss can be achieved by combining the mean squared error (MSE) loss with the log sub-band or bin energy [8]. The loss function derived from cross-entropy (XE) predicts the sample bits (or sample size) as classes, so the signal to be modeled needs to be quantized at a resolution not too high for processing by the DNN. 2 of the speech signal quantized at 16 bits 16Predicting classes is currently very costly to process with DNNs. Fortunately, quantizing the content of voice signals above 3.4 kHz into 8 bits does not significantly degrade the quality
[32] . Since the distribution of data learned using cross-entropy loss is preferably a Gaussian distribution rather than, for example, the Laplacian distribution of voice signals
[34] , it is usually shaped by a non-linear function. Surprisingly, the μ-law function, which has been used in
[35] ,
[32] ,
[23] ,
[24] to make voice data x more Gaussian, is exactly the same as the world's first standardized digital voice codec [1]. TIFF0007679244000001.tif18170
[0012] Regarding adversarial loss, even with today's powerful networks, the distribution of time-domain voice is very complex and difficult to model. A generative model learned with MSE or CE loss to match this complex distribution only generates a smoothed approximation of it. When applied to BBWE, this means that the resulting voice signal lacks sharpness and energy
[30] .
[0013] A Generative Adversarial Network
[36] can be regarded as a kind of extended loss function. Here, two networks, a generator and a discriminator, are competing. Figure 2 shows a Generative Adversarial Network. The generator attempts to generate realistic data, and the discriminator distinguishes between the generated data and the data from the training database. After successful learning, the discriminator is no longer necessary, and the goal is only to improve the loss of the generator. The reason adversarial learning is suitable for learning generative models such as BBWE is that it can model some modes of the distribution without smoothing or averaging over all modes.
[0014] Regarding the class of the network, it should be noted that another important aspect in the design of the DNN is the selection of the class of the network to be used. Generally, fully connected layers
[18] ,
[21] , convolutional neural networks (CNNs)
[11] ,
[37] , or recurrent neural networks (RNNs), and their subtypes such as long short-term memory (LSTM) units
[38] ,
[39] , [8], or gated recurrent units (GRUs)
[40] ,
[39] are well-known. Fully connected layers are only used in systems operating on frames
[18] ,
[21] , while RNNs and CNNs enable the processing of time-domain data in a streaming manner
[23] ,
[24] .
[0015] TIFF0007679244000002.tif146170
[0016] WaveNet (registered trademark) is also adopted in BBWE. In
[42] , it is trained with clean speech conditioned on the bitstream parameters of the encoded NB speech. Here, the network functions as a decoder and implicitly performs bandwidth expansion. In response to this, in
[24] , features calculated with the NB signal are given to WaveNet (registered trademark). When the learning is successful, only the features are given to the network, and the NB speech signal is ignored.
[0017] WaveNet (registered trademark)-based models claim very high perceptual quality, but are difficult to train and the computational complexity during evaluation is very high. This has led to several optimizations and alternative models (e.g.,
[43] ). One particular alternative is LPCNet, which is designed for either original speech synthesis
[32] or speech coding
[44] . In LPCNet, the convolutional layer of WaveNet (registered trademark) is replaced with a recurrent layer. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0018] The object of the present invention is to provide an improved concept for blind bandwidth expansion.
Means for Solving the Problem
[0019] The object of the present invention is solved by the apparatus according to claim 1, by the method according to claim 19, by the method according to claim 20, by the method according to claim 23, and by the computer program according to claim 25.
[0020] There is provided an apparatus for processing a narrowband speech input signal by performing bandwidth expansion of the narrowband speech input signal to obtain a wideband speech output signal according to an embodiment. The apparatus includes a signal envelope extrapolator including a first neural network, the first neural network being configured to receive a plurality of samples of the signal envelope of the narrowband speech input signal as input values of the first neural network, and being configured to determine a plurality of samples of the extrapolated signal envelope as output values of the first neural network. Further, the apparatus includes an excitation signal extrapolator 130 configured to receive a plurality of samples of the excitation signal of the narrowband speech input signal and configured to determine a plurality of extrapolated excitation signal samples. Further, the apparatus includes a combiner 140 configured to generate a wideband speech output signal such that the wideband speech output signal expands the bandwidth with respect to the narrowband speech input signal, depending on the plurality of samples of the extrapolated signal envelope and depending on the plurality of samples of the extrapolated excitation signal.
[0021] Furthermore, there is provided a method for processing a narrowband speech input signal by performing bandwidth expansion of the narrowband speech input signal to obtain a wideband speech output signal according to an embodiment. The method includes the following: - Receiving, as input values of a first neural network, a plurality of samples of the signal envelope of the narrowband speech input signal, and determining, as output values of the first neural network, a plurality of samples of the extrapolated signal envelope; - Receiving a plurality of samples of an excitation signal of a narrowband audio input signal and determining a plurality of extrapolated excitation signal samples, and - Generating the wideband audio output signal to expand the bandwidth with respect to the narrowband audio input signal, wherein the wideband audio input signal depends on a plurality of samples of an extrapolated signal envelope and the plurality of extrapolated excitation signal samples.
[0022] Furthermore, a method for training a neural network according to an embodiment is provided.
[0023] - The neural network receives a first plurality of line spectral frequencies of a narrowband audio input signal as input values of the neural network.
[0024] - The neural network determines a second plurality of line spectral frequencies of a wideband audio output signal as output values of the first neural network, and each of the one or more second plurality of line spectral frequencies is associated with a frequency greater than any frequency associated with any of the first plurality of line spectral frequencies.
[0025] - The second plurality of line spectral frequencies of the wideband audio output signal are converted from the line spectral frequency domain to the linear prediction coding domain to obtain a second plurality of linear prediction coding coefficients of the wideband audio output signal.
[0026] - The finite impulse response filter is used to convert the second plurality of linear prediction coding coefficients of the wideband audio output signal from the linear prediction coding domain to the finite impulse response filter domain to obtain the linear prediction coding coefficients converted by a plurality of finite impulse filters.
[0027] - The method includes training the first neural network depending on the linear prediction coding coefficients converted by a plurality of finite impulse filters.
[0028] In an embodiment, when the first neural network is trained, linear prediction coding coefficients converted by a plurality of finite impulse filters, or values derived from the linear prediction coding coefficients converted by the plurality of finite impulse filters, can be fed back to the neural network, for example.
[0029] According to an embodiment, when the first neural network is trained, depending on the linear prediction coding coefficients converted by a plurality of finite impulse filters and a plurality of extrapolated excitation signal samples, for example, samples of a plurality of wideband voice output signals are generated, and the plurality of wideband voice output signals or values derived from the samples of the plurality of wideband voice output signals are fed back to the neural network, for example.
[0030] Furthermore, a method for training the first and / or second neural network according to an embodiment is provided.
[0031] - The first neural network receives, as input values of the first neural network, a plurality of samples of the signal envelope of a narrowband voice input signal, and determines, as output values of the first neural network, a plurality of samples of the extrapolated signal envelope, and / or the second neural network receives, as input values of the second neural network, a plurality of samples of the excitation signal of the narrowband voice input signal, and determines, as output values of the second neural network, a plurality of extrapolated excitation signal samples.
[0032] - The first and / or second neural network is trained using a discriminator neural network, and when the first and / or second neural network is trained, the first and / or second neural network and the discriminator neural network operate as a generative adversarial network.
[0033] - During the learning of the first and / or second neural network, the discriminator neural network receives, as input values of the discriminator neural network, the output values of the first and / or second neural network, or receives, as input values of the discriminator network, derived values derived from the output values of the first and / or second neural network.
[0034] - When receiving the input values of the discriminator neural network, the discriminator neural network determines, as the output of the discriminator neural network, a quality indication of the input values of the discriminator neural network, and the first and / or second neural network learns depending on the quality indication.
[0035] According to an embodiment, the discriminator neural network is, for example, a first discriminator neural network. The first neural network learns using, for example, the first discriminator neural network, and the first neural network learns depending on a quality indication that is a first quality indication. The second neural network learns using, for example, a second discriminator neural network, and during the learning of the second neural network, the second neural network and the second discriminator neural network operate as a second adversarial generation network. During the learning of the second neural network, the second discriminator neural network receives, for example, the output value of the second neural network as an input value of the second discriminator neural network, or may receive, for example, a derived value derived from the output value of the second neural network as an input value of the second discriminator network. When receiving the input value of the second discriminator neural network, the second discriminator neural network determines, as an output of the second discriminator neural network, a second quality indication of the input value of the second discriminator neural network, where the second neural network is configured to learn depending on the second quality indication.
[0036] Furthermore, a computer program is provided, each of the computer programs being configured to implement one of the above-described methods when executed on a computer or a signal processing device.
[0037] As already explained, blind bandwidth expansion improves the perceptual quality and intelligibility of telephone-quality voice by artificially reproducing missing frequency content that is transmitted without being encoded by a voice codec. Embodiments provide a new approach based on deep neural networks to solve this problem. These embodiments are based on convolutional architectures or recurrent architectures. All operate in the time domain. Motivated by the source-filter model of human voice generation, two of the provided systems decompose the voice signal into a spectral envelope and an excitation signal; each of them has a bandwidth individually extended by a dedicated DNN. All systems are trained such that adversarial loss and perceptual loss are mixed. To avoid mode collapse and more stable adversarial learning, spectral normalization can be employed, for example, in the discriminator.
[0038] In an embodiment, two BBWEs based on deep neural networks using adversarial learning for a voice coding scenario are provided.
[0039] According to an embodiment, two novel deep network structures for blind bandwidth expansion are provided, one based on a convolutional kernel and the other based on a recurrent kernel.
[0040] Both networks can be trained, for example, with a mixture of adversarial loss and spectral loss.
[0041] These two systems are adversarially trained BBWEs, meaning "end-to-end", that is, the input is voice in the time domain and the output is also voice in the time domain.
[0042] In an embodiment, to improve the performance of the GAN, for example, hinge loss and spectral normalization can be applied.
[0043] In an embodiment, a new approach to BBWE based on a generative model used for bandwidth expansion of an audio signal is provided.
[0044] In the two presented systems, paradigms established in the world of audio coding are adopted, i.e., the decomposition of an audio signal known as the envelope and source-filter model can be applied, for example, to a GAN model. As a result, the computational complexity can be reduced, for example, to about one-third. This approach has been tested and evaluated within the application of BBWE, but is not limited thereto. The system according to the embodiment significantly improves the speech recognition error rate of NB speech.
[0045] In some embodiments, a generative model for generating enhanced audio from encoded audio, bandwidth-limited audio, or damaged audio is provided.
[0046] According to an embodiment, the target audio for learning can be decomposed, for example, into an envelope and an excitation. The envelope can be, for example, LPC coefficients. The excitation can be, for example, an LPC residual.
[0047] In some embodiments, the envelope and the excitation can be learned separately, for example. Each of the envelope and the excitation can be learned, for example, with a mixture of an adversarial loss (known from an adversarial generative network (GAN)) and an L1 loss. A characteristic loss is also added to the learning of the excitation signal.
[0048] According to an embodiment, the envelope can be learned, for example, with the representation of the encoded and / or bandwidth-limited and / or damaged envelope as input and the original envelope as target. A possible representation of the envelope can be, for example, LPC coefficients.
[0049] In an embodiment, the input for learning the excitation signal can be, for example, encoded and / or band-limited and / or corrupted time-domain audio and / or compressed feature representations. The target can be, for example, the original clean audio.
[0050] According to an embodiment, to learn the excitation signal, the loss can be propagated, for example, through the envelope. This can be done, for example, by considering the envelope as a DNN layer that propagates the loss. If the envelope is represented by an LPC filter, this filter can be, for example, a pure IIR filter. In this case, the loss may, for example, propagate slowly or not at all (also known as the vanishing gradient problem). In an embodiment, the IIR filter can be approximated by an FIR filter, for example, by truncating the impulse response. As a result, the envelope can be implemented, for example, as a convolutional layer (CNN layer) within the network.
[0051] Some embodiments are based on the decomposition of an audio signal into an excitation signal and an envelope, similar to voice coders [2], [4]. This is achieved using linear predictive coding (LPC). Since the recurrent layer simply models the excitation signal, prediction is easy. In some embodiments, LPCNet is also adopted in BBWE
[33] .
[0052] Hereinafter, embodiments of the present invention will be described in more detail with reference to the drawings.
Brief Description of the Drawings
[0053]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Mode for Carrying Out the Invention
[0054] FIG. 1 shows an apparatus for processing a narrowband speech input signal by performing bandwidth expansion of the narrowband speech input signal to obtain a wideband speech output signal according to an embodiment.
[0055] The apparatus includes a signal envelope extrapolator 120 including a first neural network 125. The first neural network 125 is configured to receive a plurality of samples of a signal envelope of a narrowband speech input signal as input values of the first neural network 125, and is configured to determine a plurality of samples of an extrapolated signal envelope as output values of the first neural network 125.
[0056] Furthermore, the apparatus includes an excitation signal extrapolator 130 configured to receive a plurality of samples of an excitation signal of the narrowband speech input signal and configured to determine a plurality of extrapolated excitation signal samples.
[0057] Furthermore, the apparatus includes a combiner 140 configured to generate a wideband speech output signal such that the wideband speech output signal expands the bandwidth with respect to the narrowband speech input signal, depending on the plurality of samples of the extrapolated signal envelope and depending on the plurality of extrapolated excitation signal samples.
[0058] According to an embodiment, the input values of the first neural network 125 are a first plurality of line spectral frequencies of the narrowband speech input signal, and the first neural network 125 may be configured to determine a second plurality of line spectral frequencies of the wideband speech output signal as output values of the first neural network 125, for example, and each of the one or more second plurality of line spectral frequencies is associated with a frequency greater than any frequency associated with any of the first plurality of line spectral frequencies.
[0059] In an embodiment, when the first neural network 125 learns, the signal envelope extrapolator 120 may be configured to convert a plurality of wideband linear prediction coding coefficients derived from an original audio signal into finite impulse response filter coefficients, for example, by calculating an impulse response and truncating the impulse response.
[0060] For example, wideband LPC filter coefficients are converted into finite impulse response filter coefficients by calculating and truncating an impulse response, even if they are, for example, IIR filter coefficients. Since this is done during learning, the wideband LPC filter coefficients used for conversion into finite impulse response filter coefficients can be derived from, for example, the original wideband audio.
[0061] According to an embodiment, when the first neural network 125 is made to learn, the signal envelope extrapolator 120 is configured, for example, to feedback an error or a gradient of the error between the wideband audio output signal and the original wideband audio signal.
[0062] As outlined above, in the embodiment, the gradient of the error is backpropagated. Here, the error is the difference between the generated wideband audio and the true wideband audio.
[0063] Generally, excitation is generated from narrowband audio, and an envelope is generated therefrom. Finally, wideband audio is derived.
[0064] During application, the output of the excitation signal extrapolator 130 may be supplied to, for example, the signal envelope extrapolator 120.
[0065] In the embodiment, during learning involving backpropagation, the gradient of the error is first passed in the reverse direction to the signal envelope extrapolator 120 and then to the excitation signal extrapolator 130.
[0066] When the signal envelope is an IIR structure or filter, the gradient cannot be passed through. For this reason, the signal envelope is converted into a finite impulse response filter.
[0067] According to an embodiment, the first neural network 125 is trained using, for example, a first discriminator neural network. When the first neural network 125 is trained, for example, the first neural network 125 and the first discriminator neural network are configured to operate as a generative adversarial network. During the training of the first neural network 125, the first discriminator neural network is configured to receive, for example, the output value of the first neural network 125 as an input value of the first discriminator neural network, or to receive a derived value derived from the output value of the first neural network 125 as an input value of the first discriminator network. When receiving the input value of the first discriminator neural network, the first discriminator neural network is configured to determine, for example, a first quality indication of the input value of the first discriminator neural network as an output of the first discriminator neural network. Then, the first neural network 125 is configured to learn, for example, depending on the first quality indication.
[0068] In an embodiment, when receiving the input value of the first discriminator neural network, the first discriminator neural network is configured to determine a quality indication such that the quality indication indicates a probability that the input value of the first discriminator neural network is related to a recorded audio signal rather than an artificially generated audio signal, or the quality indication indicates a value for estimating whether the output value of the first discriminator neural network is related to a recorded signal or an artificially generated signal.
[0069] According to an embodiment, the first neural network 125 or the second neural network 135 learns using a loss function that depends on, for example, a quality indication determined by the first discriminator neural network.
[0070] In an embodiment, the loss function depends on, for example, a hinge loss, or a Wasserstein distance, or an entropy-based loss.
[0071] TIFF0007679244000003.tif36170
[0072] In an embodiment, the loss function depends on an (additional) Lp loss.
[0073] TIFF0007679244000004.tif30170
[0074] In an embodiment, the first discriminator neural network can be learned using, for example, recorded speech.
[0075] According to an embodiment, the excitation signal extrapolator 130 includes a second neural network 135, and the second neural network 135 is configured to receive, for example, a plurality of samples of the excitation signal of the narrowband speech input signal as input values of the second neural network 135, and / or is the narrowband speech input signal, and / or is a shaped version of the narrowband speech input signal. The second neural network 135 is configured to determine, for example, a plurality of extrapolated excitation signal samples as output values of the second neural network 135.
[0076] In an embodiment, the input value of the second neural network 135 is, for example, a first plurality of time-domain signal samples of an excitation signal of a narrowband voice input signal, and / or, for example, a narrowband voice input signal, and / or, for example, a shaped version of a narrowband voice input signal. Here, the second neural network 135 is configured, for example, such that the output value of the second neural network 135 is determined so that a plurality of samples of the extrapolated excitation signal are samples of a second plurality of time-domain signals of an extended time-domain excitation signal with an extended bandwidth with respect to the excitation signal of the narrowband voice input signal.
[0077] According to an embodiment, the second neural network 135 is trained, for example, using a second discriminator neural network, and during the training of the second neural network 135, the second neural network 135 and the second discriminator neural network are configured to operate as a second adversarial generation network. During the training of the second neural network 135, the second discriminator neural network is configured, for example, to receive the output value of the second neural network 135 as an input value of the second discriminator neural network, or, for example, to receive a derived value derived from the output value of the second neural network 135 as an input value of the second discriminator network. And / or, the second discriminator neural network is configured, for example, to receive the output value of the combiner 140 as an input value of the second discriminator neural network.
[0078] Upon receiving the input value of the second discriminator neural network, the second discriminator neural network is configured, for example, to determine a second quality indication of the input value of the second discriminator neural network as the output of the second discriminator neural network; wherein the second neural network 135 is configured, for example, to learn depending on the second quality indication.
[0079] In an embodiment, the apparatus includes, for example, a signal analyzer 110 configured to generate from the narrowband speech input signal a plurality of samples of the signal envelope of the narrowband speech input signal and a plurality of samples of the excitation signal of the narrowband speech input signal.
[0080] According to an embodiment, the first neural network 125 includes, for example, one or more convolutional neural networks.
[0081] In an embodiment, the first neural network 125 includes, for example, one or more deep neural networks.
[0082] Next, specific embodiments will be described.
[0083] Hereinafter, three BBWEs based on DNN according to embodiments will be described: two are based on a convolutional architecture and the other is based on a mixture of a convolutional architecture and a recurrent architecture. All can be adversarially learned, for example, using the same discriminator, the same perceptual loss, and the same optimization algorithm. The architecture of the first BBWE refers to WaveNet®, and the other architectures refer to LPCNet. First, since all the generation networks are presented and all the systems share the same discriminator, it will be described below.
[0084] First, the convolutional BBWE described in the embodiment will be explained.
[0085] The first architectural solution for this task is a stack of convolutional neural networks (CNNs), which are currently a standard component of GANs. Using CNNs enables high-speed processing, especially on GPUs.
[0086] The convolutional generative model adopted a structure such as WaveNet®. Specifically, it is a stack of 20 layers, and in all layers, each layer uses causal convolutions with a kernel size of 33 and softmax-gated activation
[45] . Biases are omitted. One of these layers is shown in Figure 3.
[0087] Figure 3 shows a single layer of the CNN-GAN with softmax-gated activations. The CNN layer has a 1D kernel with 32 input channels and 64 output channels. Half of the output channels are fed to the tanh activation, and the other half are fed to the softmax activation. Residual connections avoid vanishing gradients and maintain stable and effective learning.
[0088] As outlined, each CNN layer has 32 input channels and 64 output channels. Half of the output channels are fed to the tanh activation, and the other half are fed to the softmax activation. To form the 32-channel output of each layer, both activations are multiplied in the channel dimension. This type of activation is more robust to reconstruction artifacts than both ReLU and sigmoid-gated activations.
[0089] The additional input layer maps the 1D input signal to a 32D signal, and the additional output layer maps the 32D signal to a 1D output signal.
[0090] The weights of the convolutional kernel are normalized using weight normalization
[46] to enable stable learning operations. Also, to speed up the learning process, batch normalization is applied to the output function from the CNN layer.
[0091] Therefore, a complete convolutional layer is composed of a causal convolution followed by batch normalization, and finally a softmax gate activation to obtain the final output. To avoid the vanishing gradient and maintain stable and effective learning, there is a connection or shortcut from the input to the output
[47] .
[0092] In this convolutional BBWE, the model is run on the raw audio waveform in the time domain. The input signal is first resampled from NB to WB using simple Sinc interpolation and then fed into the generator model. The generator ensures that the original bandwidth of this upsampled signal is extended to obtain a complete WB structure with clearly higher perceptual quality.
[0093] This system is called CNN-GAN.
[0094] Here, the LPC-GAN according to the embodiment will be described.
[0095] In particular, two systems according to the embodiment are provided, which are different from the convolutional ones in two aspects: First, the architecture of the DNN is different. Second, the audio signal is decomposed into an excitation signal and an envelope. This is inspired by the BBWE based on LPCNet
[33] , and in some embodiments, this is applied. The motivation for decomposing the signal into excitation and envelope is the same as that of the BBWE based on LPCnet
[33] , which is to reduce the computational complexity of the entire system.
[0096] Figure 4 shows the block diagram of the system, and Figure 6 details one of the DNN bandwidths that expand the excitation signal. In particular, Figure 4 shows the proposed system based on decomposing the audio signal into an excitation signal and an LPC envelope. All the solid-line paths operate on samples, and all the dashed-line paths operate on 15-ms frames.
[0097] In Figure 4, the input NB audio signal is separated into an LPC representing the spectral envelope and an excitation signal (also known as the residue). The excitation signal and the input signal are fed to the first DNN for extrapolation to the WB excitation signal. This path operates on samples, shown here as a solid line. The LPC is extrapolated to the WB envelope using the second DNN in the upper path. This path operates on 15-ms frames and is shown here as a dashed line. The LPC coefficients are IIR filter coefficients, and since the filter can become unstable by operations such as extrapolation, they are extrapolated in the LSF domain
[48] . The LSF is a bijective transformation of the LPC with several advantages: First, it is less sensitive to noise perturbations, and a stable LPC filter is always guaranteed by the ordered set of LSFs with the minimum distance between coefficients. Second, since the spectral envelope at a specific frequency depends only on one of the LSFs, an incorrect extrapolation of a single LSF coefficient affects mainly the spectral envelope in a limited frequency range. These properties are suitable for extrapolating to the set representing the WB envelope. The extrapolated LSF coefficients are converted to the LPC domain to form the extrapolated excitation signal and the output signal. This is realized in various ways for learning and evaluation.
[0098] The extrapolated excitation signal formed by the LPC envelope forms the output WB signal. When training a DNN that extrapolates the excitation signal, it is necessary to propagate the gradient through the LPC filter, which can be achieved by performing LPC filtering in an additional DNN layer. Since the LPC filter is a pure IIR filter, this DNN layer must be a layer with recurrent units. Unfortunately, when backpropagating the gradient through the recurrent layer, the gradient vanishes (also known as the vanishing gradient problem
[38] ), resulting in insufficient learning. As a solution to this problem, the IIR filter coefficients are converted to FIR filter coefficients by calculating the truncated impulse response from the IIR filter. From signal processing, it is known that any IIR filter can be approximated by a FIR filter by truncating the infinite impulse response
[34] . And LPC shaping can be realized in the convolutional layer. Figure 5 shows the effect when truncated to 64 samples.
[0099] Figure 5 shows the transfer functions of the resulting FIR filter from the 12th-order IIR LPC filter and the truncated impulse response.
[0100] The IIR LPC envelope is smooth, but the truncated FIR envelope has many ripples and does not follow the IIR envelope well at high frequencies. Therefore, before calculating the truncated impulse response, the LPC coefficients are multiplied by an exponential function: TIFF0007679244000005.tif14170
[0101] TIFF0007679244000006.tif23170
[0102] In Figure 5, the impulse response is truncated to 64 samples. The green filter is the result of processing the IIR LPC coefficients with Equation (4), and the red filter is the one without any processing.
[0103] In the initial experiment, it has been shown that the FIR-shaped signal contains artifacts, which can be easily identified by the discriminator. As a result, the balance of adversarial loss is not achieved, and the learning training of the generator is insufficient. This can be solved by calculating the adversarial loss of the actually generated unshaped excitation signal.
[0104] LPC shaping by the FIR filter is only performed during the learning time. During the evaluation time, since there is no need to backpropagate the gradient, the LPC coefficients are applied as an IIR filter.
[0105] Two different DNNs are used for the extrapolation of the excitation signal. The first is based on a mixture of convolutional layers and recurrent layers, and the second is based only on a convolutional architecture. Details of the first are shown in Figure 6.
[0106] Figure 6 shows the structure in which the DNN extrapolates the excitation signal. The shape of the signal in the parentheses is given with the batch dimension omitted. T is the length of the input signal.
[0107] TIFF0007679244000007.tif145170
[0108] The purpose of the first CNN layer is to add characteristic dimensions to the one-dimensional time-domain signal. These characteristic dimensions are required by the GRU layer; otherwise, the GRU matrix will collapse into a simple vector. The CNN adds characteristic dimensions by operating kernels, usually called channels, in parallel. As a result, 256 channels are required to maintain the compatibility between the CNN layer and the GRU layer.
[0109] This makes the calculation extremely complex. This can be prevented by dividing the channels into 16 groups of 16 channels each. This is equivalent to arranging 16 layers in parallel for every 16 channels. The structure of the CNN layer (kernel size, gate activation, etc.) can be the same as that described for the convolutional BBWE above, for example. Since there are still characteristic dimensions in the output of the second GRU layer, it is compressed into a one-dimensional signal using a single convolutional kernel with a kernel size of 1.
[0110] The main factor contributing to the computational complexity lies in the matrices of the first GRU. To further reduce the complexity, these matrices can be made sparse during learning
[49] .
[0111] After repeating the initial learning with dense matrices, small-sized blocks are identified and forced to be zero. The boolean matrix stores the indices of these blocks. As learning progresses, more blocks can be set to zero until the desired sparsity is obtained. Similar to
[32] , 16x1 blocks are used, including all diagonal terms. The final proportion of elements stored in the matrix is as follows: TIFF0007679244000008.tif27170
[0112] Ignoring the computational overhead for indexing, this sparsification scheme reduces the computational cost of the GRU by 90%. Figure 7 shows one of the sparse matrices after learning. This system is called LPC-RNN-GAN.
[0113] In particular, Figure 7 shows one of the matrices from the GRU after sparsification.
[0114] The DNN based on the convolutional architecture only has the same structure as that described in Sc. III-A has three structural differences. First, the size of the CNN kernel is only 17. Second, in order to compensate for the resulting smaller receptive field, this system uses dilated convolutions with an expansion factor of 2 for each layer. Third, to reduce complexity, this system utilizes the above-mentioned grouping by dividing the channel dimension into four groups. Furthermore, it will be shown below that this can reduce the computational complexity by about one-third. This system is called LPC-CNN-GAN.
[0115] The DNN that extrapolates the LPC envelope is also a combination of a CNN layer, followed by a GRU layer, and a final CNN layer. The CNN layer has a 2D kernel with a kernel size of 3, operates on the current, past, and future frames, and is the main cause of the overall algorithm delay of the system.
[0116] The discriminator according to the embodiment will be described below.
[0117] The discriminator functions as a convolutional encoder that extracts the potential representation of the input signal and evaluates the adversarial loss. In CNN-GAN, LPC-CNN-GAN, and LPC-RNN-GAN, the same discriminator architecture composed of convolutional layers is used for adversarial learning. Stable adversarial learning is achieved by applying spectral normalization to the convolutional kernels of the discriminator layers
[50] . This type of normalization enforces the Lipschitz condition on the function learned by the discriminator. It has been found to be important for an effective and stable adversarial learning method. Since the discriminator operates in a conditional setting
[51] , the input signal includes real / fake WB speech waveforms concatenated with NB speech waveforms upsampled along the channel dimension. Figure 8 shows the discriminator. It consists of six convolutional layers with a kernel size of 32 and a stride of 2 steps. Biases are omitted. For activation, Leaky ReLU with a negative slope of 0.2 is used.
[0118] Specifically, Figure 8 shows a GAN discriminator network composed of six convolutional layers with a kernel of 32 samples operating with a stride of 2 in each layer. The numbers within the layers represent the dimensions of the input and output channels of each layer.
[0119] Since the conditional input is time-domain NB speech, the discriminator rejects speech generated with waveforms different from the original waveform. In LPCNet based on BBWE, as will be described later, there are fewer constraints on the generated waveforms. To realize a GAN with fewer constraints on the generated waveforms, a second discriminator that obtains a low-dimensional feature representation as input is evaluated. The features are the Mel-frequency cepstral coefficients (MFCCs)
[52] calculated from NB speech. This discriminator, combined with the absence of the Lp loss, tends to penalize speech that is not generated with waveforms different from the original waveform.
[0120] Here, considerations regarding the purpose of learning are presented.
[0121] The adversarial metric used in this work is the hinge loss
[53] : TIFF0007679244000009.tif12158 Here, D() is the raw output of the discriminator. Lim et. al.
[53] showed that compared with the losses used in the first GAN paper
[36] and the Wasserstein distance
[54] , the hinge loss exhibits less mode collapse and a more stable learning behavior.
[0122] In the first experiment using the proposed system, it has been shown that the hinge loss functions in the same way as the feature matching. As already observed in
[30] ,
[25] , the adversarial loss can be modified by the Lp norm calculated on samples and features. Here, the L1 norm calculated on time-domain samples and the L2 norm calculated with logarithmic Mel energy are used as the feature loss L mel The total loss learning of the generator is as follows: TIFF0007679244000010.tif12158
[0123] Below, the experimental setup will be described.
[0124] As learning materials, several publicly available speech databases
[55] ,
[56] ,
[57] and speech items in other languages were used. A total of 13 hours of learning materials were used, and all of them were resampled to a sampling frequency of 16 kHz. The silent parts of the learning data were removed using voice-activation-detection
[58] . The input signal of NB was encoded at 10.2 kbps in AMR-NB. The target clean speech signal was pre-emphasized with the primary filter E shown below. TIFF0007679244000011.tif14159
[0125] Inverse (de-emphasis) filter D TIFF0007679244000012.tif18159 was applied to the generated voice. This is to correct the slope of the spectrum of the voice where high frequencies may not be emphasized enough in the generated voice. The 12th-order LPC envelope is extracted using the Levinson recursion after calculating the autocorrelation in the time domain for a 128-sample frame windowed with a Hann window. Then, as described for the LPC-GAN above, it is converted to, for example, an FIR filter. The DNN is trained in batches of 8 items, and each item contains 1 second of voice.
[0126] The optimization algorithms for both the generator and the discriminator are Adam
[59] . The learning rate of the generator is 0.0001, and the learning rate of the discriminator is 0.0004. For more stable adversarial loss, the coefficients (beta parameters) used to calculate the moving average of the gradient and its square are set to 0.5 and 0.99, respectively. Since the RNN of the LPC-RNN-GAN (see the description of the LPC-GAN above) usually learns slower than the CNN, the learning rates of the generator and the discriminator are set to 0.0001. The beta parameters for training the generator are set to 0.7 and 0.99. The coefficient λ that controls the amount of adversarial loss in Equation (10) is set to 0.0015. The sparsification of the GRU layer starts from the 160th batch, and the final sparsification is achieved at the 10000th batch. All CNN layers are trained with batch normalization to speed up learning and prevent the network from falling into mode collapse.
[0127] The additional frame rate network for extrapolating the LPC coefficients in the LSF domain has 10 CNN layers, followed by a single GRU and a final CNN layer. The initial CNN layers are 2D convolutions using a kernel size of 3x3, 16 channels, a tanh activation function, and residual connections. The matrix size of the GRU is 16x16, and the final convolutional layer is 5 channels, with the number of missing LSF coefficients concatenated to the NB LSF coefficients to form the WB LSF coefficients.
[0128] Below, the presented system is compared with the LCPNet based on the BBWE of
[33] . In contrast to the published systems, the DNN used for extrapolating the LPC envelope is adversarially trained here. For this purpose, the same discriminator architecture is used with only the input dimension adapted.
[0129] All DNNs are implemented and trained using PyTorch
[60] .
[0130] Below, it is considered from the perspective of evaluation. The system provided based on the embodiment is compared with the previously published systems by objective metrics and subjective metrics by listening tests. An estimate of the computational complexity is given and compared with the state-of-the-art speech coding technology. The objective and subjective tests show that the proposed system provides substantially better quality than the previous techniques. The system according to the embodiment is shown to reduce the Word Error Rate of the speech recognition system.
[0131] The perceived quality of the presented BBWE is evaluated by objective metrics that have been used so far to access the quality of speech and subjective metrics by listening tests. Furthermore, the latency and computational complexity of the algorithms are given for each BBWE. The correlation between the objective and subjective results examines whether the subjective evaluation is predictive enough.
[0132] Regarding the computational complexity, the computational complexity of the proposed BBWE is the estimated value of WMOPS (Weighted Million Operations per Second) per second for each audio sample. WMOPS is an ITU unit for calculating the computational complexity of standardized speech processing tools
[61] . Additions (ADD), multiplications (MUL), and multiply-accumulate (MAC) operations are each counted as one operation, while complex operations such as tanh, sigmoid, or softmax operations are each counted as 25 operations. Below, a numerical value is calculated for each audio sample. This numerical value is multiplied by the sampling frequency to calculate the estimated value of WMOPS. This should be regarded as a rough estimate that does not take into account the advantages of today's parallel processing architectures. The results are summarized in Table 1 together with the computational complexity of the state-of-the-art standardized audio codec EVS [2],
[62] .
[0133] In particular, Table 1 shows the computational complexity and algorithmic latency of the systems provided by several embodiments, LPCNet-BBWE
[33] and EVS [2],
[62] (EVS is a state-of-the-art standardized audio codec). WMOPS is an ITU standard
[61] for calculating computational complexity and is calculated at a sampling frequency of 16 kHz.
[0134]
Table 1
[0135] TIFF0007679244000014.tif70170
[0136] Regarding the computational complexity of LPC-RNN-GAN and LPC-CNN-GAN: As described for the above LPC-GAN, this system has an initial CNN layer that divides a one-dimensional signal into 256 channels. These layers are the same CNN layers as above, but differ in that the channels are grouped into blocks as described for the above LPC-GAN. Here, a total of 256 channels are grouped into 16 blocks of 16 channels each. This is the same as having 16 CNN layers with 16 channels in parallel.
[0137] The operation of one RNN layer for a single audio sample is shown in Equation (5). Here, M i is the input dimension, and M h is the output (or hidden) dimension. Then, for the calculation of the reset gate and the update gate (the first two lines of the equation), MAC operations of M i *M h *2 and a sigmoid operation of M h are required. For the new gate (the third line of the equation), MAC operations of M i *M h *2 + M h and tangent hyperbolicus operations of M h are required. Finally, for the output (the last line), a MAC operation of M h *2 is required. In the first large GRU layer, because a sparsified matrix is used (see the description regarding the above LPC-GAN), the operations are calculated with a reduced matrix size. The overhead due to additional addressing-operations is ignored. In the first GRU, all matrices are square with M i = M h = 256, and in the second GRU, M i = 256 and M h = 32.
[0138] The last CNN layer only sums the output dimensions and requires 32 ADD operations. The computational complexity of LPC-CNN-GAN is calculated as having four such networks with a channel dimension of only 8 in parallel as described above.
[0139] During evaluation, the LPC filter is applied as a 12-tap IIR filter and requires 12 MAC operations per sample. The conversion from LPC to LSF coefficients and its inverse are ignored here because these conversions are performed on a frame basis and are expected to contribute little to the overall complexity. Table 1 above summarizes the number of operations using the parameterizations used.
[0140] Regarding algorithmic delay, algorithmic delay is the theoretical delay in milliseconds between the input speech and the processed output speech caused by the block processing of speech samples. The time of the CPU or GPU is not considered. The numerical values are summarized in Table 1 above.
[0141] TIFF0007679244000015.tif33170
[0142] Regarding LPC-RNN-GAN and LPC-CNN-GAN, the causes of the algorithmic delay in these systems are the initial convolutional layer and LPC processing. The GRU layer does not cause any algorithmic delay. The kernel size of the four convolutional layers is 16 taps, and since 16 taps are calculated for future samples, a delay of 4 milliseconds occurs. Therefore, the algorithmic delay of the LPC processing due to the windowed autocorrelation function is 15 milliseconds. Since this block processing is independent of the convolutional layer, the total algorithmic delay of the entire system is 15 milliseconds. LPC-CNN-GAN has the same algorithmic delay as CNN-GAN because it uses a kernel with an expansion of 2 and half the size of CNN-GAN.
[0143] A listening test by a human listener is the ultimate basis for evaluating (e.g., objective) perceptual quality, but it requires a significant amount of effort to implement. Objective metrics are easily usable alternatives. Here, four different measurements are used: Perceptual Objective Listening Quality Analysis, Fréchet Deep Speech Distance, Word Error Rate, and Short-Time Objective Intelligibility measure. All measurements except for the Word Error Rate are calculated on a multi-lingual, multi-speaker database of about one hour that is not part of the training set.
[0144] Perceptual Objective Listening Quality Analysis (POLQA) is a standardized method aimed at predicting the perceptual quality of an audio signal encoded on the same Mean Opinion Score (MOS) used in listening tests
[63] . The estimation results are summarized in Figure 9, showing that LPC-RNN-GAN achieved the highest evaluation, followed by CNN-GAN.
[0145] In particular, Figure 9 shows the Perceptual Objective Listening Quality Analysis (POLQA) of various BBWEs within a 95% confidence interval. A larger value means better quality.
[0146] Evaluating the quality of audio or images generated by a GAN is a difficult task. In typical use cases, since a GAN generates items from noise, there is no reference for comparison, so metrics based on the Lp norm cannot be used.
[0147] The Fréchet deep speech distance (FDSD) can be considered, for example. A common objective measure for evaluating the quality of images created by a GAN is the Fréchet inception distance (FID)
[64] . This metric is calculated based on the outputs of different DNNs trained to classify images or speech. In contrast to generative modeling, image and speech classification (recognition) is already quite sophisticated, and it may be possible to obtain an estimate of quality from the entropy of the output of a DNN that classifies the generated data. An item that is strongly classified as one class rather than all other classes indicates high quality, and the conditional probability of the generated item should have low entropy. Furthermore, since the GAN needs to generate a wide variety of items (without mode collapse), it is preferable for the integral value of the limiting probability distribution of the classification output to have high entropy. The inception distance (ID) of
[65] formulates this mathematically. Heusel et.al.
[65] improved this by also using the distribution of the classification results of actual data based on the Fréchet distance. TIFF0007679244000016.tif13170
[0148] TIFF0007679244000017.tif34169
[0149] Figure 10 shows the Fréchet deep speech distance (FDSD) of various BBWE. The smaller the value, the higher the quality.
[0150] Regarding the word error rate, BBWE can not only improve the perceptual quality, but also the speech intelligibility [5], [6], and even the performance of the automatic speech recognition (ASR) system. The state-of-the-art ASR systems are based on DNNs trained on speech with a fixed sampling frequency (mainly 16 kHz). As a result, when the speech is coded with the NB codec, the performance of such systems drops significantly. Evaluate the impact of AMR-NB-based speech coding on the word error rate (WER) of state-of-the-art ASR systems and how BBWE can mitigate this impact. The ASR system used here is an open implementation of Mozilla's RNN based on a deep speech system
[68] with a connectionist temporal classification (CTC) loss
[69] trained on a general speech multilingual speech corpus
[70] . The evaluation is performed on the evaluation set of this database. The WER metric is evaluated at the word level of the transcribed speech and is calculated as follows: TIFF0007679244000018.tif19169 Here, S is the number of substitutions, D is the number of deletions, I is the number of insertions, and C is the number of correctly transcribed words.
[0151] Figure 11 shows the word error rate (WER) and character error rate (CER) of various BBWEs. The smaller the value, the higher the performance. In particular, Figure 11 shows the ASR performance of AMR-NB and various BBWEs, along with the character error rate (CER), which is calculated in the same way as WER but at the character level instead of the word level.
[0152] Table 2 shows an example of one of the least performing items. Interestingly, the uncoded items on average give better results, but there are no outliers that give results worse than 0:6 when using AMR-NB-coded items from the database. The items processed with BBWE improve the average WER, but also produce outliers with a WER of 8:0 or higher.
[0153]
Table 2
[0154] Regarding the Short-Time Objective intelligibility measure (STOI), the Short-Time Objective intelligibility measure (STOI) is defined as an estimated value of the linear correlation coefficient between the clean time energy envelope and the BBWE-processed audio subbands. These subbands are obtained from dividing the audio signal into frames of length 256 samples with a Hamming window with 50% overlap, and each frame is zero-padded to 512 samples and calculated based on the time-frequency representation obtained by Fourier transform. The 15 1 / 3-octave bands are calculated by averaging the DFT bins.
[0155] Originally, this measurement value is calculated for audio sampled at a sampling frequency of 10 kHz. Since it is evaluating the quality of WB audio, this measurement value is extended to 16 kHz.
[0156] Figure 12 shows the results of the presented system.
[0157] In particular, Figure 12 shows the Short-Time Objective intelligibility measure (STOI) of the presented system. The smaller the value, the lower the quality.
[0158] According to this measurement value, LPC-RNN-GAN exhibits the best performance, followed by LPC-CNN-GAN.
[0159] Below, the subjective perceptual quality will be considered.
[0160] To finally judge the perceptual quality of the proposed system, a MUSHRA listening test
[71] was conducted. According to the MUSHRA methodology, the test items include the reference marked as such, the hidden reference, and the AMR-NB coded signal that functions as an anchor. Twelve experienced listeners participated in the test. The speech items used in the test are about 10 seconds long and are neither part of the learning nor the test set. The items are recorded with voices of native speakers of Chinese, English, French, German, and Spanish. The results are shown in box plots showing the mean value and 95% confidence interval for each item in Figure 13 and for all items averaged in Figure 14. The results are shown as bar graphs in Figure 15.
[0161] In particular, Figure 13 shows the results of the listening test evaluating various BBWEs as box plots with 95% confidence intervals for each item.
[0162] Figure 14 shows the results of the listening test evaluating various BBWEs as bar graphs with 95% confidence intervals averaged over all items.
[0163] Figure 15 shows the results of the listening test evaluating various BBWEs with the evaluations from each user shown as a warm plot.
[0164] The system marked as CNN-feat-cond is a CNN-GAN trained with a discriminator having conditional input based on the functions described above regarding the discriminator. The L1 loss is also removed from the learning objective.
[0165] The results show that all the presented systems significantly improve the quality of the AMR-NB speech for all items. Except for CNN-feat-cond, none of the presented systems is significantly better than the others. Tendentially, the best system is LPC-CNN-GAN, which is also significantly better than the CNN-feat-cond system.
[0166] Examining the results of a single item, it can be seen that the quality depends quite a lot on the item. The LPC-CNN-GAN is not always a system that exhibits the best performance. In the case of items for Spanish women, German women, and two items for men, the LPCNet-based system shows the best performance. In the case of Chinese male items, the LPC-RNN-GAN shows the best performance, and in the case of Spanish male items, the CNN-GAN shows the best performance. The CNN-GAN often has the fewest noisy artifacts, but often cannot reconstruct fricative sounds well.
[0167] In the LPCNet-based system, the quality variation is particularly large. In this system, while very high quality can be obtained, serious artifacts such as a clicking feeling and pitch instability may occur. On the other hand, the GAN-based system is not affected by such serious artifacts, but is affected by broadband crackling noise. The LPCNet-based system, and in some cases the function-adjustment-based system, change the characteristics of the voice because there are fewer constraints imposed on the waveforms generated by both systems. In the MUSHRA test, this can result in lower scores, like in various test methods such as the absolute category rating (ACR) test without a reference being given.
[0168] To check to what extent the objective measure reflects the subjective evaluation, the correlation with the MOS value of the listening test is examined. For a fair comparison, all measurements are normalized to zero mean and standard deviation. FDSD, WER, and CER give low values for better quality estimates, so these values are initially negated.
[0169] Figure 16 shows the normalized values, and Table 3 shows the correlation values.
[0170] In particular, Figure 16 shows the normalized objective and subjective measurement values.
[0171]
Table 3
[0172] STOI shows the highest correlation with the MOS value, followed by POLQA, WER, and CER. WER. It can be seen that WER is the only measurement with the same order as the result of the listening test. The difference between the WER value and the FDSD value is strange because both measurement values are based on the outputs of similar networks (DeepSpeech and DeepSpeech 2).
[0173] Two basically different approaches for performing BBWE, namely the GAN model and the autoregressive model, were compared. Both approaches rely on generative models that can model complex data distributions such as the distribution of audio in the time domain, and neither approach is troubled by the problem of smoothing.
[0174] Compared with state-of-the-art models such as WaveNet®
[35] , both approaches have a medium level of computational complexity.
[0175] BBWE based on LPCNet is the model with the lowest computational complexity. The main reason for the reduced complexity is that this model has fewer constraints on the generated waveforms. The waveforms generated by LPCNet may be very different from the original waveforms, while BBWE based on GAN maintains the original waveforms through a mixture of conditioning and adversarial loss and L1 loss. Unfortunately, even when the conditioning was changed to functional conditioning and the L1 loss was removed, the quality of the generated speech did not improve.
[0176] LPC-RNN-GAN and LPC-CNN-GAN differ in the DNN used for extrapolation of the excitation signal. The former is based on a mixture of CNN and RNN, and the latter uses only CNN.
[0177] The computational complexity of both DNNs is approximately the same. There is no significant difference in performance, but the performance of LPC-CNN-GAN is tendentially better. Furthermore, the learning time of CNN is short and is not so affected by the adjustment of hyperparameters. LPC-RNN-GAN is the first to successfully apply sparsification in the context of GAN learning.
[0178] When correlating the results of the listening test with objective measurements, ambiguous results are obtained. The authors of
[66] showed that the FDSD measurement works well for estimating the quality of adversarially generated speech, but here the small differences between the presented systems cannot be accessed. The metrics that most closely correlate with the subjective results are the STOI and WER metrics.
[0179] Although several aspects have been described in the context of the apparatus so far, these aspects are also descriptions of the corresponding methods, and it is clear that the blocks or apparatuses correspond to the steps of the method or the features of the steps of the method. Similarly, the aspects described in the context of the method steps are also descriptions of the corresponding blocks or items or functions of the corresponding apparatuses. Some or all of the method steps may be performed (or used) by a hardware apparatus such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0180] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software, or at least partially in hardware, or at least partially in software. The implementation can be, for example, a digital storage medium such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, FLASH memory, etc., having electronically readable control signals stored thereon and cooperating (or capable of cooperating) with a programmable computer system such that each method is executed. Thus, the digital storage medium can be computer-readable.
[0181] Some embodiments according to the present invention include a data carrier having electronically readable control signals and capable of cooperating with a programmable computer system such that one of the methods described herein is executed.
[0182] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code, the program code being operative to execute one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0183] Other embodiments are those in which a computer program for executing one of the methods described herein is stored on a machine-readable carrier.
[0184] In other words, embodiments of the method of the present invention are thus computer programs having program code for executing one of the methods described herein, in the case where the computer program is executed on a computer.
[0185] Accordingly, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitory.
[0186] Accordingly, a further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals can be configured to be transmitted via a data communication connection such as, for example, the Internet.
[0187] A further embodiment comprises processing means, such as, for example, a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0188] A further embodiment comprises a computer having installed thereon a computer program for performing one of the methods described herein.
[0189] A further embodiment according to the present invention comprises an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver is, for example, a computer, a mobile device, a storage device, etc. The apparatus or system can, for example, constitute a file server for transmitting the computer program to the receiver.
[0190] In some embodiments, a programmable logic device (e.g., a field programmable gate array) can be used to perform some or all of the functionality of the methods described herein. In some embodiments, the field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0191] The devices described herein can be implemented using a hardware device, a computer, or a combination of a hardware device and a computer.
[0192] The methods described herein can be performed using a hardware device, a computer, or a combination of a hardware device and a computer.
[0193] The above-described embodiments are merely illustrative of the principles of the present invention. It will be understood that improvements and modifications to the arrangements and details described herein will be apparent to those skilled in the art. Accordingly, it is intended that the present disclosure be limited only by the scope of the impending claims and not by the specific details presented by the description and illustration of the embodiments.
[0194] References [1] International Telecommunication Union, "Pulse code modulation (pcm) of voice frequencies," ITU-T Recommendation G.711, November 1988. [2] S. Bruhn, H. Pobloth, M. Schnell, B. Grill, J. Gibbs, L. Miao, K. Jaervinen, L. Laaksonen, N. Harada, N. Naka, S. Ragot, S. Proust, T. Sanda, I. Varga, C. Greer, M. Jelinek, M. Xie, and P. Usai, "Standardization of the new 3GPP EVS codec," in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, 2015, pp. 5703-5707. [Online]. Available: https: / / doi.org / 10.1109 / ICASSP.2015.7179064 [3] S. Disch, A. Niedermeier, C. R. Helmrich, C. Neukam, K. Schmidt, R. Geiger, J. Lecomte, F. Ghido, F. Nagel, and B. Edler, "Intelligent gap filling in perceptual transform coding of audio," in Audio Engineering Society Convention 141, Los Angeles, Sep 2016. [Online]. Available: http: / / www.aes.org / e-lib / browse.cfm?elib=18465 [4] 3GPP, "TS 26.090, Mandatory Speech Codec speech processing functions; Adaptive Multi-Rate (AMR) speech codec; Transcoding functions," 1999. [5] P. Bauer, R. Fischer, M. Bellanova, H. Puder, and T. Fingscheidt, "On improving telephone speech intelligibility for hearing impaired persons," in Proceedings of the 10. ITG Conference on Speech Communication, Braunschweig, Germany, September 26-28, 2012, 2012, pp. 1-4. [Online]. Available: http: / / ieeexplore.ieee.org / document / 6309632 / [6] P. Bauer, J. Jones, and T. Fingscheidt, "Impact of hearing impairment on fricative intelligibility for artificially bandwidth-extended telephone speech in noise," in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2013, Vancouver, BC, Canada, May 26-31, 2013, 2013, pp. 7039-7043. [Online]. Available: https: / / doi.org / 10.1109 / ICASSP.2013.6639027 [7] J. Abel, M. Kaniewska, C. Guillaume, W. Tirry, H. Pulakka, V. Myllylae, J. Sjoberg, P. Alku, I. Katsir, D. Malah, I. Cohen, M. A. T. Turan, E. Erzin, T. Schlien, P. Vary, A. H. Nour-Eldin, P. Kabal, and T. Fingscheidt, "A subjective listening test of six different artificial bandwidth extension approaches in english, chinese, german, and korean," in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2016, Shanghai, China, March 20-25, 2016, 2016, pp. 5915-5919. [Online]. Available: https: / / doi.org / 10.1109 / ICASSP.2016.7472812 [8] K. Schmidt and B. Edler, "Blind bandwidth extension based on convolutional and recurrent deep neural networks," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), April 2018, pp. 5444-5448. [9] K. Schmidt, "Neubildung von unterdrueckten Sprachfrequenzen durch ein nichtlinear verzerrendes Glied," Dissertation, Techn. Hochsch. Berlin, 1933.
[10] M. Schroeder, "Recent progress in speech coding at bell telephone laboratories," in Proceedings of the third international congress on acoustics, Stuttgart, 1959.
[11] Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, "Gradient-based learning applied to document recognition," Proceedings of the IEEE, vol. 86, no. 11, pp. 2278-2324, Nov 1998.
[12] C. Ledig, L. Theis, F. Huszar, J. Caballero, A. P. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi, "Photo-realistic single image super-resolution using a generative adversarial network," CoRR, vol. abs / 1609.04802, 2016. [Online]. Available: http: / / arxiv.org / abs / 1609.04802
[13] X. Li, V. Chebiyyam, and K. Kirchhoff, "Speech audio super-resolution for speech recognition," in Interspeech 2019, 20th Annual Conference of the International Speech Communication Association, Graz, Austria, September 15-19, 2019, 09 2019.
[14] P. Jax and P. Vary, "Wideband extension of telephone speech using a hidden markov model," in 2000 IEEE Workshop on Speech Coding. Proceedings., 2000, pp. 133-135.
[15] K. Schmidt and B. Edler, "Deep neural network based guided speech bandwidth extension," in Audio Engineering Society Convention 147, Oct 2019. [Online]. Available: http: / / www.aes.org / e-lib / browse.cfm? elib=20627
[16] H. Carl and U. Heute, "Bandwidth enhancement of narrow-band speech signals," in Signal Processing VII: Theories and Applications: Proceedings of EUSIPCO-94 Seventh European Signal Processing Conference, September 1994, pp. 1178-1181.
[17] H. Pulakka and P. Alku, "Bandwidth extension of telephone speech using a neural network and a filter bank implementation for highband mel spectrum," IEEE Trans. Audio, Speech & Language Processing, vol. 19, no. 7, pp. 2170-2183, 2011. [Online]. Available: https: / / doi.org / 10.1109 / TASL.2011.2118206
[18] K. Li and C. Lee, "A deep neural network approach to speech bandwidth expansion," in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, 2015, pp. 4395-4399. [Online]. Available: https: / / doi.org / 10.1109 / ICASSP.2015.7178801
[19] P. Bauer, J. Abel, and T. Fingscheidt, "Hmm-based artificial bandwidth extension supported by neural networks," in 14th International Workshop on Acoustic Signal Enhancement, IWAENC 2014, Juan-les-Pins, France, September 8-11, 2014, 2014, pp. 1-5. [Online]. Available: https: / / doi.org / 10.1109 / IWAENC.2014.6953304
[20] J. Sautter, F. Faubel, M. Buck, and G. Schmidt, "Artificial bandwidth extension using a conditional generative adversarial network with discriminative training," in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 7005-7009.
[21] J. Abel, M. Strake, and T. Fingscheidt, "A simple cepstral domain dnn approach to artificial speech bandwidth extension," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), April 2018, pp. 5469-5473.
[22] J. Abel and T. Fingscheidt, "Artificial speech bandwidth extension using deep neural networks for wideband spectral envelope estimation," IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 1, pp. 71-83, 2018.
[23] Z. Ling, Y. Ai, Y. Gu, and L. Dai, "Waveform modeling and generation using hierarchical recurrent neural networks for speech bandwidth extension," IEEE / ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 5, pp. 883-894, May 2018.
[24] A. Gupta, B. Shillingford, Y. M. Assael, and T. C. Walters, "Speech bandwidth extension with wavenet," ArXiv, vol. abs / 1907.04927, 2019.
[25] S. Kim and V. Sathe, "Bandwidth extension on raw audio via generative adversarial networks," 2019.
[26] Y. Dong, Y. Li, X. Li, S. Xu, D. Wang, Z. Zhang, and S. Xiong, "A time-frequency network with channel attention and non-local modules for artificial bandwidth extension," in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 6954-6958.
[27] J. Makhoul and M. Berouti, "High-frequency regeneration in speech coding systems," in ICASSP '79. IEEE International Conference on Acoustics, Speech, and Signal Processing, April 1979, pp. 428-431.
[28] N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, "Efficient neural audio synthesis," CoRR, vol. abs / 1802.08435, 2018. [Online]. Available: http: / / arxiv.org / abs / 1802. 08435
[29] S. Li, S. Villette, P. Ramadas, and D. J. Sinder, "Speech bandwidth extension using generative adversarial networks," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), April 2018, pp. 5029-5033.
[30] S. E. Eskimez, K. Koishida, and Z. Duan, "Adversarial training for speech super-resolution," IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 2, pp. 347-358, 2019.
[31] X. Hao, C. Xu, N. Hou, L. Xie, E. S. Chng, and H. Li, "Time-domain neural network approach for speech bandwidth extension," in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 866-870.
[32] J. Valin and J. Skoglund, "Lpcnet: Improving neural speech synthesis through linear prediction," in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 5891-5895.
[33] K. Schmidt and B. Edler, "Blind bandwidth extension of speech based on lpcnet," in 2020 28th European Signal Processing Conference (EUSIPCO).
[34] L. Rabiner and R. Schafer, Digital Processing of Speech Signals. Englewood Cliffs: Prentice Hall, 1978.
[35] A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, "Wavenet: A generative model for raw audio," in The 9th ISCA Speech Synthesis Workshop, Sunnyvale, CA, USA, 13-15 September 2016, 2016, p. 125.
[36] I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, "Generative adversarial networks," 2014.
[37] Y. Gu and Z. Ling, "Waveform modeling using stacked dilated convolutional neural networks for speech bandwidth extension," in Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20-24, 2017, 2017, pp. 1123-1127. [Online]. Available: http: / / www.isca-speech.org / archive / Interspeech 2017 / abstracts / 0336.html
[38] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Computation, vol. 9, no. 8, pp. 1735-1780, 1997. [Online]. Available: https: / / doi.org / 10.1162 / neco.1997.9.8.1735
[39] Y. Gu, Z. Ling, and L. Dai, "Speech bandwidth extension using bottleneck features and deep recurrent neural networks," in Interspeech 2016, 17th Annual Conference of the International Speech Communication Association, San Francisco, CA, USA, September 8-12, 2016, 2016, pp. 297-301. [Online]. Available: https: / / doi.org / 10.21437 / Interspeech.2016-678
[40] J. Chung, C. Guelcehre, K. Cho, and Y. Bengio, "Empirical evaluation of gated recurrent neural networks on sequence modeling," NIPS Deep Learning workshop, Montreal, Canada, 2014. [Online]. Available: http: / / arxiv.org / abs / 1412.3555
[41] A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu, "Conditional image generation with pixelcnn decoders," CoRR, vol. abs / 1606.05328, 2016. [Online]. Available: http: / / arxiv.org / abs / 1606.05328
[42] W. B. Kleijn, F. S. C. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, "Wavenet based low rate speech coding," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), April 2018, pp. 676-680.
[43] Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, "Fftnet: A real-time speaker-dependent neural vocoder," in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), April 2018, pp. 2251-2255.
[44] J.-M. Valin and J. Skoglund, "A real-time wideband neural vocoder at 1.6 kb / s using lpcnet," ArXiv, vol. abs / 1903.12087, 2019.
[45] A. Mustafa, A. Biswas, C. Bergler, J. Schottenhamml, and A. Maier, "Analysis by Adversarial Synthesis - A Novel Approach for Speech Vocoding," in Proc. Interspeech, 2019, pp. 191-195. [Online]. Available: http: / / dx.doi.org / 10.21437 / Interspeech.2019-1195
[46] T. Salimans and D. P. Kingma, "Weight normalization: A simple reparameterization to accelerate training of deep neural networks," in Advances in NeurIPS, 2016, pp. 901-909.
[47] K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770-778.
[48] Yao Tianren, Xiang Juanjuan, and Lu Wei, "The computation of line spectral frequency using the second chebyshev polynomials," in 6th International Conference on Signal Processing, 2002., vol. 1, Aug 2002, pp. 190-192 vol.1.
[49] N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, "Efficient neural audio synthesis," 2018.
[50] T. Miyato, T. Kataoka, M. Koyama, and Y. Yoshida, "Spectral normalization for generative adversarial networks," 2018.
[51] M. Mirza and S. Osindero, "Conditional generative adversarial nets," ArXiv, vol. abs / 1411.1784, 2014.
[52] A. Salman, E. Muhammad, and K. Khurshid, "Speaker verification using boosted cepstral features with gaussian distributions," in 2007 IEEE International Multitopic Conference, 2007, pp. 1-5.
[53] J. H. Lim and J. C. Ye, "Geometric gan," 2017.
[54] M. Arjovsky, S. Chintala, and L. Bottou, "Wasserstein gan," 2017.
[55] C. Veaux, J. Yamagishi, and K. Macdonald, "Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit," 2017.
[56] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, "Librispeech: An ASR corpus based on public domain audio books," in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, 2015, pp. 5206-5210. [Online]. Available: https: / / doi.org / 10.1109 / ICASSP.2015.7178964
[57] M. Soloducha, A. Raake, F. Kettler, and P. Voigt, "Lombard speech database for german language," in Proc. DAGA 2016 Aachen, 03 2016.
[58] "Webrtc vad v2.0.10," https: / / webrtc.org.
[59] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," CoRR, vol. abs / 1412.6980, 2014.
[60] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, "Pytorch: An imperative style, high-performance deep learning library," in Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett, Eds. Curran Associates, Inc., 2019, pp. 8024-8035. [Online]. Available: http: / / papers.neurips.cc / paper / 9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf
[61] ITU-T Study Group 12, Software tools for speech and audio coding standardization, Geneva, 2005.
[62] G. T. 26.445, "EVS codec; detailed algorithmic description; technical specification, release 12," Sep. 2014.
[63] ITU-T Study Group 12, P.863 : Perceptual objective listening quality prediction, Geneva, 2018.
[64] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, G. Klambauer, and S. Hochreiter, "Gans trained by a two time-scale update rule converge to a nash equilibrium," CoRR, vol. abs / 1706.08500, 2017. [Online]. Available: http: / / arxiv.org / abs / 1706.08500
[65] T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford,X. Chen, and X. Chen, "Improved techniques for training gans," in Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016, pp. 2234-2242. [Online]. Available: https: / / proceedings.neurips.cc / paper / 2016 / file / 8a3363abe792db2d8761d6403605aeb7-Paper.pdf
[66] M. Binkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, "High fidelity speech synthesis with adversarial networks," CoRR, vol. abs / 1909.11646, 2019. [Online]. Available: http: / / arxiv.org / abs / 1909.11646
[67] D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. H. Engel, L. Fan, C. Fougner, T. Han, A. Y. Hannun, B. Jun, P. LeGresley, L. Lin, S. Narang, A. Y. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, Y. Wang, Z. Wang, C. Wang, B. Xiao, D. Yogatama, J. Zhan, and Z. Zhu, "Deep speech 2: End-to-end speech recognition in english and mandarin," CoRR, vol. abs / 1512.02595, 2015. [Online]. Available: http: / / arxiv.org / abs / 1512.02595
[68] A. Y. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y. Ng, "Deep speech: Scaling up end-to-end speech recognition," CoRR, vol. abs / 1412.5567, 2014. [Online]. Available: http: / / arxiv.org / abs / 1412.5567
[69] A. Graves, S. Fernandez, and F. Gomez, "Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks," in In Proceedings of the International Conference on Machine Learning, ICML 2006, 2006, pp. 369-376.
[70] R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, "Common voice: A massively-multilingual speech corpus," CoRR, vol. abs / 1912.06670, 2019. [Online]. Available: http: / / arxiv.org / abs / 1912.06670
[71] ITU-R, Recommendation BS.1534-1 Method for subjective assessment of intermediate sound quality (MUSHRA), Geneva, 2003.
Claims
1. 1. An apparatus for processing an audio input signal by performing bandwidth extension of the audio input signal to obtain an audio output signal, the apparatus comprising: a signal envelope extrapolator (120) including a first neural network (125) configured to receive a plurality of signal envelope samples of the audio input signal as input values of the first neural network (125) and to determine a plurality of extrapolated signal envelope samples as output values of the first neural network (125); an excitation signal extrapolator (130) configured to receive a plurality of samples of an excitation signal of the audio input signal and to determine a plurality of extrapolated excitation signal samples; a combiner (140) configured to generate the audio output signal in dependence on the plurality of extrapolated signal envelope samples and the plurality of extrapolated excitation signal samples, such that the audio output signal is bandwidth extended with respect to the audio input signal; Including, the input values of the first neural network (125) being a first plurality of line spectral frequencies of the audio input signal, and the first neural network (125) being configured to determine as the output values of the first neural network (125) a second plurality of line spectral frequencies of the audio output signal, each of one or more of the second plurality of line spectral frequencies being associated with a frequency that is greater than any frequency associated with any of the first plurality of line spectral frequencies; the excitation signal extrapolator (130) comprises a second neural network (135) configured to receive a plurality of samples of the excitation signal and / or the audio input signal and / or a shaped version of the audio input signal as input values of the second neural network (135) and configured to determine the plurality of extrapolated excitation signal samples as output values of the second neural network (135); the input values of the second neural network (135) are a first plurality of time-domain signal samples of the excitation signal of the audio input signal and / or the audio input signal and / or a shaped version of the audio input signal, and wherein the second neural network (135) is configured to determine the output values of the second neural network (135) such that the plurality of extrapolated excitation signal samples are a second plurality of time-domain signal samples of an extended time-domain excitation signal that has a bandwidth extended relative to the excitation signal of the audio input signal.
2. 2. The apparatus of claim 1, wherein upon training the first neural network, the signal envelope extrapolator is configured to convert a plurality of linear predictive coding coefficients derived from an original audio signal into finite impulse response filter coefficients by calculating an impulse response and truncating the impulse response.
3. 3. The apparatus of claim 2, wherein upon training the first neural network (125), the signal envelope extrapolator (120) is configured to feed back an error between the audio output signal and the original audio signal or a gradient of the error.
4. the first neural network (125) is trained using a first discriminator neural network, and upon training the first neural network (125), the first neural network (125) and the first discriminator neural network are configured to operate as a generative adversarial network; During training of the first neural network (125), the first discriminator neural network is configured to receive the output values of the first neural network (125) as input values of the first discriminator neural network, or configured to receive derived values derived from the output values of the first neural network (125) as the input values of the first discriminator neural network; 2. The apparatus of claim 1, wherein upon receiving the input values of the first discriminator neural network, the first discriminator neural network is configured to determine, as an output of the first discriminator neural network, a first quality indication of the input values of the first discriminator neural network, and the first neural network (125) is configured to learn in dependence on the first quality indication.
5. 5. The apparatus of claim 4, wherein upon receiving the input values of the first discriminator neural network, the first discriminator neural network is configured to determine the quality indication such that the quality indication is indicative of the probability that the input values of the first discriminator neural network relate to a recorded speech signal rather than an artificially generated speech signal, or such that the quality indication is indicative of a value estimating whether the output values of the first discriminator neural network relate to a recorded signal or to an artificially generated signal.
6. 5. The apparatus of claim 4, wherein the first neural network (125) or the second neural network (135) trains using a loss function that depends on the quality indication determined by the first discriminator neural network.
7. The apparatus of claim 6 , wherein the loss function relies on a hinge loss, or a Wasserstein distance, or an entropy-based loss.
8. The loss function is The hinge loss L defined as hinge Depends on 8. The apparatus of claim 7, wherein D() denotes the output of the first discriminator neural network.
9. The apparatus of claim 6 , wherein the loss function depends on an additional Lp-loss.
10. The apparatus of claim 3 , wherein the first discriminator neural network is trained using recorded speech.
11. the second neural network (135) is trained using a second discriminator neural network, and during training of the second neural network (135), the second neural network (135) and the second discriminator neural network are configured to operate as a second generative adversarial network; During training of the second neural network (135), the second discriminator neural network receives as input values of the second discriminator neural network: configured to receive the output values of the second neural network (135) or derived values derived from the output values of the second neural network (135) as the input values of the second discriminator network; and / or The output of the combiner (140) configured to receive 2. The apparatus of claim 1, wherein upon receiving the input values of the second discriminator neural network, the second discriminator neural network is configured to determine, as an output of the second discriminator neural network, a second quality indication of the input values of the second discriminator neural network, wherein the second neural network (135) is configured to learn in dependence on the second quality indication.
12. The apparatus of claim 1 , further comprising a signal analyzer (110) configured to generate the plurality of samples of the signal envelope of the audio input signal and the plurality of samples of the excitation signal of the audio input signal from the audio input signal.
13. The apparatus of claim 1 , wherein the first neural network (125) comprises one or more convolutional neural networks.
14. The apparatus of claim 1 , wherein the first neural network (125) comprises one or more deep neural networks.
15. 2. The apparatus of claim 1, wherein the audio input signal is a narrowband audio input signal and / or the audio output signal is a wideband audio output signal.
16. 1. A method for processing an audio input signal to obtain an audio output signal by performing bandwidth extension of the audio input signal, the method comprising the steps of: receiving a plurality of signal envelope samples of the audio input signal as input values of a first neural network and determining a plurality of extrapolated signal envelope samples as output values of the first neural network; receiving a plurality of samples of an excitation signal of the audio input signal and determining a plurality of extrapolated excitation signal samples; generating the audio output signal in dependence on the plurality of extrapolated signal envelope samples and the plurality of extrapolated excitation signal samples, such that the audio output signal is bandwidth extended with respect to the audio input signal; Including, the input values of the first neural network are a first plurality of line spectral frequencies of the audio input signal, and the first neural network is configured to determine as the output values of the first neural network a second plurality of line spectral frequencies of the audio output signal, each of one or more of the second plurality of line spectral frequencies being associated with a frequency that is greater than any frequency associated with any of the first plurality of line spectral frequencies; the second neural network is configured to receive a plurality of samples of the excitation signal of the audio input signal and / or the audio input signal and / or a shaped version of the audio input signal as input values of the second neural network, and to determine the plurality of extrapolated excitation signal samples as output values of the second neural network; 2. The method of claim 1, wherein the input values of the second neural network are a first plurality of time-domain signal samples of the excitation signal of the audio input signal and / or the audio input signal and / or a shaped version of the audio input signal, and wherein the second neural network determines the output values of the second neural network such that the plurality of extrapolated excitation signal samples are a second plurality of time-domain signal samples of an extended time-domain excitation signal that has been bandwidth extended relative to the excitation signal of the audio input signal.
17. 17. Computer program for carrying out the method according to claim 16, when the computer program is running on a computer or signal processor.
Citation Information
Patent Citations
JPP2956548B