Perceptual Loss Function for Machine Learning-Based Speech Encoding and Decoding

By employing a neural network with a perception-based loss function and psychoacoustic model, the audio encoding and decoding process is enhanced to achieve superior perceptual quality, addressing the limitations of existing audio codecs.

JP7690545B2Active Publication Date: 2025-06-10DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023194046
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-04
Filing Date
2023-11-15
Publication Date
2025-06-10
Estimated Expiration
2039-04-10

AI Technical Summary

Technical Problem

Existing audio codecs struggle to efficiently encode and decode audio signals while maintaining high audio quality, particularly in terms of perceptual quality that aligns with human auditory perception.

Method used

The use of a neural network implemented by a control system, which receives an input audio signal, generates an encoded audio signal, decodes the encoded signal, and trains the neural network using a perception-based loss function that incorporates a psychoacoustic model to minimize the average noise-to-mask ratio.

Benefits of technology

This approach significantly improves the perceptual quality of the output audio signal by aligning the encoding and decoding processes with human auditory perception, outperforming traditional methods based on mean squared error or other loss functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690545000005
    Figure 0007690545000005
  • Figure 0007690545000006
    Figure 0007690545000006
  • Figure 0007690545000007
    Figure 0007690545000007
Patent Text Reader

Abstract

To provide computer-implemented methods for training a neural network, as well as for implementing audio encoders and decoders via trained neural network.SOLUTION: The neural network may receive an input audio signal, generate an encoded audio signal, and decode the encoded audio signal. A loss function generating module may receive the decoded audio signal and a ground truth audio signal, and may generate a loss function value corresponding to the decoded audio signal. Generating the loss function value may involve applying a psychoacoustic model. The neural network may be trained based on the loss function value. The training may involve updating at least one weight of the neural network.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the processing of audio signals. In particular, the present disclosure relates to the encoding and decoding of audio data.

Background Art

[0002] An audio codec is a device or computer program that can encode and / or decode digital audio data when a specific audio file or streaming media audio format is provided. The main purpose of an audio codec is usually to represent an audio signal with the minimum number of bits while maintaining a reasonable level of audio quality with that number of bits. Such audio data compression can reduce both the storage space required for audio data and the bandwidth required for transmitting audio data.

Summary of the Invention

[0003] Various audio processing methods are disclosed in this specification. Some such methods include receiving an input audio signal by a neural network implemented by a control system including one or more processors and one or more non-transitory storage media. Such methods may include generating an encoded audio signal by the neural network and based on the input audio signal. Some such methods may include decoding the encoded audio signal by the control system to generate a decoded audio signal, and receiving the decoded audio signal and a ground truth audio signal by a loss function generation module implemented by the control system. Such methods may include generating a loss function value corresponding to the decoded audio signal by the loss function generation module. The step of generating a loss function value may include applying a psychoacoustic model. Such methods may include training the neural network based on the loss function value. The step of training may include updating at least one weight of the neural network.

[0004] According to some implementations, the step of training the neural network may include backpropagation based on the loss function value. In some examples, the neural network may include an autoencoder. The step of training the neural network may include changing a physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network.

[0005] In some implementations, a first portion of the neural network may generate the encoded audio signal, and a second portion of the neural network may decode the encoded audio signal. In some such implementations, the first portion of the neural network may include an input neuron layer and a plurality of hidden neuron layers. The input neuron layer may include more neurons than the final hidden neuron layer in some examples. At least some neurons of the first portion of the neural network may be configured by a rectified linear unit (ReLU) activation function. In some examples, at least some neurons in the hidden layer of the second portion of the neural network may be configured by a ReLU activation function, and at least some neurons in the output layer of the second portion may be configured by a sigmoid activation function.

[0006] According to some examples, the psychoacoustic model may be at least partially based on one or more psychoacoustic masking thresholds. In some implementations, the psychoacoustic model may include modeling an outer ear transfer function, grouping into critical bands, frequency domain masking (including but not limited to level-dependent spreading), modeling frequency-dependent hearing thresholds, and / or calculating a noise-to-mask ratio. In some examples, the loss function may include calculating an average noise-to-mask ratio, and the training step may include minimizing the average noise-to-mask ratio.

[0007] Several voice encoding methods and apparatuses are disclosed in the present specification. In some examples, the voice encoding method may include receiving a current input voice signal by a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement a voice encoder including a neural network trained according to any of the methods disclosed in the present specification. Such a model may include encoding, by the voice encoder, the current input voice signal into a compressed voice format, and outputting an encoded voice signal in the compressed voice format.

[0008] Several voice decoding methods and apparatuses are disclosed in the present specification. In some examples, the voice decoding method may include receiving a current input compressed voice signal by a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement a voice decoder including a neural network trained according to any of the methods disclosed in the present specification. Such a method may include decoding, by the voice decoder, the current input compressed voice signal, and outputting a decoded voice signal. Some such methods may include playing back the decoded voice signal by one or more transducers.

[0009] Some or all of the methods described in this specification may be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include, but are not limited to, memory devices as described in this specification, such as random access memory (RAM), read-only memory (ROM), and the like. Accordingly, various novel aspects of the subject matter described in this disclosure may be implemented on non-transitory media storing software. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable, for example, by one or more components of a control system as disclosed in this specification. The software may include, for example, instructions for performing one or more of the methods disclosed in this specification.

[0010] At least some aspects of this disclosure may be implemented via apparatus. For example, one or more devices may be configured to at least partially execute the methods disclosed in this specification. In some implementations, the apparatus may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic element, discrete gate or transistor logic, or discrete hardware components. Accordingly, in some implementations, the control system may include one or more processors and one or more non-transitory storage media operably coupled to the one or more processors.

[0011] According to some such examples, the device may include an interface system and a control system. The control system may be configured to implement, for example, one or more of the methods disclosed in the present specification. For example, the control system may be configured to implement an audio encoder. The audio encoder may include a neural network trained according to one or more of the methods disclosed in the present specification. The control system may be configured to receive a current input audio signal, encode the current input audio signal into a compressed audio format, and output an encoded audio signal in the compressed audio format (e.g., by the interface system).

[0012] Alternatively or additionally, the control system may be configured to implement an audio decoder. The audio decoder may include a neural network trained according to a process including receiving an input training audio signal by the neural network and by the interface system, and generating an encoded training audio signal based on the neural network and the input training audio signal. The process may include, by the control system, decoding the encoded training audio signal to generate a decoded training audio signal, and receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system. The process may include generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal. The step of generating the loss function value may include applying a psychoacoustic model. The process may include training the neural network based on the loss function value.

[0013] The audio encoder may be further configured to receive a current input audio signal, encode the current input audio signal into a compressed audio format, and output an encoded audio signal in the compressed audio format.

[0014] In some implementations, the disclosed system may include an audio decoder device. The audio decoder device may include an interface system and a control system, and the control system may include one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement an audio decoder.

[0015] The audio decoder may include a neural network trained according to a process including receiving an input training audio signal by the neural network and by the interface system, and generating an encoded training audio signal based on the neural network and the input training audio signal. The process may include decoding, by the control system, the encoded training audio signal to generate a decoded training audio signal, and receiving, by a loss function generation module implemented by the control system, the decoded training audio signal and a ground truth audio signal. The process may include generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal. The step of generating the loss function value may include applying a psychoacoustic model. The process may include training, based on the loss function value, the neural network.

[0016] The audio decoder may be further configured to receive a current input encoded audio signal in a compressed audio format, decode the current input encoded audio signal into an uncompressed audio format, and output a decoded audio signal in the uncompressed audio format. According to some implementations, the system may include one or more transducers configured to reproduce the decoded audio signal.

[0017] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. The relative dimensions in the following drawings may not be drawn to scale. Similar numbers and designations in the various drawings generally represent like elements. **Brief Description of the Drawings**

[0018]

Figure 1

[0019]

Figure 2

[0020]

Figure 3

[0021]

Figure 4-1

Figure 4-2

[0022]

Figure 5A

[0023]

Figure 5B

[0024]

Figure 5C

[0025]

Figure 6

[0026]

Figure 7A

[0027]

Figure 7B

[0028]

Figure 8

[0029]

Figure 9A

[0030]

Figure 9B

[0031]

Figure 10

[0032]

Figure 11

[0033]

Figure 12

[0034]

Figure 13

Mode for Carrying Out the Invention

[0035] The following description is directed to specific implementations for the purpose of explaining some novel aspects of the present disclosure and examples of contexts in which such novel aspects may be implemented. However, the teachings in this specification can be applied in various different ways. Further, the described embodiments may be implemented in various hardware, software, firmware, etc. For example, aspects of the present application may be realized, at least in part, by a device, a system including more than one device, a method, a computer program product, etc. Thus, aspects of the present application may take the form of an embodiment of hardware, an embodiment of software (including firmware, resident software, microcode, etc.), and / or an embodiment combining both software and hardware aspects. Such embodiments may be referred to herein as "circuits," "modules," or "engines." Some aspects of the present application may take the form of a computer program product embodied in one or more non-transitory media having computer-readable program code implemented thereon. Such non-transitory media may include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. Thus, the teachings of the present disclosure are not limited to the implementations illustrated and / or described in this specification, but rather have broad applicability.

[0036] The inventor has studied various machine learning methods related to audio data processing, including but not limited to audio data encoding and decoding. In particular, the inventor has studied various methods of training different types of neural networks using loss functions related to how humans begin to perceive sound. The effects of each of these loss functions were evaluated according to the audio data generated by neural network encoding. The audio data was evaluated according to subjective and objective criteria. In some examples, audio data processed by a neural network trained using a loss function based on mean squared error was used as a basis for evaluating audio data generated according to the methods disclosed herein. In some examples, the process of evaluation according to subjective criteria includes the steps of having a human listener evaluate the resulting audio data and obtaining the listener's feedback.

[0037] The technology disclosed herein is based on the above research. The present disclosure provides various examples of using perception-based loss functions for training neural networks for audio data encoding and / or decoding. In some examples, the perception-based loss function is based on a psychoacoustic model. The psychoacoustic model may be based, for example, at least in part on one or more psychoacoustic masking thresholds. In some implementations, the psychoacoustic model may include the steps of modeling the outer ear transfer function, grouping into critical bands, frequency domain masking (including but not limited to level-dependent spreading), modeling the frequency-dependent hearing threshold, and / or calculating the noise-to-mask ratio. In some implementations, the loss function may include the step of calculating the average noise-to-mask ratio. In some such examples, the training process may include the step of minimizing the average noise-to-mask ratio.

[0038] FIG. 1 is a block diagram showing an example of components of a device that may be configured to perform at least a portion of the methods disclosed in this specification. In some examples, device 105 may be or include a personal computer, a desktop computer, or other local device configured to provide audio processing. In some examples, device 105 may be or include a server. According to some examples, device 105 may be a client device configured to communicate with a server via a network interface. The components of device 105 may be implemented by hardware, by software stored on a non-transitory medium, by firmware, and / or by combinations thereof. The types and numbers of components shown in FIG. 1 and other drawings disclosed in this specification are shown by way of example only. Alternative implementations may include more, fewer, and / or different components.

[0039] In this example, device 105 includes interface system 110 and control system 115. Interface system 110 may include one or more network interfaces, one or more interfaces between control system 115 and the memory system, and / or one or more external interfaces (e.g., one or more Universal Serial Bus (USB) interfaces). In some implementations, interface system 110 may include a user interface system. The user interface system may be configured to receive input from a user. In some implementations, the user interface system may be configured to provide feedback to the user. For example, the user interface system may include one or more displays corresponding to a touch and / or gesture detection system. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include devices that provide tactile feedback, such as motors, vibrators, etc. Control system 115 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic elements, discrete gates or transistor logic, and / or discrete hardware components.

[0040] In some examples, device 105 may be implemented in a single device. However, in some implementations, device 105 may be implemented in more than one device. In some such implementations, the functions of control system 115 may be included in more than one device. In some examples, device 105 may be a component of another device.

[0041] Figure 2 shows a block implementing the processing of machine learning according to a perception-based loss function, according to one example. In this example, the input audio signal 205 is provided to the machine learning module 210. The input audio signal 205 may, in some examples, be suitable for human conversations. However, in other examples, the input audio signal 205 may be suitable for other sounds such as music.

[0042] According to some examples, the elements of the system 200 may include, but are not limited to, the machine learning module 210 and may be implemented by one or more control systems such as the control system 115. The machine learning module 210 may receive the input audio signal 205, for example, by an interface system such as the interface system 110. In some examples, the machine learning module 210 may be configured to implement one or more neural networks such as the neural networks disclosed in this specification. However, in other implementations, the machine learning module 210 may be configured to implement one or more other types of machine learning such as Non-negative Matrix Factorization, Robust Principal Component Analysis, Sparse Coding, Probabilistic Latent Component Analysis, etc.

[0043] In the example shown in Figure 2, the machine learning module 210 provides the output audio signal 215 to the loss function generation module 220. The loss function generation module 225 and the optional ground truth module 220 may be implemented, for example, by a control system such as the control system 115. In some examples, the loss function generation module 225, the machine learning module 210, and the optional ground truth module 220 may be implemented by the same device. On the other hand, in other examples, the loss function generation module 225, the optional ground truth module 220, and the machine learning module 210 may be implemented by different devices.

[0044] According to this example, the loss function generation module 225 receives the input audio signal 205 and uses the input audio signal 205 as the "ground truth" for error determination. However, in some alternative implementations, the loss function generation module 225 may receive ground truth data from an optional ground truth module 220. Such implementations may include tasks such as conversation enhancement or conversation noise removal where the ground truth is not the original input audio signal. Regardless of whether the ground truth data is the input audio signal 205 or data received from an optional ground truth module, the loss function generation module 225 evaluates the output audio signal according to the loss function algorithm and the ground truth data, and provides the loss function value 230 to the machine learning module 210. In some such implementations, the machine learning module 210 includes an implementation of the optimization module 315 described later with reference to FIG. 3. In other examples, the system 200 includes an implementation of the optimization module 315 that is separate from but communicates with the machine learning module 210 and the loss function generation module 225. Various examples of loss functions are disclosed in the present specification. In this example, the loss function generation module 225 applies a perception-based loss function that may be based on a psychoacoustic model. According to this example, the machine learning process implemented by the machine learning module 210 (e.g., the process of training a neural network) is at least partially based on the loss function value 230.

[0045] Using a perception-based loss function, such as a loss function based on a psychoacoustic model, for machine learning (e.g., for training a neural network) can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of the output audio signal generated by a machine learning process using a traditional loss function based on mean squared error (MSE), L1-norm, etc. For example, a neural network trained for a given time duration with a loss function based on a psychoacoustic model can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of the output audio signal generated by a neural network having the same architecture trained with a loss function based on MSE for the same time duration. Further, a neural network trained to converge with a loss function based on a psychoacoustic model can typically generate an output audio signal with a higher perceptual quality than the output audio signal of a neural network having the same architecture trained to converge with a loss function based on MSE.

[0046] Some of the disclosed loss functions utilize psychoacoustic principles to determine which differences in the output audio signal 215 are audible to an average person and which are not. In some examples, a loss function based on a psychoacoustic model may utilize psychoacoustic phenomena such as temporal masking, frequency masking, loudness equalization curves, level-dependent masking, and / or the human hearing threshold. In some implementations, the perceptual loss function may operate in the time domain, and in other implementations, the perceptual loss function may operate in the frequency domain. In alternative implementations, the perceptual loss function may include both time domain and frequency domain operation. In some examples, the loss function may calculate the loss function using one input frame, and in other examples, the loss function may calculate the loss function using multiple input frames.

[0047] Figure 3 shows an example of the training process of a neural network according to several implementations disclosed in the present specification. Similar to other figures provided in the present specification, the number and types of elements are merely examples. According to some examples, the elements of system 301 may be implemented by one or more control systems such as control system 115. In the example shown in Figure 3, neural network 300 is an autoencoder. The technology for designing autoencoders is described in chapter 14 of Goodfellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning (MIT Press, 2016), which is incorporated herein by reference.

[0048] Neural network 300 includes a layer of nodes, also referred to herein as "neurons". Each neuron has a real-valued activation function. Its output, when given an input or a set of inputs, is generally referred to as an "activation" that defines the output of the neuron. According to some examples, the neurons of neural network 300 may utilize a sigmoid activation function, an ELU activation function, and / or a tanh activation function. Alternatively or additionally, the neurons of neural network 300 may utilize a rectified linear unit (ReLU) activation function.

[0049] Each connection between neurons (also called a "synapse") has a modifiable real-valued weight. A neuron can be an input neuron (receiving data from outside the network), an output neuron, or a hidden neuron that modifies data on the way from an input neuron to an output neuron. In the example shown in FIG. 3, the neurons in neuron layer 1 are input neurons, the neurons in neuron layer 7 are output neurons, and the neurons in neuron layers 2-6 are hidden neurons. Five hidden neuron layers are shown in FIG. 3, but some implementations may include more or fewer hidden layers. Some implementations of neural network 300 may include more or fewer hidden layers, for example, 10 or more hidden layers. For example, some implementations may include 10, 20, 30, 40, 50, 60, 70, 80, 90, or more hidden layers.

[0050] Here, the first part of neural network 300 (encoder part 305) is configured to generate an encoded audio signal, and the second part of neural network 300 (decoder part 310) is configured to decode the encoded audio signal. In this example, the encoded audio signal is a compressed audio signal, and the decoded audio signal is an uncompressed (decompressed) audio signal. Thus, the input audio signal 205 is compressed by the encoder part 305, as suggested by the reduced size of the blocks used to illustrate neuron layers 1-4. In some examples, the input neuron layer may include more neurons than at least one of the hidden neuron layers of the encoder part 305. However, in an alternative implementation, neuron layers 1-4 may all have the same number of neurons, or substantially the same number of neurons.

[0051] Therefore, the compressed audio signal provided by the encoder portion 305 is then decoded by the neuron layer of the decoder portion 310 to constitute the output signal 215, and the output signal 215 is an estimate of the input audio signal 205. A perceptual loss function, such as a psychoacoustics-based loss function, may be used to determine the update of the parameters of the neural network 300 during the training phase. These parameters can later be used to decode (e.g., decompress) any encoded (e.g., compressed) audio signal using the weights determined by the parameters received from the training algorithm. In other words, encoding and decoding may be performed separately from the training process after satisfactory weights of the neural network 300 have been determined.

[0052] According to this example, the loss function generation module 225 receives at least a part of the audio signal 205 and uses this as the ground truth data. Here, the loss function generation module 225 evaluates the output audio signal according to the loss function algorithm and the ground truth data, and provides the loss function value 230 to the optimization module 315. In this example, the optimization module 315 is initialized by information regarding the neural network and the loss function used by the loss function generation module 225. According to this example, the optimization module 315 uses this information together with the loss value received by the optimization module 315 from the loss function generation module 225 to calculate the gradient of the loss function with respect to the weights of the neural network. Once this gradient is known, the optimization module 315 uses an optimization algorithm to generate an update 320 for the weights of the neural network. According to some implementations, the optimization module 315 may utilize an optimization algorithm such as the Stochastic Gradient Descent or Adam optimization algorithm. The Adam optimization algorithm is disclosed in D. P. Kingma and J. L. Ba, “Adam: a Method for Stochastic Optimization,” in Proceedings of the International Conference on Learning Representations (ICLR), 2015, pp. 1-15, which is incorporated herein by reference. In the example shown in FIG. 3, the optimization module 315 is configured to provide an update 320 for the neural network 300. In this example, the loss function generation module 225 applies a perception-based loss function that may be based on a psychoacoustic model. According to this example, the process of training the neural network 300 is at least partially based on backpropagation. This backpropagation is shown in FIG. 3 by the dotted arrows between the neuron layers. Backpropagation (also known as "backpropagation") is a method used in neural networks to calculate the error contribution of each neuron after a batch of data has been processed.Backpropagation technology, sometimes also called backpropagation of error, because errors (errors) can be propagated backward through neural network layers calculated in the output.

[0053] Neural network 300 may be implemented by a control system such as control system 115 described above with reference to FIG. 1. Accordingly, the step of training neural network 300 may include changing the physical state of a non-transitory storage medium location corresponding to the weights in neural network 300. The storage medium location may be a portion of one or more storage media accessible by the control system or a portion thereof. Weights as described above correspond to connections between neurons. The step of training neural network 300 may include changing the physical state of a non-transitory storage medium location corresponding to the value of the activation function of the neurons.

[0054] Figures 4-1(A)-(C) show alternative examples of neural networks suitable for implementing some of the methods disclosed in this specification. According to these examples, input neurons and hidden neurons utilize the rectified linear unit (ReLU) activation function, and output neurons utilize the sigmoid activation function. However, alternative implementations of neural network 300 may include other activation functions and / or other combinations of activation functions including, but not limited to, ELU (Exponential Linear Unit) and / or tanh activation functions.

[0055] According to these examples, the input voice data is 256-bit voice data. In the example shown in FIG. 4-1(A), the encoder portion 305 compresses the input voice data into 32-bit voice data, providing a maximum of 8x compression (reduction). According to the example shown in FIG. 4-1(B), the encoder portion 305 compresses the input voice data into 16-bit voice data, providing a maximum of 16x compression (reduction). The neural network 300 shown in FIG. 4-1(C) includes an encoder portion 305 that compresses the input voice data into 8-bit voice data, providing a maximum of 32x compression (reduction). The inventor conducted a hearing test based on the type of neural network shown in FIG. 4-1(B). Some of the results are described below.

[0056] FIG. 4-2(D) shows an example of a block of the encoder portion of an autoencoder according to an alternative example. The encoder portion 305 may be implemented by a control system such as the control system 115 described above with reference to FIG. 1. The encoder portion 305 may be implemented, for example, by one or more processors of a control system according to software stored on one or more non-transitory storage media. The number and type of elements shown in FIG. 4-2(D) are merely examples. Other implementations of the encoder portion 305 may include more, fewer, or different elements.

[0057] In this example, the encoder portion 305 includes three neuron layers. According to some examples, the neurons of the encoder portion 305 may utilize the ReLU activation function. However, according to some alternative examples, the neurons of the encoder portion 305 may utilize the sigmoid activation function and / or the tanh activation function. The neurons in neuron layers 1-3 maintain their N-dimensional state while processing the N-dimensional input data. Layer 450 is configured to receive the output of neuron layer 3 and apply a pooling algorithm. Pooling is a form of non-linear downsampling. According to this example, layer 450 is configured to divide the output of neuron layer 3 into a set of M non-overlapping parts or "sub-regions" and apply a max-pooling function that outputs the maximum value for each sub-region.

[0058] FIG. 5A is a flow diagram outlining blocks of a method for training a neural network for voice encoding and decoding, according to an example. Method 500 may be performed, in some examples, by the apparatus of FIG. 1 or by another type of apparatus. In some examples, the blocks of method 500 may be implemented by software stored on one or more non-transitory media. The blocks of method 500 are not necessarily executed in the order shown, any more than the other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described.

[0059] Here, block 505 includes the step of receiving an input audio signal by a neural network implemented by a control system including one or more processors and one or more non-transitory memory media. In some examples, the neural network may or may not include an autoencoder. According to some examples, block 505 may include the control system of FIG. 1 that receives the input audio signal via interface system 110. In some examples, block 505 may include neural network 300 that receives input audio signal 205 as described above with reference to FIGS. 2-4C. In some implementations, input audio signal 205 may include at least a portion of a conversation dataset such as a publicly available conversation dataset known as TIMIT. TIMIT is a dataset of conversations of phoneme and word utterances of American English speakers of different genders and dialects. TIMIT was commissioned by DARPA (Defense Advanced Research Projects Agency). The corpus design of TIMIT is a joint research among Texas Instruments (TI), Massachusetts Institute of Technology (MIT), and SRI International. According to some examples, method 500 may include the step of converting input audio signal 205 from the time domain to the frequency domain, for example, by fast Fourier transform (FFT), discrete cosine transform (DCT), or short-time Fourier transform (STFT). In some implementations, min / max scaling may be applied to input audio signal 205 prior to block 510.

[0060] According to this example, block 510 includes the step of generating an encoded audio signal by a neural network and based on an input audio signal. The encoded audio signal may be or include a compressed audio signal. Block 510 may be performed by an encoder portion of a neural network, such as the encoder portion 305 of the neural network 300 described in this specification. However, in other examples, block 510 may include the step of generating an encoded audio signal by an encoder that is not part of the neural network. In some such examples, a control system implementing the neural network may also include an encoder that is not part of the neural network. For example, the neural network may include a decoding portion but not necessarily an encoding portion.

[0061] In this example, block 515 includes the step of decoding the encoded audio signal by a control system to generate a decoded audio signal. The decoded audio signal may be or include an uncompressed audio signal. In some implementations, block 515 may include the step of generating decoding conversion coefficients. Block 515 may be performed by a decoder portion of a neural network, such as the decoder portion 310 of the neural network 300 described in this specification. However, in other examples, block 510 may include the step of generating a decoded audio signal and / or decoding conversion coefficients by a decoder that is not part of the neural network. In some such examples, a control system implementing the neural network may also include a decoder that is not part of the neural network. For example, the neural network may include an encoding portion but not necessarily a decoding portion.

[0062] Accordingly, in some implementations, the first part of the neural network may be configured to generate an encoded audio signal, and the second part of the neural network may be configured to decode the encoded audio signal. In some such implementations, the first part of the neural network may include an input neuron layer and a plurality of hidden neuron layers. In some examples, the input neuron layer may include more neurons than at least one of the hidden neuron layers of the first part. However, in alternative implementations, the input neuron layer may have the same number of neurons as, or a substantially similar number of neurons to, the hidden neuron layers of the first part.

[0063] According to some examples, at least some neurons of the first part of the neural network may be configured by a rectified linear unit (ReLU) activation function. In some implementations, at least some neurons in the hidden layer of the second part of the neural network may be configured by a rectified linear unit (ReLU) activation function. According to some such implementations, at least some neurons in the output layer of the second part may be configured by a sigmoid activation function.

[0064] In some implementations, block 520 may include receiving, by a loss function generation module implemented by a control system, a decoded audio signal and / or decoded conversion coefficients, and a ground truth signal. The ground truth signal may include, for example, a ground truth audio signal and / or ground truth conversion coefficients. In some such examples, the ground truth signal may be received from a ground truth module such as the ground truth module 220 shown and described above with respect to FIG. 2. However, in some implementations, the ground truth signal may be, or may include, an input audio signal or a portion of the input audio signal. The loss function generation module may be, for example, an instance of the loss function generation module 225 disclosed herein.

[0065] According to some implementations, block 525 may include the step of generating, by a loss function generation module, a loss function value corresponding to the decoded audio signal and / or the decoded transform coefficients. In some such implementations, the step of generating the loss function value may include the step of applying a psychoacoustic model. In the example shown in FIG. 5A, block 530 includes the step of training a neural network based on the loss function value. The step of training may include the step of updating at least one weight in the neural network. In some such examples, an optimizer such as the optimization module 315 described above with reference to FIG. 3 may be initialized with information regarding the neural network and the loss function used by the loss function generation module 225. The optimization module 315 uses the information, together with the loss function value received by the optimization module 315 from the loss function generation module 225, to calculate the gradient of the loss function with respect to the weights of the neural network. After calculating the gradient, the optimization module 315 may use an optimization algorithm to generate updates to the weights of the neural network and provide these updates to the neural network. The step of training the neural network may include backpropagation based on the updates provided by the optimization module 315. Techniques for detecting and addressing overfitting during the process of training the neural network are described in chapter 5 and 7 of Goodfellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning (MIT Press, 2016), which is hereby incorporated by reference. The step of training the neural network may include the step of changing the physical state of at least one non-transitory storage medium location corresponding to at least one weight or at least one activation function value of the neural network.

[0066] The psychoacoustic model may vary according to a specific implementation. According to some examples, the psychoacoustic model may be at least partially based on one or more psychoacoustic masking thresholds. In some implementations, the step of applying the psychoacoustic model may include the step of modeling the outer ear transfer function, the step of grouping into critical bands, the step of frequency domain masking (including but not limited to level-dependent spreading), the modeling of frequency-dependent hearing thresholds, and / or the calculation of the noise-to-mask ratio. Some examples are described below with reference to FIGS. 6 to 10B.

[0067] In some implementations, the determination of the loss function generation module of the loss function may include the step of calculating a noise-to-mask ratio such as the average noise-to-masking ratio (NMR). The training process may include the step of minimizing the average NMR. Some examples are described below.

[0068] According to some examples, the step of training the neural network may continue until the loss function becomes relatively "flat". As a result, the difference between the current loss function value and the previous loss function value (such as the previous loss function value) becomes equal to or lower than the threshold. In the example shown in FIG. 5, the step of training the neural network may include repeating at least some of blocks 505 to 535 until the difference between the current loss function value and the previous loss function value is less than or equal to a predetermined value.

[0069] After the neural network has been trained, the neural network (or a portion thereof) may be used to process audio data, e.g., to encode or decode audio data. FIG. 5B is a flow diagram outlining blocks of a method of using a neural network trained for audio encoding, according to one example. Method 540 may be performed, in some examples, by the apparatus of FIG. 1 or by another type of apparatus. In some examples, the blocks of method 540 may be implemented by software stored on one or more non-transitory media. The blocks of method 540 are not necessarily performed in the order shown, as with the other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described.

[0070] In this example, block 545 includes receiving the currently input audio signal. In this example, block 545 includes receiving the currently input audio signal by a control system, the control system including one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors. Here, the control system is configured to implement an audio encoder including a neural network trained according to one or more of the methods disclosed herein.

[0071] In some examples, the training process includes receiving an input training audio signal by a neural network and via an interface system, generating an encoded training audio signal by the neural network based on the input training audio signal, decoding the encoded training audio signal by a control system to generate a decoded training audio signal, receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system, generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal, where the step of generating the loss function value includes applying a psychoacoustic model, and training the neural network based on the loss function value.

[0072] According to this implementation, block 550 includes encoding the current input audio signal into a compressed audio format by an audio encoder. Here, block 555 includes outputting the encoded audio signal in the compressed audio format.

[0073] FIG. 5C is a flowchart outlining blocks of a method of using a neural network trained for audio decoding according to an example. Method 560 may be performed in some examples by the device of FIG. 1 or by another type of device. In some examples, the blocks of method 560 may be implemented by software stored on one or more non-transitory media. The blocks of method 560 are not necessarily performed in the order shown, as with the other methods described herein. Further, such a method may include more or fewer blocks than those illustrated and / or described.

[0074] In this example, block 565 includes the step of receiving the currently input compressed audio signal. In some such examples, the currently input compressed audio signal may be generated according to method 540 or a similar method. In this example, block 565 includes the step of receiving the currently input compressed input audio signal by a control system, the control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. Here, the control system is configured to implement an audio decoder including a neural network trained according to one or more of the methods disclosed herein.

[0075] According to this implementation, block 570 includes the step of decoding the currently input compressed audio signal by the audio decoder. For example, block 570 may include the step of decompressing the currently input compressed audio signal. Here, block 575 includes the step of outputting the decoded audio signal. According to some examples, method 540 may include the step of playing back the decoded audio signal by one or more transducers.

[0076] As described above, the inventors have studied various ways of training different types of neural networks using loss functions related to how humans begin to perceive sound. The effect of each of these loss functions was evaluated according to the audio data generated by neural network encoding. In some examples, audio data processed by a neural network trained using a loss function based on mean squared error (MSE) was used as a basis for evaluating audio data generated according to the methods disclosed herein.

[0077] FIG. 6 is a block diagram showing a loss function generation module configured to generate a loss function based on the mean squared error. Here, the estimated magnitude of the audio signal generated by the neural network and the magnitude of the ground truth / true audio signal are both provided to the loss function generation module 225. The loss function generation module 225 generates a loss function value 230 based on the MSE value. The loss function value 230 may be provided to an optimization module configured to generate an update to the weights of the neural network for training.

[0078] The inventors evaluated several implementations of the loss function based at least in part on an acoustic response model of one or more parts of the human ear, which may also be referred to as an “ear model”. FIG. 7A is a graph of a function that approximates the standard acoustic response of the human external auditory canal.

[0079] FIG. 7B shows a loss function generation module configured to generate a loss function based on the standard acoustic response of the human external auditory canal. In this example, the function W is applied to both the audio signal generated by the neural network and the ground truth / true audio signal.

[0080] In some examples, the function W may be as follows.

Equation

[0081] Equation 1 is used in an implementation of the Perceptual Evaluation of Audio Quality (PEAQ) algorithm that models the acoustic response of the human external auditory canal. In Equation 1, f represents the frequency of the audio signal. In this example, the loss function generation module 225 generates a loss function value 230 based on the difference between the two resulting values. The loss function value 230 may be provided to an optimization module configured to generate an update to the weights of the neural network for training.

[0082] When compared with the audio signal generated by a neural network trained according to the loss function based on MSE, the audio signal generated by training the neural network using the loss function as shown in FIG. 7B provides only a slight improvement. For example, using an objective standard based on Perceptual Objective Listening Quality Analysis (POLQA), the audio data based on MSE reached a score of 3.41. On the other hand, the audio data generated by training a neural network using the loss function as shown in FIG. 7B achieved a score of 3.48.

[0083] In several experiments, the inventors tested the audio signal generated by a neural network trained according to the loss function based on band operation. FIG. 8 shows a loss function generation module configured to generate a loss function based on band operation. In this example, the loss function generation module 225 is configured to perform a band operation on the audio signal generated by the neural network and the ground truth / true audio signal, and to calculate the difference between the results.

[0084] In some implementations, the band operation is based on the "Zwicker" bands, which are critical bands defined according to chapter 6 (Critical Bands and Excitation) of Fastl, H., & Zwicker, E. (2007), Psychoacoustics: Facts and Models (3rd ed., Springer), incorporated herein by reference. In an alternative implementation, the band operation is based on the "Moore" bands, which are critical bands defined according to chapter 3 (Frequency Selectivity, Masking, and the Critical Band) of Moore, B. C.J. (2012), An Introduction to the Psychology of Hearing (Emerald Group Publishing), incorporated herein by reference. However, other examples may include other types of band operations known to those skilled in the art.

[0085] Based on the inventors' experiments, the inventors concluded that band operation alone is unlikely to provide satisfactory results. For example, using an objective criterion based on POLQA, the MSE-based speech data achieved a score of 3.41. On the other hand, the speech data generated by a neural network using band operation achieved only a score of 1.62.

[0086] In several experiments, the inventors tested audio signals generated by a neural network trained according to a loss function based at least in part on frequency masking. FIG. 9A shows the process involved in frequency masking according to several examples. In this example, the spreading function is calculated in the frequency domain. This spreading function may be, for example, a level and frequency dependent function that can be estimated from the input audio signal, for example from an input audio frame to be written. Next, convolution with the frequency spectrum of the input audio signal may be performed, which generates an Excitation pattern. The result of the convolution between the input audio data and the spreading function is an approximation of how the human auditory filter responds to the Excitation of the incoming audio. Thus, the process is a simulation of the human hearing mechanism. In some implementations, the audio data is grouped into frequency bins, and the convolution process includes, for each frequency bin, convolving the spreading function with the corresponding audio data of that frequency bin.

[0087] The Excitation pattern may be adjusted to generate a masking pattern. In some examples, the Excitation pattern may be adjusted downward, for example by 20 dB, to generate a masking pattern.

[0088] FIG. 9B shows an example of a spreading function. According to this example, the spreading function is a simplified asymmetric triangular function that can be pre-calculated for efficient implementation. In this simple example, the vertical axis represents decibels and the horizontal axis represents Bark sub-bands. According to one such example, the spreading function is calculated as follows.

Equation

[0089] In Equations 2 and 3, S l represents the slope of the portion of the spreading function in FIG. 9B to the left of the peak frequency, and S urepresents the slope of the portion of the spread function on the right side of the peak frequency. The unit of the slope is dB / Bark. In Equation 3, fc represents the center or peak frequency of the spread function, and L represents the level or amplitude of the audio data. In some examples, for the sake of simplifying the calculation of the spread function, L may be considered a constant. According to some such examples, L may be 70 dB.

[0090] In some such implementations, the Excitation pattern may be calculated as follows.

Equation

[0091] In Equation 4, E represents the Excitation function (also referred to as the Excitation pattern in this specification), SF represents the spread function, and BP represents the banded pattern of the frequency-binned audio data. In some implementations, the Excitation pattern may be adjusted to generate a masking pattern. In some examples, the Excitation pattern may be adjusted downward, for example, by 20 dB, 24 dB, 27 dB, etc., to generate a masking pattern.

[0092] FIG. 10 shows an example of an alternative implementation of the loss function generation module. The elements of the loss function generation module 225 may be implemented by a control system such as the control system 115 described above with reference to FIG. 1.

[0093] In this example, the reference audio signal x ref is referred to elsewhere in this specification as an instance of the ground truth signal and is provided to the fast Fourier transform (FFT) block 1005a of the loss function generation module 225. The test audio signal x is generated by a neural network such as one of those disclosed in this specification and is provided to the FFT block 1005b of the loss function generation module 225.

[0094] According to this example, the output of the FFT block 1005a is provided to the ear model block 1010a, and the output of the FFT block 1005a is provided to the ear model block 1010b. The ear model blocks 1010a and 1010b may be configured to provide functions based on the standard acoustic responses of one or more parts of a human ear canal. In one such example, the ear model blocks 1010a and 1010b may be configured to apply the functions described above in Equation 1.

[0095] According to this implementation, the outputs of the ear model blocks 1010a and 1010b are provided to the difference calculation block 1015. The difference calculation block 1015 is configured to calculate the difference between the output of the ear model block 1010a and the output of the ear model block 1010b. The output of the difference calculation block 1015 may be considered as an approximation of the noise in the test signal x.

[0096] In this example, the output of the ear model block 1010a is provided to the band processing block 1020a, and the output of the difference calculation block 1015 is provided to the band processing block 1020b. The band processing blocks 1020a and 1020b are configured to apply the same type of band processing, which may be one of the above-described band processing (e.g., Zwicker or Moore band processing). However, in an alternative implementation, the band processing blocks 1020a and 1020b may be configured to apply any suitable band processing known to those skilled in the art.

[0097] The output of the band processing block 1020a is provided to the frequency masking block 1025, and the frequency masking block 1025 is configured to apply frequency masking processing. The masking block 1025 may be configured to apply one or more of the frequency masking processes disclosed in the present specification, for example. As described above with reference to FIG. 9B, the use of the simplified frequency masking process can provide potential advantages. However, in an alternative implementation, the masking block 1025 may be configured to apply one or more other frequency masking processes known to those skilled in the art.

[0098] According to this example, the output of the masking block 1025 and the output of the band processing block 1020b are both provided to the noise-to-mask ratio (NMR) calculation block 1030. As described above, the output of the difference calculation block 1015 may be considered as an approximation of the noise in the test signal x. Therefore, the output of the band processing block 1020b may be considered as a frequency band-processed version of the noise in the test signal x. According to one example, the NMR calculation block 1030 may calculate the NMR as follows.

Equation

[0099] In Equation 5, BP noise represents the output of the band processing block 1020b, and MP represents the output of the masking block 1025. According to some examples, the NMR calculated by the NMR calculation block 1030 may be the average NMR over all the frequency bands output by the band processing blocks 1020a and 1020b. The NMR calculated by the NMR calculation block 1030 may be used as the loss function value 230, for example, to train a neural network as described above. For example, the loss function value 230 may be provided to an optimization module configured to generate updated weights of the neural network.

[0100] FIG. 11 shows an example of objective test results of some disclosed implementations. FIG. 11 shows a comparison between PESQ scores n of audio data generated by a neural network trained using loss functions based on mean squared error (MSE), power law, NMR-Zwicker (NMR based on a band processing such as Zwicker band processing but with a band slightly narrower than that defined by Zwicker), and NMR-Moore (NMR based on Moore band processing). These results are based on the output of the neural network described above with reference to FIG. 4-1(B), indicating that both the NMR-Zwicker and NMR-Moore results are somewhat better than the MSE and power law results.

[0101] FIG. 12 shows an example of subjective test results of audio data corresponding to a male speaker generated by a neural network trained using various types of loss functions. In this example, the subjective test results are MUSHRA (MUltiple Stimulus test with Hidden Reference and Anchor) evaluations. MUSHRA is described in ITU-R BS.1534 and is a well-known method for conducting codec listening tests to evaluate the perceptual quality of the output from lossy audio compression algorithms. The MUSHRA method has the advantage of presenting multiple stimuli simultaneously. As a result, the subjects can directly perform any comparison between them. The time required to execute the test using the MUSHRA method can be significantly shortened compared to other methods. This is partly true because the results from all codecs are presented simultaneously for the same samples. As a result, the paired t-test or repeated measures of variance can be used for statistical analysis. The numerical values along the x-axis in FIG. 12 are the identification numbers of different audio files.

[0102] More specifically, FIG. 12 shows a comparison between the audio data generated by the same neural network trained using a loss function based on MSE, a loss function based on the power law, a loss function based on NMR-Zwicker, and a loss function based on NMR-Moore, the audio data generated by applying a 3.5 kHz low-pass filter (one of the standard "anchors" of the MUSHRA method), and the reference audio data, in terms of the MUSHRA evaluation. In this example, the MUSHRA evaluation was obtained from 11 different listeners. As shown in FIG. 12, the average MUSHRA evaluation of the audio data generated by the neural network trained using the loss function based on NMR-Moore was significantly higher than any of the others. The difference was approximately 30 MUSHRA points, representing a large effect that is rarely seen. The second-highest average MUSHRA evaluation was for the audio data generated by the neural network trained using the loss function based on NMR-Zwicker.

[0103] FIG. 13 shows an example of the subjective test results of the audio data corresponding to a female speaker generated by a neural network trained using the same types of loss functions shown in FIG. 12. As in FIG. 12, the numerical values along the x-axis in FIG. 13 are the identification numbers of different audio files. In this example, the highest average MUSHRA evaluation was also assigned to the audio data generated by the neural network after training using the loss function based on NMR. The perceived difference between the NMR-Moore and NMR-Zwicker audio data and the other audio data was not as clear in this example as the perceived difference shown in FIG. 12, but still, the results shown in FIG. 13 indicate a significant improvement.

[0104] The general principles defined in this specification may be applied to other implementations without departing from the scope of the present disclosure. Accordingly, the claims are not intended to limit the implementations shown in this specification, but are to be regarded as covering the broadest scope consistent with the present disclosure, the principles disclosed in this specification, and the novel features.

[0105] Various aspects of the present invention may be apparent from the enumerated example embodiments (EEE) listed below. (EEE1) A method for voice processing implemented by a computer, comprising: receiving an input voice signal by a neural network implemented by a control system including one or more processors and one or more non-transitory storage media; generating an encoded voice signal based on the neural network and the input voice signal; decoding the encoded voice signal by the control system to generate a decoded training voice signal; receiving the decoded voice signal and a ground truth voice signal by a loss function generation module implemented by the control system; generating, by the loss function generation module, a loss function value corresponding to the decoded voice signal, the step of generating the loss function value including applying a psychoacoustic model; training the neural network based on the loss function value, the step of training including updating at least one weight of the neural network. (EEE2) The method according to EEE1, wherein the neural network includes backpropagation based on the loss function value. (EEE3) The method according to EEE1 or EEE2, wherein the neural network includes an autoencoder. (EEE4) The method according to any one of EEE1 to EEE3, wherein the step of training the neural network includes changing a physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network. (EEE5) The first part of the neural network generates the encoded audio signal, and the second part of the neural network decodes the encoded audio signal, the method according to any one of EEE1 to 4. (EEE6) The first part of the neural network includes an input neuron layer and a plurality of hidden neuron layers, and the input neuron layer includes more neurons than the final hidden neuron layer, the method according to EEE5. (EEE7) At least some neurons of the first part of the neural network are constituted by a rectified linear unit (ReLU) activation function, the method according to EEE5. (EEE8) At least some neurons in the hidden layer of the second part of the neural network are constituted by a rectified linear unit (ReLU) activation function, and at least some neurons in the output layer of the second part are constituted by a sigmoid activation function, the method according to EEE5. (EEE9) The psychoacoustic model is at least partially based on one or more psychoacoustic mask thresholds, the method according to any one of EEE1 to 8. (EEE10) The psychoacoustic model is as follows: Modeling of the outer ear transfer function, Grouping into critical bands, Frequency domain masking including, but not limited to, level-dependent spreading, Modeling of frequency-dependent hearing thresholds, Or calculation of the noise-to-mask ratio, including one or more of the above, the method according to any one of EEE1 to 9. (EEE11) The loss function includes a step of calculating an average noise-to-mask ratio, and the step of training includes a step of minimizing the average noise-to-mask ratio, the method according to any one of EEE1 to 10. (EEE12) A voice encoding method, A control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, the step of receiving a current input audio signal, wherein the control system is configured to implement an audio encoder including a neural network trained according to any one of the methods described in EEE1 to 11, the step; Encoding, by the audio encoder, the current input audio signal into a compressed audio format; Outputting the encoded audio signal in the compressed audio format. A method including the steps. (EEE13) An audio decoding method, A control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, the step of receiving a current input compressed audio signal, wherein the control system is configured to implement an audio decoder including a neural network trained according to any one of the methods described in EEE1 to 11, the step; Decoding, by the audio decoder, the current input compressed audio signal; Outputting the decoded audio signal. A method including the steps. (EEE14) The method according to EEE13, further including the step of reproducing the decoded audio signal by one or more transducers. (EEE15) An apparatus, An interface system; A control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, wherein the control system is configured to implement the method according to any one of EEE1 to 14, the control system; An apparatus including. (EEE16) One or more non-transitory media storing software, wherein the software includes instructions for controlling one or more devices to execute the method according to any one of EEE1 to 14, the non-transitory media. (EEE17) An audio encoding device, An interface system and, A control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, the control system being configured to implement an audio encoder, the audio encoder including a neural network trained according to the method described in any one of EEE1 to 11, a control system, and, The control system is configured to receive a current input audio signal, encode the current input audio signal into a compressed audio format, and output an encoded audio signal in the compressed audio format, a device configured as such. (EEE18) An audio encoding device, including an interface system and, a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, the control system being configured to implement an audio encoder, the audio encoder including a neural network trained according to a process, the process including receiving an input training audio signal by the neural network and via the interface system, generating an encoded training audio signal by the neural network and based on the input training audio signal, decoding the encoded training audio signal by the control system to generate a decoded training audio signal, receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system, generating, by the loss function generation module, a loss function value corresponding to the decoded audio signal, the step of generating the loss function value including applying a psychoacoustic model, a step, The step of training the neural network based on the loss function value; comprising, wherein the voice encoder encodes a current input voice signal into a compressed voice format, and is further configured to output an encoded voice signal in the compressed voice format, a device. (EEE19) A system including a voice decoder device, an interface system, and a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, wherein the control system is configured to implement a voice decoder, the voice decoder includes a neural network trained according to a process, and the process includes the step of receiving an input training voice signal by the neural network and via the interface system, the step of generating an encoded training voice signal by the neural network based on the input training voice signal, the step of decoding the encoded training voice signal by the control system to generate a decoded training voice signal, the step of receiving the decoded training voice signal and a ground truth voice signal by a loss function generation module implemented by the control system, the step of generating, by the loss function generation module, a loss function value corresponding to the decoded training voice signal, wherein the step of generating the loss function value includes the step of applying a psychoacoustic model, the step of training the neural network based on the loss function value, comprising, wherein the voice decoder receives a current input encoded voice signal in a compressed voice format, decodes the current input encoded voice signal into an uncompressed voice format, A device further configured to output a decoded audio signal of the non-compressed audio format. (EEE20) The system according to EEE19, further comprising one or more transducers configured to reproduce the decoded audio signal.

Explanation of Signs

[0106] 205 Input audio signal 215 Output audio signal 225 Loss function generation module 230 Loss function value 305 Encoder part 310 Decoder part 315 Optimization module

Claims

1. A method for encoding an audio signal using a neural network, the method comprising: receiving an audio signal; encoding the audio signal by the neural network; wherein the neural network is configured based on a perception-based loss function.

2. The step of encoding the audio signal by the neural network comprises: generating an encoded audio signal by the neural network based on the audio signal, according to the method of Claim 1.

3. Before the encoding step, further comprising the step of converting the audio signal from a time-domain representation to a frequency-domain representation, according to the method of Claim 1.

4. The neural network is trained based on the perception-based loss function, and the perception-based loss function models the acoustic response of the human external auditory canal, according to the method of Claim 1.

5. The perception-based loss function is based on at least one of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-mask ratio, minimization of the average noise-to-mask ratio, modeling of the external ear transfer function, grouping into critical bands, frequency-domain masking, level-dependent spreading, and / or modeling of the frequency-dependent hearing threshold, according to the method of Claim 4.

6. The audio signal includes at least one of conversation, music, and / or a mixture of voices, according to the method of Claim 1.

7. The neural network includes an autoencoder, according to the method of Claim 1.

8. Further comprising the step of outputting the encoded audio signal in a compressed audio format, according to the method of Claim 1.

9. A method for decoding an audio signal using a neural network, the method comprising: receiving an encoded audio signal; decoding the encoded audio signal via the neural network to generate at least one of a decoded audio signal and / or at least one decoded conversion coefficient; wherein the neural network is configured based on a perception-based loss function.

10. The decoding step comprises The method according to claim 9, comprising the step of generating at least one of the decoded audio signal and / or the at least one decoded conversion coefficient based on the neural network and based on the encoded audio signal.

11. The method according to claim 9, wherein the neural network is trained based on the perception-based loss function, and the perception-based loss function models the acoustic response of the human external auditory canal.

12. The method according to claim 11, wherein the perception-based loss function is based on at least one of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-mask ratio, minimization of the average noise-to-mask ratio, modeling of the outer ear transfer function, grouping into critical bands, frequency domain masking, level-dependent spreading, and / or modeling of the frequency-dependent hearing threshold.

13. The method according to claim 9, wherein the decoded audio signal includes at least one of conversation, music, and / or a mixture of voices.

14. An audio decoder, a neural network, receiving an encoded audio signal, generating a decoded audio signal and / or at least one decoded conversion coefficient based on the received encoded audio signal, a neural network configured as such, wherein the neural network is configured based on a perception-based loss function, the audio decoder.

15. The neural network is a plurality of neuron layers, an input neuron layer, a plurality of hidden neuron layers, an output neuron layer, The audio decoder according to claim 14, comprising a plurality of neuron layers including.

16. Each neuron layer of the plurality of neuron layers is configured by at least one of a rectified linear unit (ReLU) activation function, a sigmoid activation function, a tanh activation function, and / or an ELU (Exponential Linear Unit) activation function. The audio decoder according to claim 15.

17. The audio decoder according to claim 14, wherein the neural network is trained based on the perception-based loss function, and the perception-based loss function models the acoustic response of the human external auditory canal.

18. The loss function based on the perception is the audio decoder according to claim 17, based on at least one of a psychoacoustic model, a psychoacoustic masking threshold, an average noise-to-mask ratio, minimization of the average noise-to-mask ratio, modeling of an outer ear transfer function, grouping into critical bands, frequency domain masking, level-dependent spreading, and / or modeling of a frequency-dependent hearing threshold.

19. The decoded audio signal includes at least one of conversation, music, and / or a mixture of voices, and the audio decoder according to claim 14.

20. The encoded audio signal is composed of a compressed audio format, and the audio decoder according to claim 14.

Citation Information

Patent Citations

  • Information processing device and information processing method

    WO2017168870A1

  • Apparatus and method for encoding an audio signal using a compensation value

    WO2018036972A1