Perceptual-based loss functions for audio encoding and decoding using machine learning
By training the neural network based on the loss function of the psychoacoustic model, the problem of poor audio quality during the compression and decoding of the audio codec is solved, and a higher perception effect is achieved.
Patent Information
- Application Number
- CN202210834906.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-04
- Filing Date
- 2019-04-10
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2039-04-10
AI Technical Summary
Existing audio codecs have difficulty maintaining audio quality during compression and decoding, resulting in poor perception.
The loss function based on psychoacoustic model is used to train the neural network, and the audio encoding and decoding is performed through the autoencoder. The psychoacoustic model is used to evaluate and optimize the differences in audio signals and improve the perceived quality.
Through loss function training based on psychoacoustic model, the perceived quality of the audio signal is significantly improved and the effect of the audio encoding and decoding process is enhanced.
Smart Images

Figure CN115410583B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 201980030729.4, application date April 10, 2019, and invention name “Perception-based loss function for audio encoding and decoding based on machine learning”. Technical Field
[0002] The present disclosure relates to audio signal processing, and more particularly to encoding and decoding audio data. Background Art
[0003] An audio codec is a device or computer program capable of encoding and / or decoding digital audio data, given a specific audio file or streaming audio format. The primary goal of an audio codec is typically to represent the audio signal using the minimum number of bits while maintaining audio quality appropriate for that number of bits. This audio data compression reduces both the storage space required for the audio data and the bandwidth required for audio data transmission. Summary of the Invention
[0004] Various audio processing methods are disclosed herein. Some such methods may be computer-implemented audio processing methods that include receiving an input audio signal via a neural network implemented via a control system, the control system including one or more processors and one or more non-transitory storage media. Such methods may include generating an encoded audio signal based on the input audio signal via the neural network. Some such methods may include decoding the encoded audio signal via the control system to produce a decoded audio signal, and receiving the decoded audio signal and a ground truth audio signal via a loss function generation module implemented via the control system. Such methods may include generating a loss function value corresponding to the decoded audio signal via the loss function generation module. Generating the loss function value may include applying a psychoacoustic model. Such methods may include training the neural network based on the loss function value. Training may include updating at least one weight of the neural network.
[0005] According to some implementations, training the neural network may include backpropagation based on the loss function value. In some examples, the neural network may include an autoencoder. Training the neural network may include changing the physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network.
[0006] In some implementations, a first portion of the neural network can generate an encoded audio signal, and a second portion of the neural network can decode the encoded audio signal. In some such implementations, the first portion of the neural network can include an input neuron layer and multiple hidden neuron layers. In some cases, the input neuron layer can include more neurons than the final hidden neuron layer. At least some of the neurons of the first portion of the neural network can be configured with a rectified linear unit (ReLU) activation function. In some examples, at least some of the neurons in the hidden layer of the second portion of the neural network can be configured with a ReLU activation function, and at least some of the neurons in the output layer of the second portion can be configured with a sigmoidal activation function.
[0007] According to some examples, the psychoacoustic model can be based at least in part on one or more psychoacoustic masking thresholds. In some implementations, the psychoacoustic model can include modeling the outer ear transfer function, grouping into critical frequency bands, frequency domain masking (including but not limited to level-dependent expansion), modeling frequency-dependent hearing thresholds, and / or calculating a noise-masking ratio. In some examples, the loss function can involve calculating an average noise-masking ratio, and training can involve minimizing the average noise-masking ratio.
[0008] Disclosed herein are audio encoding methods and apparatus. In some examples, the audio encoding method may include receiving a currently input audio signal via a control system comprising one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system is configured to implement an audio encoder comprising a neural network trained according to any of the methods disclosed herein. Such a method may include encoding the currently input audio signal in a compressed audio format via the audio encoder and outputting the encoded audio signal in the compressed audio format.
[0009] Disclosed herein are certain audio decoding methods and apparatus. In some examples, the audio decoding methods may include receiving a currently input compressed audio signal via a control system comprising one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system is configured to implement an audio decoder comprising a neural network trained according to any of the methods disclosed herein. Such methods may include decoding the currently input compressed audio signal via the audio decoder and outputting a decoded audio signal. Some such methods may include reproducing the decoded audio signal via one or more transducers.
[0010] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include storage devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, and the like. Thus, various innovative aspects of the subject matter described in this disclosure may be implemented in a non-transitory medium having software stored thereon. The software may, for example, include instructions for controlling at least one device to process audio data. The software may, for example, be executed by one or more components of a control system such as those disclosed herein. The software may, for example, include instructions for executing one or more of the methods disclosed herein.
[0011] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be configured to at least partially perform the methods disclosed herein. In some embodiments, an apparatus may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between a control system and a memory system, one or more interfaces between a control system and another device, and / or one or more external device interfaces. The control system may include at least one of a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. Therefore, in some implementations, the control system may include one or more processors and one or more non-transitory storage media operably coupled to the one or more processors.
[0012] According to some such examples, the apparatus may include an interface system and a control system. The control system may, for example, be configured to implement one or more of the methods disclosed herein. For example, the control system may be configured to implement an audio encoder. The audio encoder may include a neural network that has been trained according to one or more of the methods disclosed herein. The control system may be configured to receive a currently input audio signal, encode the currently input audio signal in a compressed audio format, and output (e.g., via the interface system) the encoded audio signal in the compressed audio format.
[0013] Alternatively or additionally, the control system can be configured to implement an audio decoder. The audio decoder can include a neural network that has been trained according to the following process, the process including receiving an input training audio signal through the neural network and through the interface system, and generating an encoded training audio signal through the neural network and based on the input training audio signal. The process can include decoding the encoded training audio signal via the control system to produce a decoded training audio signal, and receiving the decoded training audio signal and a true audio signal via a loss function generation module implemented by the control system. The process can include generating a loss function value corresponding to the decoded training audio signal via the loss function generation module. Generating the loss function value may include applying a psychoacoustic model. The process can include training the neural network based on the loss function value.
[0014] The audio encoder may be further configured to receive a currently input audio signal, encode the currently input audio signal in a compressed audio format, and output an encoded audio signal in the compressed audio format.
[0015] In some implementations, the disclosed system may include an audio decoding device. The audio decoding device may include an interface system and a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement an audio decoder.
[0016] The audio decoder may include a neural network that has been trained according to the following process, the process including receiving an input training audio signal through the neural network and via the interface system, and generating an encoded training audio signal through the neural network and based on the input training audio signal. The process may include decoding the encoded training audio signal via the control system to produce a decoded training audio signal, and receiving the decoded training audio signal and a true audio signal through a loss function generation module implemented by the control system. The process may include generating a loss function value corresponding to the decoded training audio signal through the loss function generation module. Generating the loss function value may include applying a psychoacoustic model. The process may include training the neural network based on the loss function value.
[0017] The audio decoder may be further configured to receive a currently input encoded audio signal in a compressed audio format, to decode the currently input encoded audio signal in a decompressed audio format, and to output a decoded audio signal in a decompressed audio format. According to some implementations, the system may include one or more transducers configured to reproduce the decoded audio signal.
[0018] Details of one or more implementations of the subject matter described in the specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Please note that the relative dimensions of the following figures may not be drawn to scale. Like reference numerals and symbols in the various figures generally indicate like elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a block diagram illustrating an example of components of an apparatus that may be configured to perform at least some of the methods disclosed herein.
[0020] Figure 2 A block diagram of a process for implementing machine learning based on a perception-based loss function is shown according to one example.
[0021] Figure 3 An example of a neural network training process according to some implementations disclosed herein is shown.
[0022] Figures 4A-4D Alternative examples of neural networks suitable for implementing some of the methods disclosed herein are shown.
[0023] Figure 5A is a flowchart outlining the blocks of a method for training a neural network for audio encoding and decoding according to one example.
[0024] Figure 5B is a flowchart outlining blocks of a method for audio encoding using a trained neural network according to one example.
[0025] Figure 5C is a flowchart outlining blocks of a method for audio decoding using a trained neural network according to one example.
[0026] Figure 6 is a block diagram illustrating a loss function generation module configured to generate a loss function based on a mean square error.
[0027] Figure 7A is a graph of a function that approximates the typical acoustic response of the human ear canal.
[0028] Figure 7B A loss function generation module is shown, which is configured to generate a loss function based on the typical acoustic response of the human ear canal.
[0029] Figure 8 A loss function generation module is shown, which is configured to generate a loss function based on a banding operation.
[0030] Figure 9A The processes involved in frequency masking according to some examples are shown.
[0031] Figure 9B An example of an extension function is shown.
[0032] Figure 10 An example of an alternative implementation of the loss function generation module is shown.
[0033] Figure 11 Examples of objective test results for some publicly available implementations are shown.
[0034] Figure 12 Shown are examples of subjective test results corresponding to audio data of a male speaker produced by neural networks trained using various loss functions.
[0035] Figure 13 Shows the Figure 12 Examples of subjective test results corresponding to audio data of a female speaker produced by a neural network trained with the same type of loss function shown in . DETAILED DESCRIPTION
[0036] The following description is directed to examples of certain implementations for the purpose of describing some innovative aspects of the present disclosure and the context in which these innovative aspects can be implemented. However, the teachings herein can be applied in a variety of different ways. Moreover, the described embodiments can be implemented in various hardware, software, firmware, etc. For example, various aspects of the present application can be at least partially embodied in an apparatus, a system including more than one device, a method, a computer program product, etc. Therefore, various aspects of the present application can take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), and / or embodiments that combine both software and hardware aspects. Such embodiments may be referred to herein as "circuits," "modules," or "engines." Some aspects of the present application can take the form of a computer program product embodied in one or more non-transitory media, with computer-readable program code embodied on the non-transitory media. Such non-transitory media can, for example, include a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any appropriate combination of the foregoing. Thus, the teachings of the present disclosure are not intended to be limited to the implementations shown in the figures and / or described herein, but rather have broad applicability.
[0037] The inventors have investigated various machine learning methods related to audio data processing, including but not limited to audio data encoding and decoding. In particular, the inventors have investigated various methods for training different types of neural networks using loss functions related to the way humans perceive sound. The effectiveness of each of these loss functions was evaluated based on audio data generated by encoding the neural network. The audio data was evaluated based on both objective and subjective criteria. In some examples, audio data processed by a neural network that had been trained using a loss function based on mean squared error was used as the basis for evaluating audio data generated according to the methods disclosed herein. In some cases, the process of evaluating using subjective criteria involved having human listeners evaluate the resulting audio data and obtaining feedback from the listeners.
[0038] The technology disclosed herein is based on the above research. The present disclosure provides various examples of using a perceptual-based loss function to train a neural network for encoding and / or decoding audio data. In some examples, the perceptual-based loss function is based on a psychoacoustic model. The psychoacoustic model can, for example, be based at least in part on one or more psychoacoustic masking thresholds. In some examples, the psychoacoustic model can include modeling the outer ear transfer function, grouping the audio data into critical bands, frequency domain masking (including but not limited to level-dependent expansion), modeling frequency-dependent hearing thresholds, and / or calculation of a noise masking ratio. In some embodiments, the loss function can involve calculating an average noise masking ratio. In some such examples, the training process can include minimizing the average noise masking ratio.
[0039] Figure 1 1 is a block diagram illustrating an example of components of an apparatus that can be configured to perform at least some of the methods disclosed herein. In some examples, apparatus 105 can be or include a personal computer, a desktop computer, or other local device configured to provide audio processing. In some examples, apparatus 105 can be or include a server. According to some examples, apparatus 105 can be a client device configured to communicate with a server via a network interface. Components of apparatus 105 can be implemented via hardware, via software stored on non-transitory media, via firmware, and / or a combination thereof. Figure 1 The types and quantities of components shown and other figures disclosed herein are shown only as examples. Alternative implementations may include more, fewer and / or different components.
[0040] In this example, device 105 includes an interface system 110 and a control system 115. Interface system 110 may include one or more network interfaces, one or more interfaces between control system 115 and a storage system, and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). In some implementations, interface system 110 may include a user interface system. The user interface system may be configured to receive input from a user. In some implementations, the user interface system may be configured to provide feedback to the user. For example, the user interface system may include one or more displays with corresponding touch and / or gesture detection systems. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include a device for providing tactile feedback, such as a motor, a vibrator, etc. Control system 115 may, for example, include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic devices, and / or discrete hardware components.
[0041] In some examples, apparatus 105 may be implemented in a single device. However, in some implementations, apparatus 105 may be implemented in more than one device. In some such implementations, the functionality of control system 115 may be included in more than one device. In some examples, apparatus 105 may be a component of another device.
[0042] Figure 2 A block diagram of a process for implementing machine learning based on a perception-based loss function according to one example is shown. In this example, an input audio signal 205 is provided to a machine learning module 210. In some examples, the input audio signal 205 may correspond to human speech. However, in other examples, the input audio signal 205 may correspond to other sounds, such as music.
[0043] According to some examples, elements of system 200, including but not limited to machine learning module 210, can be implemented via one or more control systems, such as control system 115. Machine learning module 210 can receive input audio signal 205, for example, via an interface system, such as interface system 110. In some cases, machine learning module 210 can be configured to implement one or more neural networks, such as those disclosed herein. However, in other implementations, machine learning module 210 can be configured to implement one or more other types of machine learning, such as non-negative matrix factorization, robust principal component analysis, sparse coding, probabilistic latent component analysis, etc.
[0044] exist Figure 2In the example shown, the machine learning module 210 provides the output audio signal 215 to the loss function generation module 220. The loss function generation module 225 and the optional truth module 220 can be implemented, for example, via a control system such as the control system 115. In some examples, the loss function generation module 225, the machine learning module 210, and the optional truth module 220 can be implemented via the same device, while in other examples, the loss function generation module 225, the optional truth module 220, and the machine learning module 210 can be implemented via different devices.
[0045] According to this example, the loss function generation module 225 receives the input audio signal 205 and uses the input audio signal 205 as the "ground truth" for error determination. However, in some alternative implementations, the loss function generation module 225 may receive ground truth data from the optional ground truth module 220. Such implementations may, for example, involve tasks such as speech enhancement or speech denoising, where the ground truth is not the original input audio signal. Regardless of whether the ground truth data is the input audio signal 205 or data received from the optional ground truth module, the loss function generation module 225 evaluates the output audio signal based on the loss function algorithm and the ground truth data and provides a loss function value 230 to the machine learning module 210. In some such implementations, the machine learning module 210 includes an implementation of the optimizer module 315, which is described below with reference to Figure 3 . In other examples, system 200 includes an implementation of optimizer module 315 separate from, but in communication with, machine learning module 210 and loss function generation module 225. Various examples of loss functions are disclosed herein. In this example, loss function generation module 225 applies a perceptually based loss function, which may be based on a psychoacoustic model. According to this example, the machine learning process (e.g., the process of training a neural network) implemented by machine learning module 210 is based in part on loss function value 230.
[0046] Employing a perceptually based loss function (such as a psychoacoustic model-based loss function) for machine learning (e.g., for training a neural network) can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of the output audio signal produced by a machine learning process using a traditional loss function based on mean squared error (MSE), L1-norm, etc. For example, a neural network trained for a given time length using a psychoacoustic model-based loss function can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of the output audio signal produced by a neural network with the same architecture trained for the same time length using an MSE-based loss function. Moreover, a neural network trained to convergence using a psychoacoustic model-based loss function will generally produce an output audio signal of higher perceptual quality than the output audio signal of a neural network with the same architecture trained to convergence using an MSE-based loss function.
[0047] Some disclosed loss functions utilize psychoacoustic principles to determine which differences in the output audio signal 215 are audible to an average person and which differences are inaudible to an average person. In some examples, the loss function based on the psychoacoustic model can employ psychoacoustic phenomena such as time masking, frequency masking, equal loudness curves, level-dependent masking, and / or human hearing thresholds. In some implementations, the perceptual loss function can operate in the time domain, while in other implementations, the perceptual loss function can operate in the frequency domain. In alternative implementations, the perceptual loss function can involve both time domain operations and frequency domain operations. In some examples, the loss function can use one frame of input to calculate the loss function, while in other examples, the loss function can use multiple input frames to calculate the loss function.
[0048] Figure 3 An example of a neural network training process according to some implementations disclosed herein is shown. Figure 1 Again, the number and type of elements are provided as examples only. According to some examples, the elements of system 301 may be implemented via one or more control systems such as control system 115. Figure 3 In the example shown, neural network 300 is an autoencoder. Techniques for designing autoencoders are described in Chapter 14 of Goodfellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning (MIT Press, 2016), which is incorporated herein by reference.
[0049] Neural network 300 includes a layer of nodes, also referred to herein as "neurons." Each neuron has a real-valued activation function, whose output is typically referred to as an "activation," which defines the output of the neuron given an input or set of inputs. According to some examples, neurons of neural network 300 may utilize a sigmoid activation function, an ELU activation function, and / or a hyperbolic tangent activation function. Alternatively or additionally, neurons of speech neural network 300 may utilize a rectified linear unit (ReLU) activation function.
[0050] Each connection between neurons (also called a "synapse") has a real-valued weight that can be modified. Neurons can be input neurons (receive data from outside the network), output neurons, or hidden neurons that modify the data in the path from input neurons to output neurons. Figure 3 In the example shown, the neurons in neuron layer 1 are input neurons, the neurons in neuron layer 7 are output neurons, and the neurons in neuron layers 2-6 are hidden neurons. Figure 3 Five hidden layers are shown in FIG, but some implementations may include more or fewer hidden layers. Some implementations of neural network 300 may include more or fewer hidden layers, such as 10 or more hidden layers. For example, some implementations may include 10, 20, 30, 40, 50, 60, 70, 80, 90, or more hidden layers.
[0051] Here, the first portion of the neural network 300 (the encoding portion 305) is configured to generate an encoded audio signal, while the second portion of the neural network 300 (the decoding portion 310) is configured to decode the encoded audio signal. In this example, the encoded audio signal is a compressed audio signal, while the decoded audio signal is a decompressed audio signal. Thus, the input audio signal 205 is compressed by the encoding portion 305, as suggested by the reduced size of the blocks used to illustrate neuron layers 1-4. In some examples, the input neuron layer can include more neurons than at least one of the hidden neuron layers of the encoding portion 305. However, in alternative embodiments, all of the neuron layers 1-4 can have the same number of neurons, or a substantially similar number of neurons.
[0052] Thus, the compressed audio signal provided by the encoding portion 305 is then decoded via the neuron layer of the decoding portion 310 to construct the output signal 215, which is an estimate of the input audio signal 205. A perceptual loss function, such as a psychoacoustic-based loss function, can then be used to determine updates to the parameters of the neural network 300 during the training phase. These parameters can then be used to decode (e.g., decompress) any audio signal that has been encoded (e.g., compressed) using weights determined by the parameters received from the training algorithm. In other words, after satisfactory weights have been determined for the neural network 300, encoding and decoding can be performed separately from the training process.
[0053] According to this example, the loss function generation module 225 receives at least a portion of the audio input signal 205 and uses it as ground truth data. Here, the loss function generation module 225 evaluates the output audio signal based on the loss function algorithm and the ground truth data and provides the loss function value 230 to the optimizer module 315. In this example, the optimizer module 315 is initialized with the loss function used by the loss function generation module 225 and information about the neural network. According to this example, the optimizer module 315 uses this information along with the loss value received by the optimizer module 315 from the loss function generation module 225 to calculate the gradient of the loss function with respect to the neural network weights. Once this gradient is known, the optimizer module 315 uses an optimization algorithm to generate updates 320 for the neural network weights. According to some implementations, the optimizer module 315 can employ an optimization algorithm such as stochastic gradient descent or the Adam optimization algorithm. The Adam optimization algorithm is disclosed in DP Kingma and JL Ba, "Adam: a Method for Stochastic Optimization", International Conference on Learning Representations (ICLR), 2015, pp. 1-15, which is incorporated herein by reference. Figure 3 In the example shown, the optimizer module 315 is configured to provide updates 320 to the neural network 300. In this example, the loss function generation module 225 applies a perceptually based loss function, which may be based on a psychoacoustic model. According to this example, the process of training the neural network 300 is based at least in part on backpropagation. This backpropagation is performed in Figure 3 In the example above, backpropagation is represented by dash-dot arrows between neuron layers. Backpropagation (also known as "backward propagation") is a method used in neural networks to calculate the error contribution of each neuron after processing a batch of data. The backward propagation technique is sometimes called backpropagation of error because the error can be calculated at the output and distributed back through the neural network layers.
[0054] The neural network 300 may be formed by, for example, Figure 1The neural network 300 may be implemented in a control system such as the control system 115 described above. Thus, training the neural network 300 may include changing the physical state of a non-transitory storage medium location corresponding to a weight in the neural network 300. The storage medium location may be part of one or more storage media accessible to the control system or a portion of the control system. As described above, the weights correspond to connections between neurons. Training the neural network 300 may also include changing the physical state of a non-transitory storage medium location corresponding to the value of the activation function of the neuron.
[0055] Figures 4A-4C Alternative examples of neural networks suitable for implementing some of the methods disclosed herein are shown. According to these examples, the input neurons and hidden neurons use the rectified linear unit (ReLU) activation function, while the output neurons use the sigmoid activation function. However, alternative implementations of the neural network 300 may include other activation functions and / or other combinations of activation functions, including but not limited to exponential linear unit (ELU) and / or hyperbolic tangent activation functions.
[0056] According to these examples, the input audio data is 256-dimensional audio data. Figure 4A In the example shown, the encoding portion 305 compresses the input audio data into 32-dimensional audio data, providing up to 8 times reduction. Figure 4B In the example shown in , the encoding section 305 compresses the input audio data into 16-dimensional audio data, providing up to 16 times reduction. Figure 4C The neural network 300 shown in FIG4B includes an encoding portion 305 that compresses the input audio data into 8-dimensional audio data, providing a reduction of up to a factor of 32. The inventors conducted listening tests based on a neural network of the type shown in FIG4B , some results of which are described below.
[0057] Figure 4D An example of a block of an encoding portion of an autoencoder according to an alternative example is shown. The encoding portion 305 may be, for example, a block such as that described above with reference to Figure 1 The coding portion 305 may be implemented, for example, by one or more processors of the control system according to software stored in one or more non-transitory storage media. Figure 4D The number and types of elements shown are examples only. Other implementations of encoding portion 305 may include more, fewer, or different elements.
[0058] In this example, the encoding portion 305 includes three layers of neurons. According to some examples, the neurons of the encoding portion 305 can adopt a ReLU activation function. However, according to some alternative examples, the neurons of the encoding portion 305 can adopt a sigmoid activation function and / or a hyperbolic tangent activation function. The neurons in the neuron layers 1-3 process N-dimensional input data while maintaining the N-dimensional state of the N-dimensional input data. Layer 450 is configured to receive the output of the neuron layer 3 and apply a pooling algorithm. Pooling is a form of nonlinear downsampling. According to this example, layer 450 is configured to apply a maximum pooling function that divides the output of the neuron layer 3 into a set of M non-overlapping partitions or "sub-regions", and for each such sub-region, outputs a maximum value.
[0059] Figure 5A is a flowchart outlining a block diagram of a method for training a neural network for audio encoding and decoding according to one example. In some cases, the method 500 may be performed by Figure 1 The method 500 may be performed by a device or other type of device. In some examples, the blocks of method 500 may be implemented via software stored on one or more non-transitory media. Like other methods described herein, the blocks of method 500 are not necessarily executed in the order indicated. Moreover, such a method may include more or fewer blocks than shown and / or described.
[0060] Here, block 505 involves receiving an input audio signal via a neural network implemented via a control system, the control system including one or more processors and one or more non-transitory storage media. In some examples, the neural network may be or may include an autoencoder. According to some examples, block 505 may involve Figure 1 The control system 115 receives the input audio signal via the interface system 110. In some examples, block 505 may involve the neural network 300 receiving the input audio signal 205, as described above with reference to Figures 2 to 4Cdescribed. In some implementations, the input audio signal 205 may include at least a portion of a speech dataset, such as a publicly available speech dataset known as TIMIT. TIMIT is a dataset of phonemically and lexically transcribed speech from American English speakers of different genders and dialects. TIMIT was commissioned by the U.S. Defense Advanced Research Projects Agency (DARPA). The corpus design for TIMIT was a joint effort between Texas Instruments (TI), the Massachusetts Institute of Technology (MIT), and SRI International. According to some examples, method 500 may include transforming the input audio signal 205 from the time domain to the frequency domain, for example, via a fast Fourier transform (FFT), a discrete cosine transform (DCT), or a short-time Fourier transform (STFT). In some implementations, min / max scaling may be applied to the input audio signal 205 prior to block 510.
[0061] According to this example, block 510 involves generating an encoded audio signal based on an input audio signal through a neural network. The encoded audio signal may be or may include a compressed audio signal. Block 510 may be performed, for example, by an encoding portion of a neural network (such as encoding portion 305 of neural network 300 described herein). However, in other examples, block 510 may involve generating the encoded audio signal via an encoder that is not part of the neural network. In some such examples, a control system implementing the neural network may also include an encoder that is not part of the neural network. For example, the neural network may include a decoding portion but not an encoding portion.
[0062] In this example, block 515 involves decoding the encoded audio signal via a control system to produce a decoded audio signal. The decoded audio signal may be or may include a decompressed audio signal. In some implementations, block 515 may involve generating decoded transform coefficients. Block 515 may be performed, for example, by a decoding portion of a neural network (such as decoding portion 310 of neural network 300 described herein). However, in other examples, block 510 may involve generating a decoded audio signal and / or decoded transform coefficients via a decoder that is not part of the neural network. In some such examples, the control system that implements the neural network may also include a decoder that is not part of the neural network. For example, the neural network may include an encoding portion but not a decoding portion.
[0063] Thus, in some implementations, the first portion of the neural network can be configured to generate an encoded audio signal, and the second portion of the neural network can be configured to decode the encoded audio signal. In some such implementations, the first portion of the neural network can include an input neuron layer and multiple hidden neuron layers. In some examples, the input neuron layer can include more neurons than at least one of the hidden neuron layers of the first portion. However, in alternative implementations, the input neuron layer can have the same number of neurons as the hidden neuron layer of the first portion, or a substantially similar number of neurons.
[0064] According to some examples, at least some neurons of the first portion of the neural network can be configured with a rectified linear unit (ReLU) activation function. In some implementations, at least some neurons in a hidden layer of the second portion of the neural network can be configured with a rectified linear unit (ReLU) activation function. According to some such implementations, at least some neurons in an output layer of the second portion can be configured with a sigmoid activation function.
[0065] In some implementations, block 520 may include receiving a decoded audio signal and / or decoded transform coefficients, and a true value signal via a loss function generation module implemented via a control system. The true value signal may, for example, include a true valued audio signal and / or a true valued transform coefficient. In some such examples, the loss function generation module may be implemented via a control system. Figure 2 The truth value module 220 shown in FIG and described above receives the truth value signal. However, in some implementations, the truth value signal can be (or can include) the input audio signal or a portion of the input audio signal. The loss function generation module can, for example, be an instance of the loss function generation module 225 disclosed herein.
[0066] According to some implementations, block 525 may include generating, by a loss function generation module, a loss function value corresponding to the decoded audio signal and / or the decoded transform coefficients. In some such implementations, generating the loss function value may involve applying a psychoacoustic model. Figure 5A In the example shown, block 530 involves training the neural network based on the loss function value. Training can involve updating at least one weight in the neural network. In some such examples, such as those described above with reference to Figure 3An optimizer such as the described optimizer module 315 may have been initialized with information about the loss function(s) and the neural network used by the loss function generation module 225. The optimizer module 315 may be configured to use this information, along with the loss function values received by the optimizer module 315 from the loss function generation module 225, to calculate the gradients of the loss function with respect to the weights of the neural network. After calculating the gradients, the optimizer module 315 may use an optimization algorithm to generate updates to the weights of the neural network and provide these updates to the neural network. Training the neural network may involve backpropagation based on the updates provided by the optimizer module 315. Techniques for detecting and addressing overfitting during neural network training are described in Chapters 5 and 7 of Goodpellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning, (MIT Press, 2016), which is incorporated herein by reference. Training the neural network may involve changing the physical state of at least one non-transitory storage medium location corresponding to at least one weight or at least one activation function value of the neural network.
[0067] The psychoacoustic model may vary depending on the specific implementation. According to some examples, the psychoacoustic model may be based at least in part on one or more psychoacoustic masking thresholds. In some implementations, applying the psychoacoustic model may include modeling the outer ear transfer function, grouping into critical bands, frequency domain masking (including but not limited to level-dependent expansion), modeling frequency-dependent hearing thresholds, and / or calculating noise masking ratios. Figure 6-10 Describe some examples.
[0068] In some implementations, the loss function of the loss function generation module may involve calculating a noise-masking ratio (NMR), such as an average NMR. The training process may involve minimizing the average NMR. Some examples are described below.
[0069] According to some examples, training the neural network can continue until the loss function is relatively "flat," such that the difference between the current loss function value and the previous loss function value (e.g., the previous loss function value) is equal to or less than a threshold. In the example shown in FIG5 , training the neural network can include repeating at least some of blocks 505 through 535 until the difference between the current loss function value and the previous loss function value is less than or equal to a predetermined value.
[0070] After the neural network has been trained, the neural network (or a portion thereof) may be used to process audio data, for example, to encode or decode audio data. Figure 5Bis a flowchart outlining a block diagram of a method for audio encoding using a trained neural network according to one example. In some cases, method 540 may be performed by Figure 1 The method 540 may be executed by a device or other type of device. In some examples, the blocks of method 540 may be implemented via software stored on one or more non-transitory media. Like other methods described herein, the blocks of method 540 are not necessarily executed in the order indicated. Moreover, such methods may include more or fewer blocks than shown and / or described.
[0071] In this example, block 545 involves receiving a currently input audio signal. In this example, block 545 involves receiving the currently input audio signal via a control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors. Here, the control system is configured to implement an audio encoder comprising a neural network that has been trained according to one or more of the methods disclosed herein.
[0072] In some examples, the training process may include: receiving an input training audio signal through a neural network and via an interface system; generating an encoded training audio signal through the neural network and based on the input training audio signal; decoding the encoded training audio signal via a control system to produce a decoded training audio signal; receiving the decoded training audio signal and a true audio signal through a loss function generation module implemented via the control system; generating a loss function value corresponding to the decoded training audio signal through the loss function generation module, wherein generating the loss function value includes applying a psychoacoustic model; and training the neural network based on the loss function value.
[0073] According to this implementation, block 550 involves encoding the currently input audio signal in a compressed audio format via an audio encoder.Here, block 555 involves outputting the encoded audio signal in a compressed audio format.
[0074] Figure 5C is a flowchart outlining a method of audio decoding using a trained neural network according to one example. In some cases, method 560 may be performed by Figure 1 The method 560 may be performed by a device or other type of device. In some examples, the blocks of method 560 may be implemented via software stored on one or more non-transitory media. Like other methods described herein, the blocks of method 560 are not necessarily executed in the order indicated. Moreover, such methods may include more or fewer blocks than shown and / or described.
[0075] In this example, block 565 involves receiving a currently input compressed audio signal. In some such examples, the currently input compressed audio signal may have been generated according to method 540 or by a similar method. In this example, block 565 involves receiving the currently input compressed audio signal via a control system comprising one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. Here, the control system is configured to implement an audio decoder comprising a neural network that has been trained according to one or more of the methods disclosed herein.
[0076] Depending on the implementation, block 570 involves decoding the currently input compressed audio signal via an audio decoder. For example, block 570 may include decompressing the currently input compressed audio signal. Here, block 575 involves outputting the decoded audio signal. According to some examples, method 540 may include reproducing the decoded audio signal via one or more transducers.
[0077] As described above, the inventors have investigated various methods for training different types of neural networks using loss functions related to the way humans perceive sound. The effectiveness of each loss function was evaluated based on audio data generated by encoding the neural network. In some examples, audio data processed by a neural network that had been trained using a loss function based on mean squared error (MSE) was used as the basis for evaluating audio data generated according to the methods disclosed herein.
[0078] Figure 6 2 is a block diagram illustrating a loss function generation module configured to generate a loss function based on a mean square error. Here, both the estimated amplitude of the audio signal generated by the neural network and the amplitude of the true / real audio signal are provided to the loss function generation module 225. The loss function generation module 225 generates a loss function value 230 based on the MSE value. The loss function value 230 can be provided to an optimizer module configured to generate updates to the weights of the neural network for training.
[0079] The inventors have evaluated several implementations of loss functions based at least in part on a model of the acoustic response of one or more parts of the human ear (which may also be referred to as an "ear model"). Figure 7A is a graph of a function that approximates the typical acoustic response of the human ear canal.
[0080] Figure 7B A loss function generation module is shown, which is configured to generate a loss function based on the typical acoustic response of the human ear canal. In this example, the function W is applied to the audio signal generated by the neural network and the ground truth / real audio signal.
[0081] In some examples, the function W may be as follows:
[0082]
[0083] Equation 1 has been used to implement a perceptual evaluation of audio quality (PEAQ) algorithm for the purpose of modeling the acoustic response of the human ear canal. In Equation 1, f represents the frequency of the audio signal. In this example, the loss function generation module 225 generates a loss function value 230 based on the difference between the two result values. The loss function value 230 can be provided to the optimizer module, which is configured to generate updates to the weights of the neural network for training.
[0084] When compared with audio signals produced by a neural network trained according to an MSE-based loss function, by using a Figure 7B The audio signals produced by training the neural network with the loss function shown in provide only slight improvements. For example, using an objective criterion based on Perceptual Objective Listening Quality Analysis (POLQA), the score of the audio data based on MSE is 3.41, while using an objective criterion such as Figure 7B The loss function shown here produces a score of 3.48 on the audio data trained on the neural network.
[0085] In some experiments, the inventors tested audio signals generated by a neural network trained according to a loss function based on a band-wise operation. Figure 8 A loss function generation module is shown, which is configured to generate a loss function based on the banding operation. In this example, the loss function generation module 225 is configured to perform the banding operation on the audio signal generated by the neural network and the true value / real audio signal, and calculate the difference between the results.
[0086] In some embodiments, the banding operation is based on "Zwicker" bands, which are critical bands defined according to Chapter 6 (Critical Bands and Excitation) of Fastl, H. & Zwicker, E. (2007), Psychoacoustics: Facts and Models (3rd ed., Springer), which is incorporated herein by reference. In an alternative implementation, the banding operation is based on "Moore" bands, which are critical bands defined according to Chapter 3 (Frequency Selectivity, Masking, and Critical Bands) of Moore, BCJ (2012), Introduction to the Psychology of Hearing (Emerald Group Publishing), which is incorporated herein by reference. However, other examples may involve other types of banding operations known to those skilled in the art.
[0087] Based on their experiments, the inventors concluded that banding alone is unlikely to provide satisfactory results. For example, using an objective metric based on POLQA, audio data based on MSE achieved a score of 3.41, while in one example, audio data generated by training a neural network using banding only achieved a score of 1.62.
[0088] In some experiments, the inventors tested audio signals generated by a neural network trained according to a loss function based at least in part on frequency masking. Figure 9A The process involved in frequency masking according to some examples is shown. In this example, a spread function is calculated in the frequency domain. The spread function can, for example, be a function that depends on level and frequency, which can be estimated based on the input audio data (for example, based on each input audio frame). Convolution with the frequency spectrum of the input audio can then be performed, which produces an excitation pattern. The convolution result between the input audio data and the spread function is an approximation of how the human auditory filter reacts to the excitation of the incoming sound. Therefore, this process is a simulation of the human hearing mechanism. In some implementations, the audio data is grouped into frequency bins, and the convolution process includes convolving the spread function of each frequency bin with the corresponding audio data of the frequency bin.
[0089] The excitation pattern can be adjusted to generate the mask pattern. In some examples, the excitation pattern can be adjusted downward, for example, by 20 dB, to generate the mask pattern.
[0090] Figure 9B An example of an expansion function is shown. According to this example, the expansion function is a simplified asymmetric trigonometric function that can be pre-calculated for efficient implementation. In this simplified example, the vertical axis represents decibels and the horizontal axis represents the Bark subband. According to one such example, the expansion function is calculated as follows:
[0091] S l =27 (Formula 2)
[0092]
[0093] In Equations 2 and 3, S l express Figure 9B The slope of the portion of the spread function to the left of the peak frequency, and S u represents the slope of the portion of the spread spectrum function to the right of the peak frequency. The slope unit is dB / Bark. In Equation 3, f represents the center frequency or peak frequency of the spread function, and L represents the level or amplitude of the audio data. In some examples, to simplify the calculation of the spread function, L can be assumed to be a constant. According to some such examples, L can be 70dB.
[0094] In some such implementations, the excitation pattern may be calculated as follows:
[0095]
[0096] In Equation 4, E represents an excitation function (also referred to herein as an excitation pattern), SF represents a spreading function, and BP represents a band pattern of audio data divided into frequency bins. In some implementations, the excitation pattern can be adjusted to generate a masking pattern. In some examples, the excitation pattern can be adjusted downward, for example, by 20 dB, 24 dB, 27 dB, etc., to generate a masking pattern.
[0097] Figure 10 An example of an alternative implementation of the loss function generation module is shown. The elements of the loss function generation module 225 can be implemented by, for example, a reference Figure 1 This may be accomplished using a control system such as control system 115 described above.
[0098] In this example, the reference audio signal x ref Provided to the Fast Fourier Transform (FFT) block 1005a of the loss function generation module 225, the reference audio signal x ref is an example of a true value signal as referenced elsewhere herein. A test audio signal x generated by a neural network such as one of those disclosed herein is provided to the FFT block 1005b of the loss function generation module 225 .
[0099] According to this example, the output of FFT block 1005a is provided to ear model block 1010a, and the output of FFT block 1005b is provided to ear model block 101b. Ear model blocks 1010a and 1010b can, for example, be configured to apply a function based on the typical acoustic response of one or more parts of the human ear. In one such example, ear model blocks 1010a and 1010b can be configured to apply the function shown in Equation 1 above.
[0100] According to this implementation, the outputs of ear model blocks 1010a and 1010b are provided to a difference calculation block 1015, which is configured to calculate the difference between the output of ear model block 1010a and the output of ear model block 1010b. The output of difference calculation block 1015 can be considered an approximation of the noise in the test signal x.
[0101] In this example, the output of the ear model block 1010a is provided to the banding block 1020a, and the output of the difference calculation block 1015 is provided to the banding block 1020b. The banding blocks 1020a and 1020b are configured to apply the same type of banding process, which may be one of the banding processes disclosed above (e.g., the Zwicker or Moore banding processes). However, in alternative embodiments, the banding blocks 1020a and 1020b may be configured to apply any suitable banding process known to those skilled in the art.
[0102] The output of the band splitting block 1020a is provided to a frequency masking block 1025, which is configured to apply a frequency masking operation. The masking block 1025 can be configured to apply one or more of the frequency masking operations disclosed herein. Figure 9B As described above, using a simplified frequency masking process may provide potential advantages. However, in alternative implementations, masking block 1025 may be configured to apply one or more other frequency masking operations known to those skilled in the art.
[0103] According to this example, both the output of masking block 1025 and the output of banding block 1020b are provided to noise-masking ratio (NMR) calculation block 1030. As described above, the output of difference calculation block 1015 can be considered an approximation of the noise in test signal x. Therefore, the output of banding block 1020b can be considered a frequency band representation of the noise in test signal x. According to one example, NMR calculation block 1030 can calculate NMR as follows:
[0104]
[0105] In Equation 5, BP noise represents the output of banding block 1020b, and MP represents the output of masking block 1025. According to some examples, the NMR calculated by NMR calculation block 1030 may be the average NMR across all frequency bands output by banding blocks 1020a and 1020b. The NMR calculated by NMR calculation block 1030 may be used as loss function value 230 for training a neural network, such as the neural network described above. For example, loss function value 230 may be provided to an optimizer module configured to generate updated weights for the neural network.
[0106] Figure 11 Examples of objective test results for some disclosed implementations are shown. Figure 11A comparison is shown between PESQ scores for audio data produced by neural networks using methods based on MSE, power law, NMR-Zwicker (NMR based on a banding process like the Zwicker banding process, but with slightly narrower frequency bands than those defined by Zwicker), and NMR-Moore (NMR based on the Moore banding process). Figure 4B These results for the output of the neural network show that both the NMR-Zwicker and NMR-Moore results perform better than the MSE and power law results.
[0107] Figure 12 Examples of subjective test results corresponding to audio data from a male speaker, generated by neural networks trained using various loss functions, are shown. In this example, the subjective test results are ratings from the Multiple Stimulus Assessment with Hidden Reference and Anchor (MUSHRA). MUSHRA, described in ITU-R BS.1534, is a well-known method for conducting codec listening tests to evaluate the perceptual quality of the output of lossy audio compression algorithms. The advantage of the MUSHRA method is that many stimuli can be presented simultaneously, allowing subjects to directly compare them. The time required to perform tests using the MUSHRA method can be significantly reduced compared to other methods. This is partly because the results for all codecs are presented simultaneously on the same samples, allowing paired t-tests or analyses to be used for statistical analysis. Figure 12 The numbers along the x-axis are the identification numbers of the different audio files.
[0108] More specifically, Figure 12 Comparisons between audio data generated by the same neural network trained using an MSE-based loss function, a power law-based loss function, an NMR-Zwicker-based loss function, and an NMR-Moore-based loss function, audio data generated by applying a 3.5kHz low-pass filter (one of the standard "anchors" of the MUSHRA technique), and MUSHRA scores for reference audio data are shown. In this example, the MUSHRA scores were obtained from 11 different listeners. Figure 12 As shown, the average Mushra score for audio data generated by a neural network trained using the NMR-Moore-based loss function is significantly higher than any other score. The difference between the two is approximately 30 Mushra points, a rare large effect. The second-highest average Mushra score is for audio data generated by a neural network trained using the NMR-Zwicker-based loss function.
[0109] Figure 13 Shown by Figure 12 An example of subjective test results corresponding to audio data of a female speaker generated by a neural network trained with the same type of loss function as shown in FIG12 . Figure 13 The numbers on the x-axis are the identification numbers of the different audio files. In this example, the highest average MUSHRA score was again assigned to the audio data generated by the neural network trained using the NMR-based loss function. Although in this example, there was no perceptual difference between the NMR-Moore and NMR-Zwicker audio data and the other audio data, Figure 12 The perceptual differences shown in Figure 13 The results shown in still indicate a significant improvement.
[0110] The general principles defined herein may be applied to other implementations without departing from the scope of the present disclosure. Thus, the claims are not intended to be limited to the embodiments shown herein, but should be accorded the widest scope consistent with this disclosure, principles and novel features.
[0111] Various aspects of the present invention may be understood from the following enumerated exemplary embodiments (EEE):
[0112] 1. A computer-implemented audio processing method, comprising:
[0113] receiving, via a neural network implemented via a control system including one or more processors and one or more non-transitory storage media, an input audio signal;
[0114] generating an encoded audio signal based on the input audio signal through a neural network;
[0115] decoding, via the control system, the encoded audio signal to produce a decoded audio signal;
[0116] Receiving, by a loss function generation module implemented via the control system, a decoded audio signal and a true audio signal;
[0117] generating, by a loss function generating module, a loss function value corresponding to the decoded audio signal, wherein generating the loss function value comprises applying a psychoacoustic model; and
[0118] The neural network is trained based on the loss function value, wherein the training includes updating at least one weight of the neural network.
[0119] 2. The method of EEE 1, wherein training the neural network includes backpropagation based on the value of the loss function.
[0120] 3. The method of EEE 1 or EEE 2, wherein the neural network includes an autoencoder.
[0121] 4. The method of any of EEE 1-3, wherein training the neural network comprises changing a physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network.
[0122] 5. The method of any one of EEEs 1-4, wherein the first part of the neural network generates an encoded audio signal and the second part of the neural network decodes the encoded audio signal.
[0123] 6. The method of EEE 5, wherein the first part of the neural network comprises an input neuron layer and multiple hidden neuron layers, wherein the input neuron layer comprises more neurons than the final hidden neuron layer.
[0124] 7. The method of EEE 5, wherein at least some of the neurons of the first part of the neural network are configured with a rectified linear unit (ReLU) activation function.
[0125] 8. The method of EEE 5, wherein at least some neurons in the hidden layer of the second part of the neural network are configured with a rectified linear unit (ReLU) activation function, and at least some neurons in the output layer of the second part are configured with a sigmoid activation function.
[0126] 9. The method of any one of EEE 1 to 8, wherein the psychoacoustic model is based at least in part on one or more psychoacoustic masking thresholds.
[0127] 10. The method of any one of EEEs 1-9, wherein the psychoacoustic model involves one or more of: modeling outer ear transfer function; grouping into critical frequency bands; frequency domain masking, including but not limited to level-dependent spreading; modeling of frequency-dependent hearing thresholds; or calculating noise masking ratios.
[0128] 11. The method of any one of EEE 1-10, wherein the loss function comprises calculating an average noise-mask ratio, and wherein the training comprises minimizing the average noise-mask ratio.
[0129] 12. An audio encoding method, comprising:
[0130] receiving, by a control system, a currently input audio signal, the control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system being configured to implement an audio encoder comprising a neural network trained according to any of the methods of EEEs 1-11;
[0131] encoding the currently input audio signal in a compressed audio format by the audio encoder; and
[0132] Outputs the encoded audio signal in a compressed audio format.
[0133] 13. An audio decoding method, comprising:
[0134] receiving a currently input compressed audio signal via a control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system being configured to: implement an audio decoder comprising a neural network that has been trained according to any of the methods of EEEs 1-11;
[0135] Decoding the currently input compressed audio signal through the audio decoder; and
[0136] Outputs the decoded audio signal.
[0137] 14. The method of EEE 13, further comprising reproducing the decoded audio signal via one or more transducers.
[0138] 15. An apparatus comprising:
[0139] interface systems; and
[0140] A control system comprising one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, the control system being configured to implement the method of any one of EEE 1-14.
[0141] 16. One or more non-transitory media having software stored thereon, the software comprising instructions for controlling one or more devices to perform the method of any one of EEE 1-14.
[0142] 17. An audio encoding device, comprising:
[0143] interface systems; and
[0144] A control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement an audio encoder comprising a neural network trained according to any of the methods described in EEE 1-11, wherein the control system is configured to:
[0145] Receive the current input audio signal;
[0146] Encode the current input audio signal in a compressed audio format;
[0147] Outputs the encoded audio signal in a compressed audio format.
[0148] 18. An audio encoding device, comprising:
[0149] interface systems; and
[0150] A control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement an audio encoder comprising a neural network trained according to the following operations, the operations comprising:
[0151] receiving an input training audio signal through the neural network and through the interface system;
[0152] generating an encoded training audio signal based on the input training audio signal through a neural network;
[0153] decoding, via a control system, the encoded training audio signal to produce a decoded training audio signal;
[0154] Receiving, by a loss function generation module implemented via the control system, a decoded training audio signal and a true audio signal;
[0155] generating a loss function value corresponding to the decoded training audio signal through the loss function generation module, wherein generating the loss function value comprises: applying a psychoacoustic model; and
[0156] Train the neural network based on the loss function value;
[0157] The audio encoder is further configured to:
[0158] Receive the current input audio signal;
[0159] Encode the current input audio signal in a compressed audio format;
[0160] Outputs the encoded audio signal in a compressed audio format.
[0161] 19. A system comprising an audio decoding device, comprising:
[0162] Interface system;
[0163] A control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement an audio decoder comprising a neural network trained according to the following operations, the operations comprising:
[0164] receiving an input training audio signal through the neural network and through the interface system;
[0165] generating an encoded training audio signal based on the input training audio signal through a neural network;
[0166] decoding, via a control system, the encoded training audio signal to generate a decoded training audio signal;
[0167] receiving, by a loss function generation module implemented via a control system, a decoded training audio signal and a ground-truth audio signal;
[0168] generating a loss function value corresponding to the decoded training audio signal through a loss function generation module, wherein generating the loss function value includes applying a psychoacoustic model; and
[0169] Train the neural network based on the loss function value;
[0170] The audio decoder is further configured as follows:
[0171] receiving a currently input encoded audio signal in a compressed audio format;
[0172] Decoding the currently input encoded audio signal in a decompressed audio format; and
[0173] Outputs the decoded audio signal in decompressed audio format.
[0174] 20. The system of EEE 19, wherein the system further comprises one or more transducers configured to reproduce the decoded audio signal.
Claims
1. A computer-implemented method for training an encoding portion of a neural network implemented via a control system, the control system comprising one or more processors and one or more non-transitory storage media, the method comprising: receiving an input audio signal comprising audio data; generating an encoded audio signal based on the input audio signal through an encoding portion of the neural network; training the encoding portion based on a loss function value, wherein the loss function value is generated based on at least one of a decoded audio signal and / or a decoded transform coefficient obtained by decoding the encoded audio signal, and a true-valued audio signal, wherein the training involves updating at least one weight of the encoding portion, Generating the loss function value includes applying a psychoacoustic model.
2. The method of claim 1, wherein generating the loss function value by applying the psychoacoustic model further comprises calculating a noise masking ratio.
3. The method of claim 1 or 2, wherein training the encoding portion of the neural network comprises backpropagation based on the loss function value.
4. An audio encoder comprising an encoding portion of a neural network trained according to the method of any one of claims 1 to 3, wherein The audio encoder is further configured to: Receive the current input audio signal; encoding the currently input audio signal in a compressed audio format; and Outputs the encoded signal in a compressed audio format.
5. An audio encoding device comprising: Interface system; as well as A control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system being configured to implement the audio encoder of claim 4.
6. A computer-implemented method for training a decoding portion of a neural network implemented via a control system, the control system comprising one or more processors and one or more non-transitory storage media, the method comprising: decoding the encoded audio signal by a decoding portion of the neural network to produce at least one of a decoded audio signal and / or decoded transform coefficients; receiving, by a loss function generation module implemented via the control system, the at least one of a decoded audio signal and / or the decoded transform coefficients, and a true-valued audio signal; generating, by the loss function generating module, a loss function value corresponding to the at least one of the decoded audio signal and / or the decoded transform coefficient; training the decoding portion of the neural network based on the loss function value, wherein training involves updating at least one weight of the decoding portion of the neural network, Generating the loss function value includes applying a psychoacoustic model.
7. The method of claim 6, wherein generating the loss function value by applying the psychoacoustic model further comprises calculating a noise masking ratio.
8. An audio decoder comprising a decoding portion of a neural network trained according to the method of any one of claims 6 to 7, wherein: The audio decoder is further configured to: receiving a currently input encoded audio signal in a compressed audio format; decoding the currently input encoded audio signal in a decompressed audio format; and Outputs the decoded audio signal in decompressed audio format.
9. An audio decoding device, comprising: Interface system; A control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system being configured to implement an audio decoder comprising a decoding portion of a neural network trained according to the method of any one of claims 6-7, The audio decoder is further configured to: receiving a currently input encoded audio signal in a compressed audio format; decoding the currently input encoded audio signal in a decompressed audio format; and Outputs the decoded audio signal in decompressed audio format.
10. A system comprising the audio decoding device according to claim 9, wherein: The system also includes one or more transducers configured to reproduce the decoded audio signal.
11. A non-transitory storage medium having software stored thereon, the software comprising instructions for controlling one or more devices to execute the method according to any one of claims 1 to 3 and 6 to 7.
12. An apparatus for training an encoding portion of a neural network, comprising: one or more processors; and One or more non-transitory storage media having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 3.
13. An apparatus for training a decoding portion of a neural network, comprising: one or more processors; and One or more non-transitory storage media having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 6 to 7.
14. An apparatus comprising means for performing the method of any one of claims 1-3 and 6-7.
15. A computer program product having instructions which, when executed by a computing device or system, cause the computing device or system to perform the method of any one of claims 1-3 and 6-7.
Citation Information
Patent Citations
A bandwidth extender
CN103026407A
Device specific multi-channel data compression
US20180018990A1