Perceptually-based loss functions for audio encoding and decoding based on machine learning

A perceptually based loss function trained neural networks improve audio encoding and decoding by considering human auditory perception, enhancing quality beyond traditional methods.

JP2025126177APending Publication Date: 2025-08-28DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025089514
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-04
Filing Date
2025-05-29
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing audio encoding and decoding technologies fail to effectively maintain audio quality while minimizing bit usage, as they do not adequately consider human auditory perception.

Method used

Implementing a perceptually based loss function, such as a psychoacoustic model, to train neural networks for audio encoding and decoding, which models human hearing thresholds and masking effects to improve perceptual quality.

Benefits of technology

Enhances the perceptual quality of audio signals by minimizing audible differences, outperforming traditional methods like MSE-based training in subjective and objective evaluations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126177000001_ABST
    Figure 2025126177000001_ABST
Patent Text Reader

Abstract

To provide computer-implemented methods for training a neural network, as well as for implementing audio encoders and decoders via trained neural network.SOLUTION: The neural network may receive an input audio signal, generate an encoded audio signal, and decode the encoded audio signal. A loss function generating module may receive the decoded audio signal and a ground truth audio signal, and may generate a loss function value corresponding to the decoded audio signal. Generating the loss function value may involve applying a psychoacoustic model. The neural network may be trained based on the loss function value. The training may involve updating at least one weight of the neural network.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to processing audio signals, and in particular to encoding and decoding audio data. [Background technology]

[0002] An audio codec is a device or computer program that can encode and / or decode digital audio data, given a particular audio file or streaming media audio format. The primary goal of an audio codec is usually to represent the audio signal using the minimum number of bits while maintaining a reasonable degree of audio quality within those bits. Such audio data compression can reduce both the storage space required for the audio data and the bandwidth required to transmit the audio data. Summary of the Invention

[0003] Various audio processing methods are disclosed herein. Some such methods include receiving an input audio signal by a neural network implemented by a control system including one or more processors and one or more non-transitory storage media. Such methods may include generating an encoded audio signal by the neural network and based on the input audio signal. Some such methods may include decoding the encoded audio signal by the control system to generate a decoded audio signal, and receiving the decoded audio signal and a ground truth audio signal by a loss function generation module implemented by the control system. Such methods may include generating, by the loss function generation module, a loss function value corresponding to the decoded audio signal. Generating the loss function value may include applying a psychoacoustic model. Such methods may include training the neural network based on the loss function value. The training step may include updating at least one weight of the neural network.

[0004] According to some implementations, training the neural network may include backpropagation based on the loss function values. In some examples, the neural network may include an autoencoder. Training the neural network may include varying a physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network.

[0005] In some implementations, a first portion of the neural network may generate the encoded audio signal, and a second portion of the neural network may decode the encoded audio signal. In some such implementations, the first portion of the neural network may include an input neuron layer and multiple hidden neuron layers. The input neuron layer may, in some examples, include more neurons than the final hidden neuron layer. At least some neurons in the first portion of the neural network may be configured with a rectified linear unit (ReLU) activation function. In some examples, at least some neurons in the hidden layer of the second portion of the neural network may be configured with a ReLU activation function, and at least some neurons in the output layer of the second portion may be configured with a sigmoid activation function.

[0006] According to some examples, the psychoacoustic model may be based at least in part on one or more psychoacoustic masking thresholds. In some implementations, the psychoacoustic model may include modeling an outer ear transfer function, grouping into critical bands, frequency-domain masking (including but not limited to level-dependent spread), modeling frequency-dependent hearing thresholds, and / or calculating a noise-to-mask ratio. In some examples, the loss function may include calculating an average noise-to-mask ratio, and the training may include minimizing the average noise-to-mask ratio.

[0007] Several audio encoding methods and apparatus are disclosed herein. In some examples, an audio encoding method may include receiving a currently input audio signal by a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement an audio encoder including a neural network trained according to any of the methods disclosed herein. Such a model may include encoding, by the audio encoder, the currently input audio signal into a compressed audio format and outputting the encoded audio signal in the compressed audio format.

[0008] Several audio decoding methods and apparatuses are disclosed herein. In some examples, an audio decoding method may include receiving a currently input compressed audio signal by a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement an audio decoder including a neural network trained according to any of the methods disclosed herein. Such a method may include decoding the currently input compressed audio signal by the audio decoder and outputting a decoded audio signal. Some such methods may include reproducing the decoded audio signal by one or more transducers.

[0009] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including, but not limited to, random access memory (RAM), read-only memory (ROM), etc. Accordingly, various novel aspects of the subject matter described in this disclosure may be implemented in a non-transitory medium having software stored thereon. The software may include, for example, instructions for controlling at least one device to process audio data. The software may be executable by, for example, one or more components of a control system as disclosed herein. The software may include, for example, instructions for performing one or more of the methods disclosed herein.

[0010] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more apparatuses may be configured to at least partially perform the methods disclosed herein. In some implementations, an apparatus may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces. The control system may include at least one of a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic element, discrete gate or transistor logic, or discrete hardware components. Thus, in some implementations, the control system may include one or more processors and one or more non-transitory storage media operably coupled to the one or more processors.

[0011] According to some such examples, a device may include an interface system and a control system. The control system may be configured, for example, to perform one or more of the methods disclosed herein. For example, the control system may be configured to implement a speech encoder. The speech encoder may include a neural network trained according to one or more of the methods disclosed herein. The control system may be configured to receive a currently input speech signal, encode the currently input speech signal into a compressed speech format, and output (e.g., via the interface system) the encoded speech signal in the compressed speech format.

[0012] Alternatively or additionally, the control system may be configured to implement an audio decoder. The audio decoder may include a neural network trained according to a process including receiving an input training audio signal by the neural network and by the interface system, and generating an encoded training audio signal by the neural network and based on the input training audio signal. The process may include decoding the encoded training audio signal by the control system to generate a decoded training audio signal, and receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system. The process may include generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal. Generating the loss function value may include applying a psychoacoustic model. The process may include training the neural network based on the loss function value.

[0013] The audio encoder may be further configured to receive a currently input audio signal, encode the currently input audio signal into a compressed audio format, and output the encoded audio signal in the compressed audio format.

[0014] In some implementations, the disclosed system may include an audio decoding device. The audio decoding device may include an interface system and a control system, the control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors. The control system may be configured to implement the audio decoder.

[0015] The audio decoder may include a neural network trained according to a process including receiving an input training audio signal by the neural network and by the interface system, and generating an encoded training audio signal by the neural network and based on the input training audio signal. The process may include decoding the encoded training audio signal by the control system to generate a decoded training audio signal, and receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system. The process may include generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal. Generating the loss function value may include applying a psychoacoustic model. The process may include training the neural network based on the loss function value.

[0016] The audio decoder may be further configured to receive a currently input encoded audio signal in a compressed audio format, decode the currently input encoded audio signal into an uncompressed audio format, and output a decoded audio signal in the uncompressed audio format. According to some implementations, the system may include one or more transducers configured to reproduce the decoded audio signal.

[0017] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. The relative dimensions of the following drawings may not be drawn to scale. Like numbers and designations in the various drawings generally indicate similar elements. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a block diagram illustrating example components of a device that may be configured to perform at least some of the methods disclosed herein.

[0019] [Figure 2] 1 illustrates blocks for implementing a machine learning process according to a perceptually based loss function, according to one example.

[0020] [Figure 3] 1 illustrates an example of a neural network training process according to some implementations disclosed herein.

[0021] [Figure 4-1] 1 shows an alternative example of a neural network suitable for implementing some of the methods disclosed herein. [Figure 4-2] 1 shows an alternative example of a neural network suitable for implementing some of the methods disclosed herein.

[0022] [Figure 5A] 1 is a flow diagram outlining blocks of a method for training a neural network for audio encoding and decoding, according to an example.

[0023] [Figure 5B] FIG. 1 is a flow diagram outlining blocks of a method for using a trained neural network for audio coding, according to an example.

[0024] [Figure 5C] FIG. 1 is a flow diagram outlining blocks of a method for using a trained neural network for speech decoding, according to an example.

[0025] [Figure 6] FIG. 2 is a block diagram illustrating a loss function generation module configured to generate a loss function based on mean squared error.

[0026] [Figure 7A] 1 is a graph of a function that approximates the typical acoustic response of the human ear canal.

[0027] [Figure 7B] 1 illustrates a loss function generation module configured to generate a loss function based on the typical acoustic response of the human ear canal.

[0028] [Figure 8] 1 illustrates a loss function generation module configured to generate a loss function based on a band operation.

[0029] [Figure 9A] 1 illustrates the processes involved in frequency masking according to some examples.

[0030] [Figure 9B] An example of a spreading function is shown below.

[0031] [Figure 10] 10 shows an example of an alternative implementation of the loss function generation module.

[0032] [Figure 11] 10 is an example of objective test results for some of the disclosed implementations.

[0033] [Figure 12] We present examples of subjective test results for speech data corresponding to a male speaker generated by neural networks trained with different types of loss functions.

[0034] [Figure 13] 12. FIG. 13 shows an example of subjective test results for speech data corresponding to a female speaker generated by a neural network trained with the same type of loss function shown in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0035] The following description is directed to particular implementations for purposes of illustrating some novel aspects of the present disclosure and example contexts in which the novel aspects may be implemented. However, the teachings herein may be applied in a variety of different ways. Moreover, the described embodiments may be implemented in various hardware, software, firmware, etc. For example, aspects of the present application may be realized, at least in part, in an apparatus, a system including more than one device, a method, a computer program product, etc. Accordingly, aspects of the present application may take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), and / or embodiments combining both software and hardware aspects. Such embodiments may be referred to herein as "circuits," "modules," or "engines." Some aspects of the present application may take the form of a computer program product embodied in one or more non-transitory medium(s) having computer-readable program code embodied thereon. Such non-transitory media may include, for example, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. Accordingly, the teachings of the present disclosure are not limited to the implementations illustrated and / or described herein, but rather have broad applicability.

[0036] The inventors have investigated various machine learning methods related to audio data processing, including, but not limited to, audio data encoding and decoding. In particular, the inventors have investigated various methods of training different types of neural networks using loss functions related to how humans begin to perceive sound. The effectiveness of each of these loss functions was evaluated according to audio data generated by neural network encoding. The audio data was evaluated according to subjective and objective criteria. In some examples, audio data processed by a neural network trained using a loss function based on mean squared error was used as a basis for evaluating audio data generated according to the methods disclosed herein. In some examples, the subjective evaluation process includes having human listeners evaluate the resulting audio data and obtaining listener feedback.

[0037] The techniques disclosed herein build on the above-mentioned research. This disclosure provides various examples of using a perceptually based loss function to train a neural network for audio data encoding and / or decoding. In some examples, the perceptually based loss function is based on a psychoacoustic model. The psychoacoustic model may be based, for example, at least in part, on one or more psychoacoustic masking thresholds. In some implementations, the psychoacoustic model may include modeling an outer ear transfer function, grouping into critical bands, frequency-domain masking (including, but not limited to, level-dependent spread), modeling frequency-dependent hearing thresholds, and / or calculating a noise-to-mask ratio. In some implementations, the loss function may include calculating an average noise-to-mask ratio. In some such examples, the training process may include minimizing the average noise-to-mask ratio.

[0038] FIG. 1 is a block diagram illustrating example components of a device that may be configured to perform at least some of the methods disclosed herein. In some examples, device 105 may be or include a personal computer, desktop computer, or other local device configured to provide audio processing. In some examples, device 105 may be or include a server. According to some examples, device 105 may be a client device configured to communicate with a server via a network interface. The components of device 105 may be implemented by hardware, by software stored on a non-transitory medium, by firmware, and / or by a combination thereof. The types and number of components shown in FIG. 1 and other figures disclosed herein are provided by way of example only. Alternative implementations may include more, fewer, and / or different components.

[0039] In this example, device 105 includes interface system 110 and control system 115. Interface system 110 may include one or more network interfaces, one or more interfaces between control system 115 and a memory system, and / or one or more external interfaces (e.g., one or more universal serial bus (USB) interfaces). In some implementations, interface system 110 may include a user interface system. The user interface system may be configured to receive input from a user. In some implementations, the user interface system may be configured to provide feedback to a user. For example, the user interface system may include one or more displays corresponding to a touch and / or gesture detection system. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include devices that provide tactile feedback, such as motors, vibrators, etc. The control system 115 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic elements, discrete gate or transistor logic, and / or discrete hardware components.

[0040] In some examples, the device 105 may be implemented in a single device. However, in some implementations, the device 105 may be implemented in more than one device. In some such implementations, the functionality of the control system 115 may be included in more than one device. In some examples, the device 105 may be a component of another device.

[0041] 2 illustrates blocks for implementing a machine learning process according to a perceptually based loss function, according to one example. In this example, an input audio signal 205 is provided to a machine learning module 210. In some examples, the input audio signal 205 may correspond to human speech. However, in other examples, the input audio signal 205 may correspond to other sounds, such as music, etc.

[0042] According to some examples, elements of system 200 include, but are not limited to, machine learning module 210, which may be implemented by one or more control systems, such as control system 115. Machine learning module 210 may receive input audio signal 205, for example, by an interface system, such as interface system 110. In some examples, machine learning module 210 may be configured to implement one or more neural networks, such as the neural networks disclosed herein. However, in other implementations, machine learning module 210 may be configured to implement one or more other types of machine learning, such as Non-Negative Matrix Factorization, Robust Principal Component Analysis, Sparse Coding, Probabilistic Latent Component Analysis, etc.

[0043] 2, machine learning module 210 provides output audio signal 215 to loss function generation module 220. Loss function generation module 225 and optional ground truth module 220 may be implemented by a control system, such as control system 115. In some examples, loss function generation module 225, machine learning module 210, and optional ground truth module 220 may be implemented by the same device, while in other examples, loss function generation module 225, optional ground truth module 220, and machine learning module 210 may be implemented by different devices.

[0044] According to this example, the loss function generation module 225 receives the input audio signal 205 and uses the input audio signal 205 as the "ground truth" for determining the error. However, in some alternative implementations, the loss function generation module 225 may receive ground truth data from the optional ground truth module 220. Such implementations may include tasks such as speech enhancement or speech noise removal, where the ground truth is not the original input audio signal. Whether the ground truth data is the input audio signal 205 or data received from the optional ground truth module, the loss function generation module 225 evaluates the output audio signal according to the loss function algorithm and the ground truth data and provides a loss function value 230 to the machine learning module 210. In some such implementations, the machine learning module 210 includes an implementation of the optimization module 315, described below with reference to FIG. 3 . In other examples, the system 200 includes an implementation of the optimization module 315 that is separate from, but in communication with, the machine learning module 210 and the loss function generation module 225. Various example loss functions are disclosed herein. In this example, the loss function generation module 225 applies a perceptually based loss function, which may be based on a psychoacoustic model. According to this example, the machine learning process (e.g., the process of training a neural network) implemented by the machine learning module 210 is based at least in part on the loss function value 230.

[0045] Utilizing a perceptually based loss function, such as a loss function based on a psychoacoustic model, for machine learning (e.g., to train a neural network) can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of an output audio signal generated by a machine learning process using a traditional loss function based on mean squared error (MSE), L1-norm, etc. For example, a neural network trained for a given length of time with a loss function based on a psychoacoustic model can improve the perceptual quality of the output audio signal 215 compared to the perceptual quality of an output audio signal generated by a neural network having the same architecture trained for the same length of time with an MSE-based loss function. Furthermore, a neural network trained to converge with a loss function based on a psychoacoustic model can typically generate an output audio signal of higher perceptual quality than the output audio signal of a neural network having the same architecture trained to converge with an MSE-based loss function.

[0046] Some disclosed loss functions utilize psychoacoustic principles to determine which differences in the output audio signal 215 are audible and inaudible to the average person. In some examples, loss functions based on psychoacoustic models may utilize psychoacoustic phenomena such as time masking, frequency masking, loudness contours, level-dependent masking, and / or human hearing thresholds. In some implementations, the perceptual loss function may operate in the time domain, while in other implementations, the perceptual loss function may operate in the frequency domain. In alternative implementations, the perceptual loss function may include both time-domain and frequency-domain operations. In some examples, the loss function may calculate the loss function using one frame input, while in other examples, the loss function may calculate the loss function using multiple input frames.

[0047] FIG. 3 illustrates an example of a neural network training process according to some implementations disclosed herein. As with other diagrams provided herein, the number and types of elements are merely exemplary. According to some examples, elements of system 301 may be implemented by one or more control systems, such as control system 115. In the example shown in FIG. 3, neural network 300 is an autoencoder. Techniques for designing autoencoders are described in chapter 14 of Goodfellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning (MIT Press, 2016), incorporated herein by reference.

[0048] Neural network 300 includes layers of nodes, also referred to herein as "neurons." Each neuron has a real-valued activation function. Its output, commonly referred to as "activation," defines the neuron's output given an input or set of inputs. According to some examples, neurons of neural network 300 may utilize sigmoid activation functions, ELU activation functions, and / or tanh activation functions. Alternatively or additionally, neurons of speech neural network 300 may utilize rectified linear unit (ReLU) activation functions.

[0049] Each connection between neurons (also called a "synapse") has a modifiable real-valued weight. Neurons may be input neurons (receiving data from outside the network), output neurons, or hidden neurons that modify data on the way from input neurons to output neurons. In the example shown in FIG. 3, neurons in neuron layer 1 are input neurons, neurons in neuron layer 7 are output neurons, and neurons in neuron layers 2-6 are hidden neurons. While five hidden neuron layers are shown in FIG. 3, some implementations may include more or fewer hidden layers. Some implementations of neural network 300 may include more or fewer hidden layers, e.g., 10 or more hidden layers. For example, some implementations may include 10, 20, 30, 40, 50, 60, 70, 80, 90, or more hidden layers.

[0050] Here, a first portion (encoder portion 305) of neural network 300 is configured to generate an encoded audio signal, and a second portion (decoder portion 310) of neural network 300 is configured to decode the encoded audio signal. In this example, the encoded audio signal is a compressed audio signal, and the decoded audio signal is an uncompressed (decompressed) audio signal. Thus, input audio signal 205 is compressed by encoder portion 305, as suggested by the reduced size of the blocks used to describe neuron layers 1-4. In some examples, the input neuron layer may include more neurons than at least one of the hidden neuron layers of encoder portion 305. However, in alternative implementations, neuron layers 1-4 may all have the same or substantially similar number of neurons.

[0051] Thus, the compressed audio signal provided by the encoder portion 305 is then decoded by a layer of neurons in the decoder portion 310 to construct the output signal 215, which is an estimate of the input audio signal 205. A perceptual loss function, such as a psychoacoustic-based loss function, may be used during the training phase to determine updates to the parameters of the neural network 300. These parameters can later be used to decode (e.g., decompress) any encoded (e.g., compressed) audio signal, using weights determined by the parameters received from the training algorithm. In other words, encoding and decoding may be performed separately from the training process after satisfactory weights for the neural network 300 have been determined.

[0052] According to this example, the loss function generation module 225 receives at least a portion of the audio signal 205 and uses it as ground truth data. Here, the loss function generation module 225 evaluates the output audio signal according to the loss function algorithm and the ground truth data and provides a loss function value 230 to the optimization module 315. In this example, the optimization module 315 is initialized with information about the neural network and the loss function to be used by the loss function generation module 225. According to this example, the optimization module 315 uses this information, along with the loss values ​​it receives from the loss function generation module 225, to calculate the gradient of the loss function with respect to the neural network weights. Once this gradient is known, the optimization module 315 uses an optimization algorithm to generate updates 320 to the neural network weights. According to some implementations, the optimization module 315 may utilize an optimization algorithm such as a Stochastic Gradient Descent or Adam optimization algorithm. The Adam optimization algorithm is disclosed in D.P. Kingma and J.L. Ba, “Adam: a Method for Stochastic Optimization,” in Proceedings of the International Conference on Learning Representations (ICLR), 2015, pp. 1-15, which is incorporated herein by reference. In the example shown in FIG. 3 , optimization module 315 is configured to provide updates 320 to neural network 300. In this example, loss function generation module 225 applies a perceptually based loss function, which may be based on a psychoacoustic model. According to this example, the process of training neural network 300 is based at least in part on backpropagation. This backpropagation is indicated in FIG. 3 by the dotted arrows between neuron layers. Backpropagation (also known as “backpropagation”) is a method used in neural networks to calculate the error contribution of each neuron after a batch of data has been processed.The backpropagation technique is sometimes called backpropagation of errors because the errors can be distributed backward through the neural network layers computed at the output.

[0053] Neural network 300 may be implemented by a control system, such as control system 115 described above with reference to FIG. 1 . Thus, training neural network 300 may include varying the physical states of non-transitory storage media locations corresponding to weights in neural network 300. The storage media locations may be portions of one or more storage media accessible by or portions of the control system. Weights, as described above, correspond to connections between neurons. Training neural network 300 may include varying the physical states of non-transitory storage media locations corresponding to values ​​of activation functions of neurons.

[0054] 4-1(A)-(C) illustrate alternative examples of neural networks suitable for implementing some of the methods disclosed herein. According to these examples, the input and hidden neurons utilize rectified linear unit (ReLU) activation functions, and the output neurons utilize sigmoid activation functions. However, alternative implementations of neural network 300 may include other activation functions and / or other combinations of activation functions, including, but not limited to, Exponential Linear Unit (ELU) and / or tanh activation functions.

[0055] In these examples, the input audio data is 256-dimensional audio data. In the example shown in FIG. 4-1(A), the encoder portion 305 compresses the input audio data into 32-dimensional audio data, providing a maximum 8x compression (reduction). In the example shown in FIG. 4-1(B), the encoder portion 305 compresses the input audio data into 16-dimensional audio data, providing a maximum 16x compression (reduction). The neural network 300 shown in FIG. 4-1(C) includes an encoder portion 305 that compresses the input audio data into 8-dimensional audio data, providing a maximum 32x compression (reduction). The inventors have conducted hearing tests based on neural networks of the type shown in FIG. 4-1(B). Some results are described below.

[0056] FIG. 4-2(D) shows an example block diagram of an encoder portion of an autoencoder according to an alternative example. The encoder portion 305 may be implemented by a control system, such as the control system 115 described above with reference to FIG. 1. The encoder portion 305 may be implemented by one or more processors of the control system in accordance with software stored on one or more non-transitory storage media, for example. The number and types of elements shown in FIG. 4-2(D) are merely examples. Other implementations of the encoder portion 305 may include more, fewer, or different elements.

[0057] In this example, encoder portion 305 includes three neuron layers. According to some examples, the neurons in encoder portion 305 may utilize a ReLU activation function. However, according to some alternative examples, the neurons in encoder portion 305 may utilize a sigmoid activation function and / or a tanh activation function. The neurons in neuron layers 1-3 maintain their N-dimensional state while processing the N-dimensional input data. Layer 450 is configured to receive the output of neuron layer 3 and apply a pooling algorithm. Pooling is a form of nonlinear downsampling. According to this example, layer 450 is configured to divide the output of neuron layer 3 into M sets of non-overlapping portions or "sub-regions" and apply a max-pooling function that outputs the maximum value for each sub-region.

[0058] FIG. 5A is a flow diagram outlining the blocks of a method for training a neural network for audio encoding and decoding, according to one example. Method 500 may, in some examples, be performed by the device of FIG. 1 or by another type of device. In some examples, the blocks of method 500 may be implemented by software stored on one or more non-transitory media. The blocks of method 500, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.

[0059] Here, block 505 includes receiving an input speech signal by a neural network implemented by a control system including one or more processors and one or more non-transitory storage media. In some examples, the neural network may or may not include an autoencoder. According to some examples, block 505 may include the control system of FIG. 1 receiving the input speech signal via interface system 110. In some examples, block 505 may include neural network 300 receiving input speech signal 205 as described above with reference to FIGS. 2-4C. In some implementations, input speech signal 205 may include at least a portion of a publicly available speech dataset, such as the TIMIT speech dataset. TIMIT is a dataset of phoneme and word transcription speech of American English speakers of different genders and dialects. TIMIT was commissioned by the Defense Advanced Research Projects Agency (DARPA). The TIMIT corpus design is a collaboration between Texas Instruments (TI), Massachusetts Institute of Technology (MIT), and SRI International. According to some examples, method 500 may include transforming input audio signal 205 from the time domain to the frequency domain, for example, with a fast Fourier transform (FFT), a discrete cosine transform (DCT), or a short-time Fourier transform (STFT). In some implementations, min / max scaling may be applied to input audio signal 205 before block 510.

[0060] According to this example, block 510 includes generating an encoded audio signal by a neural network and based on the input audio signal. The encoded audio signal may be or include a compressed audio signal. Block 510 may be performed by an encoder portion of a neural network, such as, for example, encoder portion 305 of neural network 300 described herein. However, in other examples, block 510 may include generating the encoded audio signal by an encoder that is not part of the neural network. In some such examples, a control system implementing the neural network may also include an encoder that is not part of the neural network. For example, a neural network may include a decoding portion but not an encoding portion.

[0061] In this example, block 515 includes decoding, by a control system, the encoded audio signal to generate a decoded audio signal. The decoded audio signal may be or may include an uncompressed audio signal. In some implementations, block 515 may include generating decoded transform coefficients. Block 515 may be performed by a decoder portion of a neural network, such as, for example, decoder portion 310 of neural network 300 described herein. However, in other examples, block 510 may include generating the decoded audio signal and / or the decoded transform coefficients by a decoder that is not part of the neural network. In some such examples, a control system implementing a neural network may also include a decoder that is not part of the neural network. For example, a neural network may include an encoding portion but not a decoding portion.

[0062] Thus, in some implementations, a first portion of the neural network may be configured to generate an encoded audio signal, and a second portion of the neural network may be configured to decode the encoded audio signal. In some such implementations, the first portion of the neural network may include an input neuron layer and multiple hidden neuron layers. In some examples, the input neuron layer may include more neurons than at least one of the hidden neuron layers of the first portion. However, in alternative implementations, the input neuron layer may have the same number of neurons as, or a substantially similar number of neurons to, the hidden neuron layer of the first portion.

[0063] According to some examples, at least some neurons in a first portion of the neural network may be configured with a rectified linear unit (ReLU) activation function. In some implementations, at least some neurons in a hidden layer of a second portion of the neural network may be configured with a rectified linear unit (ReLU) activation function. According to some such implementations, at least some neurons in an output layer of the second portion may be configured with a sigmoid activation function.

[0064] In some implementations, block 520 may include receiving, by a loss function generation module implemented by the control system, the decoded audio signal and / or the decoded transform coefficients and a ground truth signal. The ground truth signal may include, for example, the ground truth audio signal and / or the ground truth transform coefficients. In some such examples, the ground truth signal may be received from a ground truth module such as ground truth module 220 shown in FIG. 2 and described above. However, in some implementations, the ground truth signal may be (or may include) the input audio signal or a portion of the input audio signal. The loss function generation module may be, for example, an instance of loss function generation module 225 disclosed herein.

[0065] According to some implementations, block 525 may include generating, by a loss function generation module, loss function values ​​corresponding to the decoded speech signal and / or the decoded transform coefficients. In some such implementations, generating the loss function values ​​may include applying a psychoacoustic model. In the example shown in FIG. 5A , block 530 includes training a neural network based on the loss function values. The training may include updating at least one weight in the neural network. In some such examples, an optimizer, such as the optimization module 315 described above with reference to FIG. 3 , may be initialized with information about the neural network and the loss function used by the loss function generation module 225. The optimization module 315 uses this information, along with the loss function values ​​it receives from the loss function generation module 225, to calculate the gradient of the loss function with respect to the neural network weights. After calculating the gradient, the optimization module 315 may use an optimization algorithm to generate updates to the neural network weights and provide these updates to the neural network. Training the neural network may include backpropagation based on updates provided by optimization module 315. Techniques for detecting and addressing overfitting during the process of training a neural network are described in chapters 5 and 7 of Goodfellow, Ian, Yoshua Bengio, and Aaron Courville, Deep Learning (MIT Press, 2016), which is incorporated herein by reference. Training the neural network may include varying the physical state of at least one non-transitory storage location corresponding to at least one weight or at least one activation function value of the neural network.

[0066] The psychoacoustic model may vary depending on the particular implementation. According to some examples, the psychoacoustic model may be based at least in part on one or more psychoacoustic masking thresholds. In some implementations, applying the psychoacoustic model may include modeling an ear transfer function, grouping into critical bands, frequency-domain masking (including but not limited to level-dependent diffusion), modeling frequency-dependent hearing thresholds, and / or calculating a noise-to-masking ratio. Some examples are described below with reference to FIGS. 6-10B.

[0067] In some implementations, the loss function generation module's determination of the loss function may include calculating a noise-to-masking ratio, such as an average noise-to-masking ratio (NMR). The training process may include minimizing the average NMR. Some examples are described below.

[0068] According to some examples, training the neural network may continue until the loss function becomes relatively "flat," such that the difference between the current loss function value and a previous loss function value (such as the previous loss function value) is at or below a threshold. In the example shown in FIG. 5, training the neural network may include repeating at least some of blocks 505-535 until the difference between the current loss function value and the previous loss function value is less than or equal to a predetermined value.

[0069] After the neural network is trained, the neural network (or portions thereof) may be used to process audio data, such as to encode or decode audio data. FIG. 5B is a flow diagram outlining blocks of a method of using a trained neural network for audio coding, according to one example. Method 540 may, in some examples, be performed by the device of FIG. 1 or by another type of device. In some examples, the blocks of method 540 may be implemented by software stored on one or more non-transitory media. The blocks of method 540, as well as other methodologies described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.

[0070] In this example, block 545 includes receiving a currently input audio signal. In this example, block 545 includes receiving a currently input audio signal by a control system including one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, where the control system is configured to implement a speech encoder including a neural network trained according to one or more of the methods disclosed herein.

[0071] In some examples, the training process may include receiving an input training audio signal by the neural network and via the interface system; generating an encoded training audio signal by the neural network and based on the input training audio signal; decoding the encoded training audio signal by the control system to generate a decoded training audio signal; receiving the decoded training audio signal and a ground truth audio signal by a loss function generation module implemented by the control system; and generating, by the loss function generation module, a loss function value corresponding to the decoded training audio signal, wherein generating the loss function value includes applying a psychoacoustic model; and training the neural network based on the loss function value.

[0072] According to this implementation, block 550 includes encoding, by an audio encoder, the currently input audio signal into a compressed audio format, where block 555 includes outputting the encoded audio signal in the compressed audio format.

[0073] 5C is a flow diagram outlining the blocks of a method for using a trained neural network for audio decoding, according to one example. Method 560 may, in some examples, be performed by the device of FIG. 1 or by another type of device. In some examples, the blocks of method 560 may be implemented by software stored on one or more non-transitory media. The blocks of method 560, as well as other methods described herein, are not necessarily performed in the order shown. Furthermore, such methods may include more or fewer blocks than those shown and / or described.

[0074] In this example, block 565 includes receiving a currently input compressed audio signal. In some such examples, the currently input compressed audio signal may be generated according to method 540 or a similar method. In this example, block 565 includes receiving the currently input compressed audio signal by a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors, where the control system is configured to implement an audio decoder including a neural network trained according to one or more of the methods disclosed herein.

[0075] According to this implementation, block 570 includes decoding the currently input compressed audio signal with an audio decoder. For example, block 570 may include decompressing the currently input compressed audio signal. Here, block 575 includes outputting the decoded audio signal. According to some examples, method 540 may include reproducing the decoded audio signal with one or more transducers.

[0076] As mentioned above, the inventors investigated various methods of training different types of neural networks using loss functions related to how humans begin to perceive sound. The effectiveness of each of these loss functions was evaluated on speech data generated by neural network coding. In some instances, speech data processed by neural networks trained using loss functions based on mean squared error (MSE) was used as a basis for evaluating speech data generated according to the methods disclosed herein.

[0077] 6 is a block diagram illustrating a loss function generation module configured to generate a loss function based on mean squared error, where the estimated magnitude of the audio signal generated by the neural network and the magnitude of the ground truth / true audio signal are both provided to the loss function generation module 225. The loss function generation module 225 generates a loss function value 230 based on the MSE value. The loss function value 230 may be provided to an optimization module configured to generate updates to the neural network weights for training.

[0078] The inventors evaluated several implementations of loss functions based at least in part on models of the acoustic response of one or more parts of the human ear, which may be referred to as “ear models.” Figure 7A is a graph of a function that approximates the typical acoustic response of the human ear canal.

[0079] 7B shows a loss function generation module configured to generate a loss function based on the typical acoustic response of the human ear canal. In this example, the function W is applied to both the neural network generated audio signal and the ground truth / true audio signal.

[0080] In some examples, the function W may be:

number

[0081] Equation 1 is used in the implementation of the Perceptual Evaluation of Audio Quality (PEAQ) algorithm, which models the acoustic response of the human ear canal. In Equation 1, f represents the frequency of the audio signal. In this example, the loss function generation module 225 generates a loss function value 230 based on the difference between the two resulting values. The loss function value 230 may be provided to an optimization module configured to generate updates to the neural network weights for training.

[0082] Compared to the audio signals generated by a neural network trained according to an MSE-based loss function, the audio signals generated by training a neural network with a loss function such as that shown in Figure 7B provide only a slight improvement. For example, using an objective standard based on Perceptual Objective Listening Quality Analysis (POLQA), the MSE-based audio data achieved a score of 3.41, while the audio data generated by training a neural network using the loss function such as that shown in Figure 7B achieved a score of 3.48.

[0083] In some experiments, the inventors tested audio signals generated by a neural network trained according to a loss function based on band manipulation. Figure 8 shows a loss function generation module configured to generate a loss function based on band manipulation. In this example, the loss function generation module 225 is configured to perform band manipulation on the audio signal generated by the neural network and the ground truth / true audio signal, and to calculate the difference between the results.

[0084] In some implementations, the band manipulation is based on "Zwicker" bands, which are critical bands defined according to chapter 6 (Critical Bands and Excitation) of Fastl, H., & Zwicker, E. (2007), Psychoacoustics: Facts and Models (3rd ed., Springer), which is incorporated herein by reference. In alternative implementations, the band manipulation is based on "Moore" bands, which are critical bands defined according to chapter 3 (Frequency Selectivity, Masking, and the Critical Band) of Moore, BCJ (2012), An Introduction to the Psychology of Hearing (Emerald Group Publishing), which is incorporated herein by reference. However, other examples may include other types of band manipulation known to those skilled in the art.

[0085] Based on the inventors' experiments, the inventors concluded that band manipulation alone is unlikely to provide satisfactory results. For example, using an objective standard based on POLQA, MSE-based speech data achieved a score of 3.41, while speech data generated by a neural network using band manipulation only achieved a score of 1.62.

[0086] In several experiments, the inventors tested audio signals generated by neural networks trained according to a loss function based at least in part on frequency masking. Figure 9A illustrates the process involved in frequency masking according to some examples. In this example, a spreading function is calculated in the frequency domain. This spreading function may be a level- and frequency-dependent function that can be estimated from the input audio signal, e.g., from each input audio frame. Next, convolution with the frequency spectrum of the input audio signal may be performed, which generates an excitation pattern. The result of the convolution between the input audio data and the spreading function is an approximation of how the human auditory filter responds to the excitation of the incoming audio. Thus, the process simulates the human hearing mechanism. In some implementations, the audio data is grouped into frequency bins, and the convolution process includes, for each frequency bin, convolving the spreading function with the corresponding audio data of that frequency bin.

[0087] The excitation pattern may be adjusted to produce a masking pattern. In some examples, the excitation pattern may be adjusted downwards, for example, by 20 dB, to produce a masking pattern.

[0088] 9B shows an example of a spreading function. According to this example, the spreading function is a simplified asymmetric trigonometric function that can be pre-calculated for efficient implementation. In this simple example, the vertical axis represents decibels and the horizontal axis represents Bark sub-bands. According to one such example, the spreading function is calculated as follows:

number

[0089] In equations 2 and 3, S l represents the slope of the portion of the spread function in Figure 9B to the left of the peak frequency, and S urepresents the slope of the portion of the spreading function to the right of the peak frequency. The units of slope are dB / Bark. In Equation 3, fc represents the center or peak frequency of the spreading function, and L represents the level or amplitude of the audio data. In some examples, to simplify the calculation of the spreading function, L may be considered a constant. According to some such examples, L may be 70 dB.

[0090] In some such implementations, the Excitation pattern may be calculated as follows:

number

[0091] In Equation 4, E represents the excitation function (also referred to herein as the excitation pattern), SF represents the spreading function, and BP represents the banded pattern of frequency-binned audio data. In some implementations, the excitation pattern may be adjusted to generate a masking pattern. In some examples, the excitation pattern may be adjusted downward, for example, by 20 dB, 24 dB, 27 dB, etc., to generate the masking pattern.

[0092] 10 shows an example of an alternative implementation of the loss function generation module. Elements of the loss function generation module 225 may be implemented by a control system, such as, for example, control system 115 described above with reference to FIG.

[0093] In this example, the reference speech signal x ref , referred to elsewhere herein as an instance of the ground truth signal, is provided to a fast Fourier transform (FFT) block 1005a of the loss function generation module 225. A test speech signal x is generated by a neural network, such as one of those disclosed herein, and is provided to an FFT block 1005b of the loss function generation module 225.

[0094] According to this example, the output of FFT block 1005a is provided to ear model block 1010a, and the output of FFT block 1005a is provided to ear model block 1010b. Ear model blocks 1010a and 1010b may be configured to provide a function based on, for example, the typical acoustic response of one or more portions of the human ear canal. In one such example, ear model blocks 1010a and 1010b may be configured to apply the function set forth above in Equation 1.

[0095] According to this implementation, the outputs of ear model blocks 1010a and 1010b are provided to a difference calculation block 1015. The difference calculation block 1015 is configured to calculate the difference between the output of ear model block 1010a and the output of ear model block 1010b. The output of difference calculation block 1015 may be thought of as an approximation of the noise present in the test signal x.

[0096] In this example, the output of ear model block 1010a is provided to band processing block 1020a, and the output of difference calculation block 1015 is provided to band processing block 1020b. Band processing blocks 1020a and 1020b are configured to apply the same type of band processing, which may be one of the band processings mentioned above (e.g., Zwicker or Moore band processing). However, in alternative implementations, band processing blocks 1020a and 1020b may be configured to apply any suitable band processing known to those skilled in the art.

[0097] The output of the band processing block 1020a is provided to a frequency masking block 1025, which is configured to apply a frequency masking process. The masking block 1025 may be configured to apply, for example, one or more of the frequency masking processes disclosed herein. As discussed above with reference to FIG. 9B, the use of a simplified frequency masking process can offer potential advantages. However, in alternative implementations, the masking block 1025 may be configured to apply one or more other frequency masking processes known to those skilled in the art.

[0098] According to this example, the output of the masking block 1025 and the output of the band processing block 1020b are both provided to a noise-to-mask ratio (NMR) calculation block 1030. As mentioned above, the output of the difference calculation block 1015 may be thought of as an approximation of the noise present in the test signal x. Therefore, the output of the band processing block 1020b may be thought of as a frequency-band processed version of the noise present in the test signal x. According to one example, the NMR calculation block 1030 may calculate the NMR as follows:

number

[0099] In Equation 5, BP noise represents the output of the band processing block 1020b, and MP represents the output of the masking block 1025. According to some examples, the NMR calculated by the NMR calculation block 1030 may be an average NMR over all frequency bands output by the band processing blocks 1020a and 1020b. The NMR calculated by the NMR calculation block 1030 may be used as a loss function value 230 to train a neural network, for example, as described above. For example, the loss function value 230 may be provided to an optimization module configured to generate updated weights for the neural network.

[0100] Figure 11 shows an example of objective test results for some of the disclosed implementations. Figure 11 shows a comparison between PESQ scores n for speech data generated by neural networks trained using loss functions based on MSE, power law, NMR-Zwicker (NMR based on band processing like Zwicker band processing but with fractionally narrower bands than those defined by Zwicker), and NMR-Moore (NMR based on Moore band processing). These results are based on the outputs of the neural networks described above with reference to Figure 4-1(B), and show that both the NMR-Zwicker and NMR-Moore results are somewhat better than the MSE and power law results.

[0101] Figure 12 shows an example of subjective test results for speech data corresponding to a male speaker generated by neural networks trained with various loss functions. In this example, the subjective test results are MUSHRA (MUltiple Stimulus Test with Hidden Reference and Anchor) ratings. MUSHRA, described in ITU-R BS.1534, is a well-known method for conducting codec listening tests to evaluate the perceptual quality of the output from lossy audio compression algorithms. The MUSHRA method has the advantage of simultaneously displaying multiple stimuli. As a result, subjects can directly perform arbitrary comparisons between them. The time required to perform a test using the MUSHRA method can be significantly reduced compared to other methods. This is partially true because results from all codecs are simultaneously expressed for the same sample. As a result, a t-test or repeated measures of variance can be used for statistical analysis. The numbers along the x-axis in Figure 12 are the identification numbers of different audio files.

[0102] More specifically, Figure 12 shows a comparison between the MUSHRA ratings of speech data generated by the same neural network trained using an MSE-based loss function, a power-law-based loss function, an NMR-Zwicker-based loss function, and an NMR-Moore-based loss function, as well as speech data generated by applying a 3.5 kHz low-pass filter (one of the standard "anchors" for the MUSHRA method), and reference speech data. In this example, MUSHRA ratings were obtained from 11 different listeners. As shown in Figure 12, the average MUSHRA rating for speech data generated by the neural network trained using the NMR-Moore-based loss function was significantly higher than any of the others. The difference was nearly 30 MUSHRA points, a rare large effect. The second highest average MUSHRA rating was for speech data generated by the neural network trained using the NMR-Zwicker-based loss function.

[0103] FIG. 13 shows an example of subjective test results for speech data corresponding to a female speaker generated by a neural network trained using the same type of loss function shown in FIG. 12. As in FIG. 12, the numbers along the x-axis in FIG. 13 are identification numbers of different speech files. In this example, the highest average MUSHRA rating was again assigned to the speech data generated by the neural network after training with the NMR-based loss function. Although the perceived differences between the NMR-Moore and NMR-Zwicker speech data and the other speech data were not as pronounced in this example as those shown in FIG. 12, the results shown in FIG. 13 still show a significant improvement.

[0104] The generic principles defined herein may be applied to other implementations without departing from the scope of the disclosure. Thus, the scope of the claims is not intended to be limited to the implementations shown herein but is to be accorded the widest scope consistent with this disclosure, the principles and novel features disclosed herein.

[0105] Various aspects of the invention may be apparent from the following enumerated example embodiments (EEE). (EEE1) A computer-implemented method of speech processing, comprising: receiving an input audio signal by a neural network implemented by a control system including one or more processors and one or more non-transitory storage media; generating an encoded speech signal by said neural network and based on said input speech signal; decoding, by the control system, the encoded speech signal to generate a decoded training speech signal; receiving, by a loss function generation module implemented by the control system, the decoded speech signal and a ground truth speech signal; generating, by a loss function generation module, a loss function value corresponding to the decoded speech signal, wherein generating the loss function value comprises applying a psychoacoustic model; training the neural network based on the loss function value, wherein the training includes updating at least one weight of the neural network. (EEE2) The method described in EEE1, wherein the neural network includes backpropagation based on the loss function value. (EEE3) The method described in EEE1 or EEE2, wherein the neural network includes an autoencoder. (EEE4) A method according to any one of EEE1 to EEE3, wherein the step of training the neural network includes a step of varying the physical state of at least one non-transitory storage medium location corresponding to at least one weight of the neural network. (EEE5) The method according to any one of EEE1 to EEE4, wherein a first part of the neural network generates the encoded audio signal and a second part of the neural network decodes the encoded audio signal. (EEE6) The method according to EEE5, wherein the first portion of the neural network includes an input neuron layer and a plurality of hidden neuron layers, and the input neuron layer includes more neurons than the final hidden neuron layer. (EEE7) The method according to EEE5, wherein at least some neurons of the first part of the neural network are configured with a rectified linear unit (ReLU) activation function. (EEE8) The method described in EEE5, wherein at least some neurons in the hidden layer of the second portion of the neural network are configured with a rectified linear unit (ReLU) activation function and at least some neurons in the output layer of the second portion are configured with a sigmoid activation function. (EEE9) The method of any one of EEE1-8, wherein the psychoacoustic model is based at least in part on one or more psychoacoustic masking thresholds. (EEE10) The psychoacoustic model is: Modeling the ear transfer function, grouping into critical bands, Frequency domain masking, including level-dependent diffusion but not limitation Modeling frequency-dependent hearing thresholds, or noise-to-mask ratio calculation, The method of any one of EEE1-9, comprising one or more of: (EEE11) The method according to any one of EEE1 to 10, wherein the loss function comprises calculating an average noise-to-mask ratio, and wherein the training comprises minimizing the average noise-to-mask ratio. (EEE12) A method of encoding speech, comprising: receiving a currently input audio signal by a control system including one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement a speech encoder including a neural network trained according to any one of the methods described in EEE1-11; encoding, by the audio encoder, the currently input audio signal into a compressed audio format; and outputting the encoded audio signal in the compressed audio format. (EEE13) A method of decoding audio, comprising the steps of: receiving a currently input compressed audio signal by a control system including one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement an audio decoder including a neural network trained according to any one of the methods described in EEE1-11; decoding, by said audio decoder, said currently input compressed audio signal; and outputting the decoded speech signal. (EEE14) The method of EEE13 further comprising the step of reproducing the decoded audio signal by one or more transducers. (EEE15) Equipment, an interface system; a control system including one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement a method according to any one of EEE1-14; Equipment including. (EEE16) One or more non-transitory media storing software, the software including instructions for controlling one or more devices to perform the method of any one of EEE1 to EEE14. (EEE17) Audio coding equipment, an interface system; a control system comprising one or more processors and one or more non-transitory storage media operatively coupled to the one or more processors, the control system configured to implement a speech encoder, the speech encoder comprising a neural network trained according to a method of any one of EEE1-11; The control system includes: Receives the currently input audio signal, encoding the currently input audio signal into a compressed audio format; outputting an encoded audio signal in the compressed audio format; The equipment is configured to: (EEE18) Audio coding equipment, an interface system; a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors; The control system is configured to implement a speech encoder, the speech encoder including a neural network trained according to a process, the process comprising: receiving an input training speech signal by said neural network and via said interface system; generating an encoded training speech signal by the neural network and based on the input training speech signal; decoding, by the control system, the encoded training speech signal to generate a decoded training speech signal; receiving, by a loss function generation module implemented by the control system, the decoded training speech signal and a ground truth speech signal; generating, by the loss function generation module, a loss function value corresponding to the decoded speech signal, wherein generating the loss function value includes applying a psychoacoustic model; training the neural network based on the loss function values; wherein the speech encoder comprises: Encoding the current input audio signal into a compressed audio format; The device is further configured to output an encoded audio signal in the compressed audio format. (EEE19) A system including an audio decoding device, an interface system; a control system including one or more processors and one or more non-transitory storage media operably coupled to the one or more processors; The control system is configured to implement an audio decoder, the audio decoder including a neural network trained according to a process, the process comprising: receiving an input training speech signal by said neural network and via said interface system; generating an encoded training speech signal by the neural network and based on the input training speech signal; decoding, by the control system, the encoded training speech signal to generate a decoded training speech signal; receiving, by a loss function generation module implemented by the control system, the decoded training speech signal and a ground truth speech signal; generating, by the loss function generation module, loss function values ​​corresponding to the decoded training speech signal, wherein generating the loss function values ​​comprises applying a psychoacoustic model; training the neural network based on the loss function values; the audio decoder includes: receiving a currently input encoded audio signal in a compressed audio format; decoding the currently input encoded audio signal into an uncompressed audio format; The device is further configured to output the decoded audio signal in the uncompressed audio format. (EEE20) The system described in EEE19, wherein the system further comprises one or more transducers configured to reproduce the decoded audio signal. [Explanation of symbols]

[0106] 205 Input Audio Signal 215 output audio signal 225 Loss function generation module 230 Loss function value 305 Encoder part 310 Decoder part 315 Optimization Module

Claims

[Claim 1] 1. A method for encoding an audio signal using a neural network, the method comprising: receiving an audio signal; encoding the audio signal with the neural network; A method comprising:

Citation Information

Patent Citations

  • Information processing device and information processing method

    WO2017168870A1

  • Apparatus and method for encoding an audio signal using a compensation value

    WO2018036972A1