Audio processing method and apparatus using a neural network

By transforming audio signals into the perceptual domain and using neural networks with psychoacoustic models and non-perceptual loss functions, the challenges of encoding and decoding general audio signals are addressed, resulting in efficient and high-performance audio processing.

JP7897228B2Active Publication Date: 2026-07-29DOLBY LABORATORIES LICENSING CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
DOLBY LABORATORIES LICENSING CORP
Filing Date
2021-10-14
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing neural networks face challenges in encoding and decoding general audio signals, particularly in the perceptual domain, due to the complexity of exploiting human auditory system limitations and high dynamic ranges in audio signals, which complicates training with non-perceptual loss functions.

Method used

Transforming audio signals into the perceptual domain using psychoacoustic models to minimize audible noise, allowing the use of non-perceptual loss functions like L1 and L2 for training, and employing first and second neural networks for encoding and decoding with conditional masking thresholds.

Benefits of technology

The method significantly reduces dynamic range and enables effective training and processing of audio signals, achieving high-performance encoding and decoding by leveraging the human auditory system's limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897228000004
    Figure 0007897228000004
  • Figure 0007897228000005
    Figure 0007897228000005
  • Figure 0007897228000006
    Figure 0007897228000006
Patent Text Reader

Abstract

This application describes a method for processing an audio signal using a neural network or using first and second neural networks. It also describes a method for training the neural network or for jointly training a set of the first and second neural networks. It also describes a method for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, and a method for deriving an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority from the following priority applications: U.S. Provisional Application No. 63 / 092,118, filed October 15, 2020, and European Patent Application No. 20210968.2, filed December 1, 2020. These are hereby incorporated by reference.

[0002] Technique Broadly speaking, the present disclosure relates to a method of processing an audio signal using a neural network or using first and second neural networks, and in particular to a method of processing an audio signal in a perceptual domain using a neural network or using first and second neural networks. The present disclosure further relates to a method of training the neural network or jointly training a set of the first and second neural networks. Further, the present disclosure relates to a method of obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, and a method of obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network. The present disclosure also relates to respective devices and computer program products.

[0003] In this document, several embodiments will be described with specific reference to their disclosure, but it will be understood that the present disclosure is not limited to such fields of use and is applicable in a broader context.

Background Art

[0004] Any discussion of background art throughout the disclosure should in no way be construed as an admission that such technology is widely known in the art or part of the common general knowledge in the art.

[0005] High-performance audio encoders and decoders exploit the limitations of the human auditory system to remove irrelevant information that humans cannot hear. Typically, encoding systems use psychoacoustic or perceptual models to calculate their respective masking thresholds. These masking thresholds are then used to control the encoding process so that the impact of introduced noise on hearing is minimized. [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] To date, neural networks have shown promise in many applications, including encoding and / or decoding images, videos, and even speech. However, there is still an existing need for the application of neural networks in general audio encoding and / or audio decoding applications using common training techniques, particularly in encoding and / or decoding applications involving perceptual domain audio signals. [Means for solving the problem]

[0007] According to a first aspect of this disclosure, a method for processing an audio signal using a neural network is provided. This method may include (a) obtaining a perceptual domain audio signal. This method may further include (b) inputting the perceptual domain audio signal into a neural network for processing the perceptual domain audio signal. This method may further include (c) obtaining a processed perceptual domain audio signal as an output from the neural network. This method may then include (d) converting the processed perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model.

[0008] In some embodiments, processing perceptual domain audio signals by a neural network may be performed in the time domain.

[0009] In some embodiments, this method may further include converting the audio signal to the frequency domain before step (d).

[0010] In some embodiments, the neural network may be conditional on information indicating a mask.

[0011] In some embodiments, the neural network may be conditional on perceptual domain audio signals.

[0012] In some embodiments, processing perceptual domain audio signals by a neural network may include predicting the processed perceptual domain audio signals across time.

[0013] In some embodiments, processing perceptual domain audio signals by a neural network may include predicting the processed perceptual domain audio signals across frequencies.

[0014] In some embodiments, processing perceptual domain audio signals by a neural network may include predicting the processed perceptual domain audio signals across time and frequency.

[0015] In some embodiments, the perceptual domain audio signal may be obtained by (a) converting the audio signal from the original signal domain to the perceptual domain by applying a mask; (b) encoding the perceptual domain audio signal; and (c) decoding the perceptual domain audio signal.

[0016] In some embodiments, quantization may be applied to the perceptual domain audio signal before encoding, and inverse quantization may be applied to the perceptual domain audio signal after decoding.

[0017] A second aspect of this disclosure provides a method for processing an audio signal using first and second neural networks. This method may include the step of (a) obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain using a first device. This method may further include the step of (b) inputting the perceptual domain audio signal into a first neural network to map the perceptual domain audio signal to a latent feature space representation. This method may further include the step of (c) obtaining a latent feature space representation as an output from the first neural network. This method may further include the step of (d) transmitting the latent feature space representation and mask of the perceptual domain audio signal to a second device. This method may further include the step of (e) receiving the latent feature space representation and mask of the perceptual domain audio signal by the second device. This method may further include the step of (f) inputting the latent feature space representation into a second neural network to generate an approximated perceptual domain audio signal. This method may further include the step of (g) obtaining an approximated perceptual domain audio signal as the output from a second neural network. This method may also include the step of (h) converting the approximated perceptual domain audio signal back to the original signal domain based on a mask.

[0018] In some embodiments, the method may further include encoding a latent feature space representation and mask of a perceptual domain audio signal into a bitstream and transmitting the bitstream to a second device, the method may further include the second device receiving the bitstream and decoding the bitstream to obtain a latent feature space representation and mask of a perceptual domain audio signal.

[0019] In some embodiments, the latent feature space representation and mask of the perceptual domain audio signal may be quantized before encoding into a bitstream and dequantized before processing by a second neural network.

[0020] In some embodiments, the second neural network may be conditional on the latent feature space representation and / or mask of the perceptual domain audio signal.

[0021] In some embodiments, the mapping of perceptual domain audio signals to latent feature space representations by a first neural network and the generation of approximate perceptual domain audio signals by a second neural network may be performed in the time domain.

[0022] In some embodiments, obtaining a perceptual domain signal in step (a) and transforming the approximated perceptual domain signal in step (h) may be performed in the frequency domain.

[0023] A third aspect of this disclosure provides a method for jointly training a set of first and second neural networks. This method may include (a) inputting the perceptual domain audio training signal into a first neural network to map the perceptual domain audio training signal to a latent feature space representation. This method may further include (b) obtaining the latent feature space representation of the perceptual domain audio training signal as an output from the first neural network. This method may further include (c) inputting the latent feature space representation of the perceptual domain audio training signal into a second neural network to generate an approximated perceptual domain audio training signal. This method may further include (d) obtaining the approximated perceptual domain audio training signal as an output from the second neural network. This method may also include (e) sequentially and iteratively adjusting the parameters of the first and second neural networks based on the difference between the approximated perceptual domain audio training signal and the original perceptual domain audio signal.

[0024] In some embodiments, the first and second neural networks may be trained in the perceptual domain based on one or more loss functions.

[0025] In some embodiments, the first and second neural networks may be trained in the perceptual domain based on the negative log-likelihood condition.

[0026] According to a fourth aspect of the present disclosure, a method for training a neural network is provided. The method may include (a) inputting a perceptual domain audio training signal into a neural network to process the perceptual domain audio training signal. The method may further include (b) obtaining the processed perceptual domain audio training signal as an output from the neural network. And the method may include (c) sequentially and iteratively adjusting the parameters of the neural network based on the difference between the processed perceptual domain audio training signal and the original perceptual domain audio signal.

[0027] In some embodiments, the neural network may be trained in the perceptual domain based on one or more loss functions.

[0028] In some embodiments, the neural network may be trained in the perceptual domain based on the negative log-likelihood condition.

[0029] A fifth aspect of this disclosure provides a method for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network. This method may include the step of (a) obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain. This method may further include the step of (b) inputting the perceptual domain audio signal into a neural network to map the perceptual domain audio signal to a latent feature space representation. This method may further include the step of (c) obtaining the latent feature space representation of the perceptual domain audio signal as an output from the neural network. This method may then include the step of (d) outputting the latent feature space representation of the perceptual domain audio signal as a bitstream.

[0030] In some embodiments, further information indicating the mask may be output as the bitstream in step (d).

[0031] In some embodiments, information indicating the latent feature space representation and / or mask of the perceptual domain audio signal may be quantized before being output as the bitstream.

[0032] In some embodiments, mapping perceptual domain audio signals to latent feature space representations using a neural network may be performed in the time domain.

[0033] In some embodiments, the acquisition of the perceptual domain audio signal may be performed in the frequency domain.

[0034] A sixth aspect of this disclosure provides a method for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network. This method may include (a) receiving the latent feature space representation of the perceptual domain audio signal as a bitstream. This method may further include (b) inputting the latent feature space representation into a neural network to generate a perceptual domain audio signal. This method may further include (c) obtaining the perceptual domain audio signal as the output from the neural network. This method may then include (d) converting the perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model.

[0035] In some embodiments, the neural network may be conditional on a latent feature space representation of a perceptual domain audio signal.

[0036] In some embodiments, in step (a), further information indicating the mask may be received as the bitstream, and the neural network may be conditional on this information.

[0037] In some embodiments, information indicating the latent feature space representation and / or mask of the perceptual domain audio signal may be received quantized, and dequantization may be performed before step (b).

[0038] In some embodiments, generating perceptual domain audio signals by a neural network may be performed in the time domain.

[0039] In some embodiments, the conversion of the perceptual domain audio signal to the original signal domain may be performed in the frequency domain.

[0040] A seventh aspect of this disclosure provides an apparatus for processing audio signals using a neural network. The apparatus may include a neural network and one or more processors, the processors configured to perform a method including: (a) obtaining a perceptual domain audio signal; (b) inputting the perceptual domain audio signal into a neural network for processing the perceptual domain audio signal; (c) obtaining a processed perceptual domain audio signal as an output from the neural network; and (d) converting the processed perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model.

[0041] According to the eighth aspect of this disclosure, an apparatus is provided for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network. The apparatus may include a neural network and one or more processors, the processors configured to perform a method including: (a) obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain; (b) inputting the perceptual domain audio signal into a neural network to map the perceptual domain audio signal to a latent feature space representation; (c) obtaining a latent feature space representation of the perceptual domain audio signal as an output from the neural network; and (d) outputting the latent feature space representation of the perceptual domain audio signal as a bitstream.

[0042] According to a ninth aspect of the present disclosure, an apparatus is provided for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network. The apparatus may include a neural network and one or more processors, the processors configured to perform a method including: (a) receiving a latent feature space representation of a perceptual domain audio signal as a bitstream; (b) inputting the latent feature space representation into a neural network to generate a perceptual domain audio signal; (c) obtaining the perceptual domain audio signal as output from a second neural network; and (d) converting the perceptual domain audio signal to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model.

[0043] According to aspects 10 through 15 of this disclosure, a computer program product is provided having a computer-readable storage medium having instructions adapted to cause a device to perform the method described herein when executed by such a device. [Brief explanation of the drawing]

[0044] Herein, with reference to the attached drawings, exemplary embodiments of the present disclosure will be described simply as examples.

[0045] [Figure 1] This example demonstrates how to process audio signals using a neural network.

[0046] [Figure 2] Further examples of how to process audio signals using neural networks are shown.

[0047] [Figure 3] This example shows a system that includes a device that processes audio signals using a neural network.

[0048] [Figure 4a]This example demonstrates how to process audio signals using the first and second neural networks. [Figure 4b] This example demonstrates how to process audio signals using the first and second neural networks.

[0049] [Figure 5] This document provides examples of systems for a device that uses a neural network to acquire and transmit a latent feature space representation of a perceptual domain audio signal, and a device that uses a neural network to acquire an audio signal from the latent feature space representation of a perceptual domain audio signal.

[0050] [Figure 6] This shows an example of how to train a neural network.

[0051] [Figure 7] This example shows how to jointly train a set of first and second neural networks.

[0052] [Figure 8] This shows an example of the original audio signal and mask as a function of level and frequency.

[0053] [Figure 9] This example shows a perceptual domain audio signal as a function of level and frequency, obtained by applying a mask to the original audio signal.

[0054] [Figure 10] This example demonstrates how to convert an audio signal into a perceptual domain and process the audio signal using a neural network.

[0055] [Figure 11]This diagram shows an example of an audio encoder and decoder operating in the perceptual domain, where a neural network is present in both the audio encoder and decoder. Because the network operates in the perceptual domain, this diagram also shows an example of using a simple loss function for training the neural network.

[0056] [Figure 12] This diagram shows an example of an audio encoder and decoder operating in the perceptual domain, where the neural network is located within the decoder. Because the network operates in the perceptual domain, this example demonstrates the use of a simple loss function for training the neural network. [Modes for carrying out the invention]

[0057] Overview While neural networks have shown promise for encoding and / or decoding images, videos, and even speech, encoding and / or decoding general audio using neural networks remains challenging. Two factors contribute to the complexity of general audio compression with neural networks: firstly, audio encoders and decoders must exploit the limitations of the human auditory system to achieve high performance. To exploit the perceptual limitations of the human auditory system, neural networks cannot be directly trained using non-perceptual loss functions such as L1 or L2 described below.

number

[0058] This disclosure describes methods and apparatus for transforming audio signals into the perceptual domain in each audio encoder and / or decoder before applying a neural network. This perceptual domain transformation of audio signals not only significantly reduces the dynamic range but also allows for the use of non-perceptual loss functions such as L1 and L2 for training the network.

[0059] A method for processing audio signals using a neural network.

[0060] Referring to the example in Figure 1, a method for processing an audio signal using a neural network is shown. In step S101, a perceptual domain audio signal is obtained. The term perceptual domain used here refers to a signal in which the relative level differences between frequency components are (approximately) proportional to their relative subjective importance. Generally, audio signals converted to the perceptual domain minimize the audible effect of adding white noise (spectrally flat noise) to the perceptual domain signal. This is because the noise is shaped to minimize audibility when the signal is converted back to the original signal domain.

[0061] Referring to the example in Figure 2, the perceptual domain audio signal may be obtained from steps S101a, S101b, and S101c, in which step S101a the audio signal can be transformed from the original signal domain to the perceptual domain by applying a mask.

[0062] One way to convert an audio signal to the perceptual domain is, for example, to estimate a mask or masking curve using a psychoacoustic model. A masking curve generally defines the minimum discernible difference (JND) level that the human auditory system can detect for a given stimulus signal. Once a masking curve is derived from the psychoacoustic model, the spectrum of the audio signal can be divided by the masking curve to generate a perceptual domain audio signal. The perceptual domain audio signal derived from multiplication by the inverse mask estimate may be converted back to the original signal by multiplying by the mask after encoding and / or decoding by a neural network. Multiplication by the mask after decoding ensures that errors introduced by the encoding and decoding processes follow the masking curve. It should be noted that this is one way to convert the original audio signal to the perceptual domain, but many other methods are possible, such as filtering in the time domain with a well-designed time-varying filter. See the examples in Figures 8 and 9 for illustrations of the conversion of the spectrum of the original audio signal to the perceptual domain. The plot in Figure 8 shows the spectrum of the original audio signal (solid line) and the estimated mask or masking curve calculated by the psychoacoustic model (dashed line). The perceptual region signal resulting from the multiplication of the inverse mask estimates is shown in the plot in Figure 9. The perceptual region signal not only allows for the use of a simple loss term during neural network training, but also exhibits a much smaller dynamic range than the original audio signal spectrum, as shown in Figure 8.

[0063] Referring again to the example in Figure 2, in step S101b, the perceptual domain audio signal is then encoded, and subsequently decoded in step S101c to obtain the perceptual domain audio signal. In some embodiments, quantization may be applied to the perceptual domain audio signal before encoding, or inverse quantization may be applied to the perceptual domain audio signal after decoding.

[0064] Referring again to the example in Figure 1, in step S102, the perceptual domain audio signal is input to the neural network for processing. The neural network used is not limited and can be selected according to the processing requirements. The neural network may operate in the frequency domain as well as the time domain, but in some embodiments, the processing of the perceptual domain audio signal by the neural network may be performed in the time domain. Furthermore, in some embodiments, the neural network may be conditioned with information indicating a mask. Alternatively or additionally, in some embodiments, the neural network may be conditioned with the perceptual domain audio signal.

[0065] In some embodiments, processing a perceptual domain audio signal by a neural network may include predicting the processed perceptual domain audio signal across time. Alternatively, in some embodiments, processing a perceptual domain audio signal by a neural network may include predicting the processed perceptual domain audio signal across frequency. Furthermore, alternatively, in some embodiments, processing a perceptual domain audio signal by a neural network may include predicting the processed perceptual domain audio signal across time and frequency.

[0066] Next, in step S103, the processed perceptual domain audio signal is obtained as the output from the neural network. In some embodiments, the processed perceptual domain audio signal may be converted to the frequency domain before the next step S104.

[0067] In step S104, the processed perceptual domain audio signal is transformed back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model. For example, to compute the mask, the psychoacoustic model may utilize frequency coefficients from a time-to-frequency transformation applied to transform the processed perceptual domain audio signal into the frequency domain. Alternatively or additionally, the mask used in step S104 may be based on the mask used to transform the original audio signal into the perceptual domain. In this case, the mask may be obtained as side information. The mask may optionally be quantized.

[0068] Thus, the term "original audio signal" used in this paper refers to each signal region of the audio signal before it is converted into the perceptual domain.

[0069] The above method can be implemented in various ways. For example, the method may be implemented by a device that processes audio signals using a neural network, which includes a neural network and one or more processors configured to perform the method.

[0070] Referring to the example in Figure 3, a system is shown that includes a device that processes audio signals using a neural network. This device may also be a decoder. In this case, the neural network is used only in the decoder.

[0071] As shown in the example in Figure 3, the perceptual domain audio signal may be quantized by the quantizer 101, or it may be (entropy) encoded by, for example, each legacy encoder 102. The quantized and encoded perceptual audio signal may then be sent to the decoder 103 as, for example, a bitstream. This is to obtain the quantized perceptual domain audio signal by, for example, (entropy) decoding the received bitstream. The quantized perceptual domain audio signal may then be inversely quantized by the inverse quantizer 104. The resulting perceptual domain audio signal is then input to a neural network (decoder-neural network) 105, and the processed perceptual domain audio signal can be obtained as the output from the neural network 105.

[0072] Alternatively or additionally, the above method may be implemented by a computer program product including a computer-readable storage medium having instructions adapted to cause a device to perform the method when executed by a device having processing capabilities.

[0073] How to process audio signals using the first and second neural networks

[0074] Referencing the examples in Figures 4a and 4b, a method for processing an audio signal using first and second neural networks is shown. For example, the first neural network may be implemented at the encoder site, and the second neural network may be implemented at the decoder site.

[0075] As shown in the example in Figure 4a, in step S201, the first device obtains a perceptual domain audio signal by applying a mask representing a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain. The first device may be, for example, an encoder. In some embodiments, obtaining the perceptual domain audio signal may be performed in the frequency domain.

[0076] In step S202, the acquired perceptual domain audio signal is then input to the first neural network to map the perceptual domain audio signal to a latent feature space representation.

[0077] In some embodiments, mapping perceptual domain audio signals to latent feature space representations by a first neural network may be performed in the time domain.

[0078] As output from the first neural network, in step S203, a latent feature space representation is obtained.

[0079] Next, in step S204, the latent feature space representation and mask of the perceptual domain audio signal are transmitted to a second device. In some embodiments, the above method may further include encoding the latent feature space representation and mask of the perceptual domain audio signal into a bitstream and transmitting the bitstream to the second device. In some embodiments, the latent feature space representation and mask of the perceptual domain audio signal may be further quantized before being encoded into a bitstream.

[0080] Referring here to the example in Figure 4b, in step S205, the latent feature space representation and mask of the perceptual domain audio signal are received by a second device. The second device may be, for example, a decoder. In some embodiments, this method may further include the second device receiving the latent feature space representation and mask of the perceptual domain audio signal as a bitstream and decoding the bitstream to obtain the latent feature space representation and mask of the perceptual domain audio signal. In some embodiments, if the latent feature space representation and mask of the perceptual domain audio signal are quantized, the latent feature space representation and mask of the perceptual domain audio signal may be dequantized before processing by the second neural network.

[0081] In step S206, the latent feature space representation is input to a second neural network to generate an approximated perceptual domain audio signal. In some embodiments, the second neural network may be conditional on the latent feature space representation and / or mask of the perceptual domain audio signal. In some embodiments, the generation of the approximated perceptual domain audio signal by the second neural network may be performed in the time domain.

[0082] In step S207, an approximated perceptual domain audio signal is obtained as the output from the second neural network.

[0083] The approximated perceptual domain audio signal is converted back to the original signal domain based on the mask in step S208. In some embodiments, the conversion of the approximated perceptual domain signal may be performed in the frequency domain.

[0084] The above methods may be implemented by systems of the respective first and second devices. Alternatively or additionally, the above methods may be implemented by computer program products, which include a computer-readable storage medium having instructions adapted to cause a device to perform the methods when executed by a device with processing capabilities.

[0085] Alternatively, the above method may be implemented in part by a device that uses a neural network to acquire and transmit a latent feature space representation of a perceptual domain audio signal, and in part by a device that uses a neural network to acquire an audio signal from the latent feature space representation of a perceptual domain audio signal. In this case, these devices may be implemented as individual devices or as a single system.

[0086] Next, a method for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network includes the following steps: In step (a), a perceptual domain audio signal is obtained by applying a mask representing a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain. In some embodiments, obtaining the perceptual domain audio signal may be performed in the frequency domain.

[0087] In step (b), the perceptual domain audio signal is input to a neural network in order to map the perceptual domain audio signal to a latent feature space representation. In some embodiments, the mapping of the perceptual domain audio signal to the latent feature space representation by the neural network may be performed in the time domain.

[0088] As output from the neural network, in step (c), a latent feature space representation of the perceptual domain audio signal is obtained. Then, in step (d), the latent feature space representation of the perceptual domain audio signal is output as a bitstream.

[0089] In some embodiments, further information indicating the mask may be output as the bitstream in step (d). In some embodiments, information indicating the latent feature space representation and / or mask of the perceptual domain audio signal may be quantized before being output as a bitstream.

[0090] The method for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network involves the following steps: Step (a) is received as a bitstream from the latent feature space representation of the perceptual domain audio signal. Step (b) is input to the neural network to generate the perceptual domain audio signal. Step (c) is obtained as the output of the neural network from the perceptual domain audio signal. Step (d) is converted back to the original signal domain based on a mask that represents a masking threshold derived from a psychoacoustic model.

[0091] In some embodiments, the neural network may be conditional on a latent feature space representation of the perceptual domain audio signal. In some embodiments, further, in step (a), mask-indicating information may be received as a bitstream, and the neural network may be conditional on this information. In some embodiments, the latent feature space representation and / or mask-indicating information of the perceptual domain audio signal may be received quantized, and dequantization may be performed before step (b). In some embodiments, the generation of the perceptual domain audio signal by the neural network may be performed in the time domain. In some embodiments, the conversion of the perceptual domain audio signal to the original signal domain may be performed in the frequency domain.

[0092] Referring to the example in Figure 5, a system is shown consisting of a device that uses a neural network to acquire and transmit a latent feature space representation of a perceptual domain audio signal (also known as the first device), and a device that uses a neural network to acquire an audio signal from the latent feature space representation of a perceptual domain audio signal (also known as the second device).

[0093] In the example of Figure 5, the perceptual domain audio signal may be input to the (first) neural network 202 in the (first) device 201 for processing as described above. The first neural network 202 may be an encoder-neuronal network. The latent feature space representation output from the (first) neural network may be quantized by the quantizer 203 and transmitted to the (second) device 204. The quantized latent feature space representation may be encoded as a bitstream and transmitted to the (second) device 204. In the (second) device 204, the received latent feature space representation may first be dequantized by the inverse quantizer 205 and optionally decoded before being input to the (second) neural network 206 to generate a perceptual domain audio signal approximated based on the latent feature space representation. The approximated perceptual domain audio signal may then be obtained as the output from the (second) neural network 206.

[0094] How to train a neural network

[0095] Referring to the example in Figure 6, a method for training a neural network is shown. In step S301, the perceptual domain audio training signal is input to the neural network for processing. The perceptual domain audio training signal is processed by the neural network, and in step S302, the processed perceptual domain audio training signal is then obtained as the output from the neural network. Based on the difference between the processed perceptual domain audio training signal and the original perceptual domain audio signal from which the perceptual domain audio training signal may have been obtained, the parameters of the neural network are then adjusted sequentially and iteratively in step S303. Based on this sequential and iterative adjustment, the neural network is trained to produce increasingly better processed perceptual domain audio training signals. The purpose of this sequential and iterative adjustment is to cause the neural network to produce processed perceptual domain audio training signals that are indistinguishable from each of the original perceptual domain audio signals.

[0096] In some embodiments, the neural network may be trained in the perceptual domain based on one or more loss functions. A neural network designed to encode an audio signal in the perceptual domain may be trained with simple loss functions such as L1 or L2, because these can introduce spectral white errors. In the case of L1 and L2, the neural network can predict the mean of the processed perceptual domain audio training signal.

[0097] Alternatively, in some embodiments, the neural network may be trained in the perceptual domain based on a negative log likelihood (NLL) condition. In the case of NLL, the neural network may predict the mean and scale as parameterizations from a pre-selected distribution. To avoid numerical instability, a logarithmic operation of the scale parameter is typically used. The pre-selected distribution may be the Laplacian. Alternatively, the pre-selected distribution may be a logistic or Gaussian distribution. In the case of a Gaussian distribution, the scale parameter may be replaced by a variance parameter. For the case of NLL, a sampling operation may be used to convert the distribution parameters into a processed perceptual domain audio training signal. The sampling operation can be written as follows:

number

[0098] For example, in the case of the Laplacian,

number

[0099] A method for jointly training a set of two neural networks, the first and the second.

[0100] Referring to the example in Figure 7, a method for jointly training the first and second sets of neural networks is shown.

[0101] In step S401, the perceptual domain audio training signal is input to the first neural network in order to map the perceptual domain audio training signal to a latent feature space representation. In step S402, the latent feature space representation of the perceptual domain audio training signal is obtained as the output from the first neural network. In step S403, the latent feature space representation of the perceptual domain audio training signal is then input to the second neural network in order to generate an approximated perceptual domain audio training signal. Then, in step S404, the approximated perceptual domain audio training signal is obtained as the output from the second neural network. Finally, in step S405, the parameters of the first and second neural networks are iteratively adjusted based on the difference between the approximated perceptual domain audio training signal and the original perceptual domain audio signal from which the perceptual domain audio training signal was derived.

[0102] In some embodiments, the first and second neural networks may be trained in the perceptual domain based on one or more loss functions. In some embodiments, the first and second neural networks may be trained in the perceptual domain based on a negative log-likelihood (NLL) condition. The goal of the sequential iterative tuning is to have the first and second neural networks generate approximate perceptual domain audio training signals that are indistinguishable from their respective original perceptual domain audio signals.

[0103] Further exemplary embodiments Further exemplary embodiments of the methods and apparatus described in this paper are shown with reference to the examples in Figures 10 to 12. The example in Figure 10 shows a schematic diagram illustrating the conversion of an audio signal to the perceptual domain for data reduction using a neural network. In the example in Figure 10, PCM audio data is used as input.

[0104] Figure 11 shows a schematic diagram of an audio encoder and decoder operating in the perceptual domain, where both the encoder and decoder have neural networks. Figure 11 also shows that a simple loss function can be used to train the neural network because it operates in the perceptual domain. In the example in Figure 11, the ground truth signal points to the original perceptual domain audio signal, and each perceptual domain audio training signal may be derived based on this, which may be compared to an approximated perceptual domain audio signal to sequentially and iteratively tune the neural network.

[0105] The example in Figure 12 shows a schematic diagram of an audio encoder and decoder operating in the perceptual domain, with a neural network inside the decoder. Figure 12 also shows that because the network operates in the perceptual domain, a simple loss function is used to train the neural network. In this case, the ground truth signal refers to the original perceptual domain audio signal, and each perceptual domain audio training signal may be derived based on this, which may be compared to the processed perceptual domain audio signal to sequentially and iteratively tune the neural network.

[0106] interpretation

[0107] Unless otherwise specified, as will be apparent from the following discussion, any discussion throughout this disclosure using terms such as “processing,” “computing,” “calculating,” “determining,” and “analyzing” is understood to refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data, expressed as physical quantities, such as electronic quantities, into other data, also expressed as physical quantities.

[0108] Similarly, the term “processor” may refer to any device or part of a device that processes electronic data from, for example, registers and / or memory to make that electronic data into other electronic data that can be stored, for example, in registers and / or memory. “Computer” or “calculator” or “computing platform” may include one or more processors.

[0109] The methods described herein are executable by one or more processors that accept computer-readable (also called machine-readable) code, which includes a set of instructions that, when executed by one or more of the processors, perform at least one of the methods described herein. This includes any processor capable of executing a set of instructions (sequential or otherwise) specifying an action to be performed. Thus, one example is a typical processing system comprising one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem, which includes main RAM and / or static RAM and / or ROM. A bus subsystem for communication between components may also be included. The processing system may further be a distributed processing system having processors connected by a network. If the processing system requires a display, such a display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data entry is required, the processing system may also include one or more input devices, such as an alphanumeric input unit, such as a keyboard, and a pointing control device, such as a mouse. The processing system may also include a storage system, such as a disk drive unit. The processing system in some configurations may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software) which, when executed by one or more processors, causes one or more of the methods described herein to be performed. Note that, if a method includes several elements, for example, several steps, the ordering of such elements is not implied unless specifically stated.The software may reside on the hard disk, or, during its execution by the computer system, may reside entirely or at least partially in RAM and / or the processor. Thus, the memory and processor also constitute a computer-readable carrier medium carrying computer-readable code. Furthermore, the computer-readable carrier medium may form or be contained within a computer program product.

[0110] In alternative exemplary embodiments, the one or more processors may operate as standalone devices or, in a networked deployment, may be connected, for example, to another processor; the one or more processors may operate as a server or user machine in a server-user network environment; or as a peer machine in a peer-to-peer or distributed network environment. The one or more processors may form a personal computer (PC), a tablet PC, a personal digital assistant (PDA), a cellular telephone, a web appliance, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) specifying the actions to be taken by that machine.

[0111] It should be noted that the term “machine” is also interpreted to include any set of machines that individually or jointly execute a set (or set) of instructions for executing one or more of the methodologies discussed herein.

[0112] Therefore, an exemplary embodiment of each method described herein may take the form of a computer-readable carrier medium carrying a set of instructions, for example, a computer program for execution on one or more processors, for example, one or more processors that are part of a web server configuration. Thus, as those skilled in the art will understand, exemplary embodiments of the disclosure may be embodied as a method, an apparatus such as a special-purpose apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, for example, a computer program product. The computer-readable carrier medium carries computer-readable code that, when executed on one or more processors, causes one or more processors to perform the method. Thus, aspects of the disclosure may take the form of a method, an exemplary embodiment entirely of hardware, an exemplary embodiment entirely of software, or an exemplary embodiment combining software and hardware aspects. Furthermore, the disclosure may take the form of a carrier medium (for example, a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0113] The software may also be transmitted and received over a network via a network interface device. While the carrier medium is a single medium in exemplary embodiments, the term “carrier medium” should be understood to include a single or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more instruction sets. The term “carrier medium” should also be understood to include any medium capable of storing, encoding, or carrying a set of instructions for execution by one or more of the processors, causing one or more of the processors to execute one or more of the methods of this disclosure. The carrier medium can take many forms, but is not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media include dynamic memory such as main memory. Transmission media include coaxial cables, copper wires, and optical fibers, including wires that constitute a bus subsystem. Transmission media can also take the form of sound waves or light waves, such as those generated during radio and infrared data communications. For example, the term “carrier medium” should be understood to include, but not be limited to, computer products embodied in solid memory, optical and magnetic media; media carrying a propagation signal detectable by at least one or more processors and representing a set of instructions that implement a method at runtime; and transmission media in a network carrying a propagation signal detectable by at least one of the one or more processors and representing a set of instructions.

[0114] It will be understood that, in some exemplary embodiments, the steps of the method discussed are performed by a suitable processor(s) of a processing (e.g., a computer) system that executes instructions (computer-readable code) stored in a memory device. It will also be understood that this disclosure is not limited to any particular implementation or programming technique, and may be implemented using any suitable technique for implementing the functionality described herein. This disclosure is not limited to any particular programming language or operating system.

[0115] Throughout this disclosure, any reference to “one embodiment,” “several embodiments,” or “a particular exemplary embodiment” means that any specific feature, structure, or characteristic described in relation to that embodiment is included in at least one embodiment of this disclosure. Therefore, the phrases “in one embodiment,” “in several embodiments,” or “in a particular exemplary embodiment” in various parts of this disclosure do not necessarily all refer to the same exemplary embodiment. Furthermore, any specific feature, structure, or characteristic can be combined in any suitable way in one or more exemplary embodiments, as will be apparent to those skilled in the art from this disclosure.

[0116] Where used herein, unless otherwise specified, the use of ordinal adjectives such as “first,” “second,” “third,” etc., to describe a common object simply indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects described in this way must be in a given order, temporally, spatially, in rank, or in any other way.

[0117] In the claims and descriptions herein, any of the terms including, containing, or having are open terms that include at least the listed elements / features but do not exclude others. Therefore, as used in the claims, the terms including / having should not be interpreted as being limited to the enumerated means, elements, or steps. For example, an apparatus having A and B should not be limited to an apparatus consisting only of elements A and B. Any of the terms including, containing, or encompassing as used herein are open terms that include at least the enumerated elements / features but do not exclude others. Therefore, including is synonymous with having and means having.

[0118] In the above description of the exemplary embodiments of this disclosure, it should be understood that, for the purpose of improving the flow of the disclosure and aiding in the understanding of one or more of the various inventive aspects, various features of the disclosure may be summarized in a single exemplary embodiment, figure, or description thereof. However, this method of disclosure should not be interpreted as reflecting an intention that the claims require more features than are expressly described in each claim. Rather, as reflected in the following claims, the aspects of the invention are fewer than all the features of a single, aforementioned exemplary embodiment. Thus, the claims following this specification are hereby expressly incorporated herein, and each claim stands alone as a separate exemplary embodiment of the disclosure.

[0119] Furthermore, while some exemplary embodiments described herein may include some features included in other exemplary embodiments, they may not include others. However, combinations of features from different exemplary embodiments are intended to be within the scope of the disclosure and constitute different exemplary embodiments, as will be understood by those skilled in the art. For example, any of the exemplary embodiments described in the following claims may be used in any combination.

[0120] Numerous specific details are described in the descriptions provided herein. However, it is understood that exemplary embodiments of this disclosure may be carried out without these specific details. On the other hand, well-known methods, structures, and techniques are not described in detail so as not to obscure the understanding of this paper.

[0121] Therefore, while what is considered to be the best form of disclosure is described, a person skilled in the art will recognize that other further modifications may be made without departing from the spirit of the disclosure, and that all such changes and modifications are intended to be requested as being included within the scope of this disclosure. For example, any of the formulas described above merely represent possible procedures. Functions may be added or removed from the block diagram, and operations may be swapped between function blocks. Steps may be added or removed from the methods described within the scope of this disclosure.

[0122] Various aspects of the present invention can be understood from the following enumerated example embodiments (EEE). [EEE1] A computer-implemented method for processing audio signals using a neural network, wherein the method is: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal in a bitstream, method. [EEE2] The method of EEE1 wherein further information indicating the mask is output in the bitstream in step (d). [EEE3] The method of EEE1 or 2, wherein the information indicating the latent feature space representation and / or the mask of the perceptual domain audio signal is quantized before the step of outputting in the bitstream. [EEE4] The mapping of the perceptual domain audio signals to the latent feature space representation by the neural network is performed in the time domain, and / or Obtaining the aforementioned perceptual domain audio signal is performed in the frequency domain. The method described in any one of EEE1 to 3. [EEE5] A computer-implemented method for decoding an audio signal using a neural network, wherein the method is: (a) The step of obtaining a representation of the audio signal in the perceptual domain; (b) the step of inputting the representation of the perceptual domain audio signal into the neural network for processing the representation of the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) a step of converting the processed perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model, method. [EEE6] The processing of the perceptual domain audio signal by the neural network is performed in the time domain; and / or The method further includes converting the audio signal to the frequency domain before step (d), Methods for EEE5. [EEE7] The neural network is conditional on the information indicating the mask; and / or The aforementioned neural network is conditional on the audio signal in the perceptual domain. The method described in EEE5 or 6. [EEE8] Processing the perceptual domain audio signal using the aforementioned neural network is: Predicting the processed perceptual domain audio signal across time; Predicting the processed perceptual domain audio signal across frequencies; and Predicting the processed perceptual domain audio signal across time and frequency, A method of EEE7 that includes at least one of the following. [EEE9] The method according to any one of EEE5 to 8, wherein the representation of the perceptual domain audio signal includes the perceptual domain audio signal. [EEE10] The aforementioned representation of the perceptual domain audio signal is: By applying the aforementioned mask, the audio signal is converted from the original signal domain to the perceptual domain; Encode the aforementioned perceptual region audio signal; This was obtained by decoding the audio signal from the perceptual domain; Optionally, Quantization is applied to the perceptual domain audio signal before encoding, and inverse quantization is applied to the perceptual domain audio signal after decoding. The method described in any one of the EEE5 to 9. [EEE11] Step (a) includes receiving the latent feature space representation of the perceptual domain audio signal in a bitstream; Step (b) includes inputting the latent feature space representation into the neural network for generating the processed perceptual domain audio signal, Methods for EEE5. [EEE12] The neural network is provided that the latent feature space representation of the perceptual domain audio signal is the method according to EEE11. [EEE13] The neural network further includes receiving additional information indicating the mask as the bitstream, and the neural network, subject to the additional information, The method described in EEE11 or 12. [EEE14] The information indicating the latent feature spatial representation and / or mask of the perceptual domain audio signal is received in a quantized form; The method further includes inverse quantization before inputting the latent feature space representation into the neural network. The method described in any one of the EEE11 to 13. [EEE15] The generation of the perceptual domain audio signal by the neural network is performed in the time domain; and / or The conversion of the aforementioned perceptual domain audio signal to the original signal domain is performed in the frequency domain. The method described in any one of the EEE11 to 14. [EEE16] A method for processing audio signals using a neural network (for example, a computer implementation), the method being: (a) The step of obtaining a perceptual audio signal; (b) The step of inputting the perceptual domain audio signal into the neural network for processing the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) a step of converting the processed perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model, method. [EEE17] The method according to EEE16, wherein the processing of the perceptual domain audio signal by the neural network is performed in the time domain. [EEE18] The method according to EEE16 or 17, further comprising converting the audio signal to the frequency domain prior to step (d). [EEE19] The method according to any one of EEE16 to 18, wherein the neural network is conditional on the information indicating the mask. [EEE20] The method according to any one of EEE16 to 19, wherein the neural network is conditioned on the perceptual domain audio signal. [EEE21] The method according to EEE19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across time. [EEE22] The method according to EEE19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across frequencies. [EEE23] The method according to EEE19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across time and frequency. [EEE24] The aforementioned audio signal in the perceptual domain is: (a) By applying the mask, the audio signal is converted from the original signal domain to the perceptual domain; (b) Encode the audio signal in the perceptual region; (c) Obtained by decoding the audio signal in the perceptual domain, The method described in any one of the EEE16 to 23. [EEE25] The method according to EEE24, wherein quantization is applied to the perceptual domain audio signal before encoding, and inverse quantization is applied to the perceptual domain audio signal after decoding. [EEE26] A method for processing an audio signal using first and second neural networks (for example, a computer implementation), wherein the method is: (a) A step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the original audio signal in the signal domain using a first device; (b) The step of inputting the perceptual domain audio signal into the first neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation as the output from the first neural network; (d) The step of transmitting the latent feature space representation and the mask of the perceptual domain audio signal to a second device; (e) The second device receives the latent feature space representation and the mask of the perceptual domain audio signal; (f) The step of inputting the latent feature space representation into the second neural network for generating an approximated perceptual domain audio signal; (g) the step of obtaining the approximated perceptual domain audio signal as the output from the second neural network; (h) a step of converting the approximated perceptual domain audio signal to the original signal domain based on the mask, method. [EEE27] The method according to EEE26, further comprising encoding the latent feature space representation and the mask of the perceptual domain audio signal into a bitstream, and transmitting the bitstream to the second device, the method further comprising the second device receiving the bitstream and decoding the bitstream to obtain the latent feature space representation and the mask of the perceptual domain audio signal. [EEE28] The method according to EEE27, wherein the latent feature space representation and the mask of the perceptual domain audio signal are quantized before encoding into the bitstream and dequantized before processing by the second neural network. [EEE29] The method according to any one of EEE26 to 28, wherein the second neural network is conditioned on the latent feature space representation and / or the mask of the perceptual domain audio signal. [EEE30] The method according to any one of EEE26 to 29, wherein the mapping of the perceptual domain audio signal to the latent feature space representation by the first neural network and the generation of the approximated perceptual domain audio signal by the second neural network are performed in the time domain. [EEE31] The method according to any one of EEE26 to 30, wherein obtaining the perceptual domain signal in step (a) and transforming the approximated perceptual domain signal in step (h) are performed in the frequency domain. [EEE32] A method for jointly training a set of first and second neural networks (for example, a computer implementation), the method being: (a) the step of inputting a perceptual domain audio training signal into the first neural network for mapping the perceptual domain audio training signal to a latent feature space representation; (b) The step of obtaining the latent feature space representation of the perceptual domain audio training signal as the output from the first neural network; (c) The step of inputting the latent feature space representation of the perceptual domain audio training signal into the second neural network for generating an approximated perceptual domain audio training signal; (d) the step of obtaining the approximated perceptual domain audio training signal as the output from the second neural network; (e) The step of iteratively adjusting the parameters of the first and second neural networks based on the difference between the approximated perceptual domain audio training signal and the original perceptual domain audio signal, method. [EEE33] The method according to EEE32, wherein the first and second neural networks are trained in a perceptual domain based on one or more loss functions. [EEE34] The method according to EEE32, wherein the first and second neural networks are trained in the perceptual domain based on a negative log-likelihood condition. [EEE35] A method for training a neural network (for example, a computer implementation), the method being: (a) The step of inputting a perceptual domain audio training signal into the neural network for processing the perceptual domain audio training signal; (b) the step of obtaining the processed perceptual domain audio training signal as the output from the neural network; (c) The step of iteratively adjusting the parameters of the neural network based on the difference between the processed perceptual domain audio training signal and the original perceptual domain audio signal, method. [EEE36] The method according to EEE35, wherein the neural network is trained in a perceptual domain based on one or more loss functions. [EEE37] The neural network is trained in the perceptual domain based on a negative log-likelihood condition, according to the method of EEE35. [EEE38] A method (e.g., a computer implementation) for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, wherein the method is: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal as a bitstream, method. [EEE39] The method according to EEE38, wherein further information indicating the mask is output as the bitstream in step (d). [EEE40] The method according to EEE38 or 39, wherein the information indicating the latent feature space representation and / or the mask of the perceptual domain audio signal is quantized before being output as the bitstream. [EEE41] The method according to any one of EEE39 to 40, wherein the mapping of the perceptual domain audio signal to the latent feature space representation by the neural network is performed in the time domain. [EEE42] The method according to any one of EEE38 to 41, wherein obtaining the aforementioned perceptual domain audio signal is performed in the frequency domain. [EEE43] A method (e.g., a computer implementation) for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network, wherein the method is: (a) The step of receiving the latent feature space representation of the perceptual domain audio signal as a bitstream; (b) The step of inputting the latent feature space representation into a neural network for generating the perceptual domain audio signal; (c) The step of obtaining the perceptual domain audio signal as the output from the neural network; (d) The step of converting the perceptual domain audio signal to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model, method. [EEE44] The method according to EEE43, wherein the neural network is conditioned on the latent feature space representation of the perceptual domain audio signal. [EEE45] In step (a), further information indicating the mask is received as the bitstream, and the neural network is conditional on the information, according to the method of EEE43 or 44. [EEE46] The method according to any one of EEE 43 to 45, wherein the information indicating the latent feature spatial representation and / or the mask of the perceptual domain audio signal is received in a quantized form, and inverse quantization is performed before step (b). [EEE47] The method according to any one of EEE 43 to 46, wherein generating the perceptual domain audio signal by the neural network is performed in the time domain. [EEE48] The method according to any one of EEE43 to 47, wherein the conversion of the aforementioned perceptual domain audio signal to the original signal domain is performed in the frequency domain. [EEE49] An apparatus for processing audio signals using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of obtaining a perceptual audio signal; (b) The step of inputting the perceptual domain audio signal into the neural network for processing the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) a step of converting the processed perceptual domain audio signal back to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model, Device. [EEE50] An apparatus for acquiring and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal in the original signal domain; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal as a bitstream, Device. [EEE51] An apparatus for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of receiving the latent feature space representation of the perceptual domain audio signal as a bitstream; (b) The step of inputting the latent feature space representation into a neural network for generating the perceptual domain audio signal; (c) The step of obtaining the perceptual domain audio signal as the output from the second neural network; (d) The step of converting the perceptual domain audio signal to the original signal domain based on a mask indicating a masking threshold derived from a psychoacoustic model, Device. [EEE52] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of EEE1 to EEE10 when executed by a device having processing capabilities. [EEE53] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of the EEE11 to 16 when executed by a device having processing capabilities. [EEE54] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of the EEE17 to 19 when executed by a device having processing capabilities. [EEE55] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of EEE20 to 22 when executed by a device having processing capabilities. [EEE56] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of EEE23 to 27 when executed by a device having processing capabilities. [EEE57] A computer program product having a computer-readable storage medium having instructions adapted to cause a device to perform the method described in any one of EEE28 to 33 when executed by a device having processing capabilities.

Claims

1. A computer-implemented method for processing audio signals using a neural network, the method being: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal in a bitstream, method.

2. Furthermore, the method according to claim 1, wherein information indicating the mask is output in the bitstream at step (d).

3. The method of claim 2, wherein the information indicating the latent feature space representation and / or the mask of the perceptual domain audio signal is quantized before the step of outputting in the bitstream.

4. The mapping of the perceptual domain audio signals to the latent feature space representation by the neural network is performed in the time domain, and / or Obtaining the aforementioned perceptual domain audio signal is performed in the frequency domain. The method according to any one of claims 1 to 3.

5. A computer-implemented method for decoding an audio signal using a neural network, the method being: (a) A step of obtaining a representation of a perceptual domain audio signal by decoding a received bitstream, wherein the perceptual domain audio signal is a signal obtained by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal; (b) the step of inputting the representation of the perceptual domain audio signal into the neural network for processing the representation of the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) The step of converting the processed perceptual region audio signal to a signal region to which the mask is not applied, based on the mask, method.

6. The processing of the perceptual domain audio signal by the neural network is performed in the time domain; and / or The method further includes converting the audio signal to the frequency domain before step (d), The method according to claim 5.

7. The neural network is conditional on the information indicating the mask; and / or The aforementioned neural network is conditional on the aforementioned perceptual domain audio signal. The method according to claim 5 or 6.

8. Processing the perceptual domain audio signal using the aforementioned neural network is: Predicting the processed perceptual domain audio signal across time; Predicting the processed perceptual domain audio signal across frequencies; and Predicting the processed perceptual domain audio signal across time and frequency, The method according to claim 7, comprising at least one of the following.

9. The method according to any one of claims 5 to 8, wherein the representation of the perceptual domain audio signal includes the perceptual domain audio signal.

10. The aforementioned representation of the perceptual domain audio signal is: Encode the aforementioned perceptual region audio signal; This is obtained by decoding the aforementioned audio signal in the perceptual domain; Optionally, Quantization is applied to the perceptual domain audio signal before encoding, and inverse quantization is applied to the perceptual domain audio signal after decoding. The method according to any one of claims 5 to 9.

11. Step (a) includes receiving the latent feature space representation of the perceptual domain audio signal in a bitstream; Step (b) includes inputting the latent feature space representation into the neural network for generating the processed perceptual domain audio signal, The method according to claim 5.

12. The method according to claim 11, wherein the neural network is conditioned on the latent feature space representation of the perceptual domain audio signal.

13. The further includes receiving additional information indicating the mask as the bitstream, The neural network is subject to the additional information, The method according to claim 11 or 12.

14. The additional information indicating the latent feature space representation and / or mask of the perceptual domain audio signal is received in a quantized form; The method further includes inverse quantization before inputting the latent feature space representation into the neural network. The method according to claim 13.

15. The generation of the perceptual domain audio signal by the neural network is performed in the time domain; and / or Converting the aforementioned perceptual domain audio signal to a signal domain to which the mask is not applied is performed in the frequency domain. The method according to any one of claims 11 to 14.

16. A computer-implemented method for processing audio signals using a neural network, the method being: (a) The step of obtaining a perceptual domain audio signal; (b) The step of inputting the perceptual domain audio signal into the neural network for processing the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) The step of converting the processed perceptual region audio signal to a signal region to which the mask is not applied, based on a mask indicating a masking threshold derived from a psychoacoustic model, method.

17. The method according to claim 16, wherein the processing of the perceptual domain audio signal by the neural network is performed in the time domain.

18. The method according to claim 16 or 17, further comprising converting the audio signal to the frequency domain prior to step (d).

19. The method according to any one of claims 16 to 18, wherein the neural network is conditional on the information indicating the mask.

20. The method according to any one of claims 16 to 19, wherein the neural network is conditional on the perceptual domain audio signal.

21. The method according to claim 19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across time.

22. The method according to claim 19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across frequencies.

23. The method according to claim 19 or 20, wherein processing the perceptual domain audio signal by the neural network includes predicting the processed perceptual domain audio signal across time and frequency.

24. The aforementioned audio signal in the perceptual domain is: (a) By applying the mask, the audio signal is converted into a perceptual domain; (b) Encode the audio signal in the perceptual domain; (c) Obtained by decoding the audio signal in the perceptual domain, The method according to any one of claims 16 to 23.

25. The method according to claim 24, wherein quantization is applied to the perceptual domain audio signal before encoding, and inverse quantization is applied to the perceptual domain audio signal after decoding.

26. A computer-implemented method for processing an audio signal using first and second neural networks, the method comprising: (a) A step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal using a first device; (b) Inputting the perceptual domain audio signal into the first neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation as the output from the first neural network; (d) Transmitting the latent feature spatial representation of the perceptual domain audio signal and the mask to a second device; (e) The second device receives the latent feature space representation of the perceptual domain audio signal and the mask; (f) The step of inputting the latent feature space representation into the second neural network for generating an approximated perceptual domain audio signal; (g) the step of obtaining the approximated perceptual domain audio signal as the output from the second neural network; (h) The step of converting the approximated perceptual region audio signal to a signal region to which the mask is not applied, based on the mask, method.

27. The method according to claim 26, wherein the method further comprises encoding the latent feature space representation and the mask of the perceptual domain audio signal into a bitstream, and transmitting the bitstream to the second device, the method further comprises the second device receiving the bitstream, and decoding the bitstream to obtain the latent feature space representation and the mask of the perceptual domain audio signal.

28. The method according to claim 27, wherein the latent feature space representation and the mask of the perceptual domain audio signal are quantized before encoding into the bitstream and dequantized before processing by the second neural network.

29. The method according to any one of claims 26 to 28, wherein the second neural network is conditioned on the latent feature spatial representation and / or the mask of the perceptual domain audio signal.

30. The method according to any one of claims 26 to 29, wherein the mapping of the perceptual domain audio signal to the latent feature space representation by the first neural network and the generation of the approximated perceptual domain audio signal by the second neural network are performed in the time domain.

31. The method according to any one of claims 26 to 30, wherein obtaining the perceptual domain audio signal in step (a) and converting the approximated perceptual domain signal in step (h) are performed in the frequency domain.

32. A computer-implemented method for jointly training a set of first and second neural networks, the method being: (a) inputting a perceptual domain audio training signal into the first neural network for mapping the perceptual domain audio training signal to a latent feature space representation; (b) The step of obtaining the latent feature space representation of the perceptual domain audio training signal as the output from the first neural network; (c) The step of inputting the latent feature space representation of the perceptual domain audio training signal into the second neural network for generating an approximated perceptual domain audio training signal; (d) the step of obtaining the approximated perceptual domain audio training signal as the output from the second neural network; (e) The step of iteratively adjusting the parameters of the first and second neural networks based on the difference between the approximated perceptual domain audio training signal and the original perceptual domain audio signal, method.

33. The method according to claim 32, wherein the first and second neural networks are trained in a perceptual domain based on one or more loss functions.

34. The method according to claim 32, wherein the first and second neural networks are trained in a perceptual domain based on a negative log-likelihood condition.

35. A computer-implemented method for training a neural network, the method being: (a) A step of inputting a perceptual domain audio training signal into the neural network for processing the perceptual domain audio training signal, wherein the perceptual domain audio training signal is a signal obtained by applying a mask indicating a masking threshold derived from a psychoacoustic model to an audio signal, and processing the perceptual domain audio training signal includes mapping the perceptual domain audio training signal to a latent feature space representation; (b) the step of obtaining the processed perceptual domain audio training signal as the output from the neural network; (c) The step of iteratively adjusting the parameters of the neural network based on the difference between the processed perceptual domain audio training signal and the original perceptual domain audio signal, method.

36. The method according to claim 35, wherein the neural network is trained in a perceptual domain based on one or more loss functions.

37. The method according to claim 35, wherein the neural network is trained in a perceptual domain based on a negative log-likelihood condition.

38. A computer-implemented method for obtaining and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, wherein the method is: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal as a bitstream, method.

39. Furthermore, the method according to claim 38, wherein information indicating the mask is output as the bitstream in step (d).

40. The method according to claim 39, wherein the information indicating the latent feature space representation and / or the mask of the perceptual domain audio signal is quantized before being output as the bitstream.

41. The method according to any one of claims 39 to 40, wherein the mapping of the perceptual domain audio signal to the latent feature space representation by the neural network is performed in the time domain.

42. The method according to any one of claims 38 to 41, wherein obtaining the aforementioned perceptual domain audio signal is performed in the frequency domain.

43. A computer-implemented method for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network, wherein the method is: (a) The step of receiving the latent feature space representation of the perceptual domain audio signal as a bitstream; (b) The step of inputting the latent feature space representation into a neural network for generating the perceptual domain audio signal; (c) The step of obtaining the perceptual domain audio signal as the output from the neural network; (d) The step of converting the perceptual domain audio signal to a signal domain to which the mask is not applied, based on a mask indicating a masking threshold derived from a psychoacoustic model, method.

44. The method according to claim 43, wherein the neural network is conditional on the latent feature space representation of the perceptual domain audio signal.

45. The method according to claim 43 or 44, wherein in step (a), information indicating the mask is further received as the bitstream, and the neural network is conditional on the information.

46. The method according to claim 45, wherein the information indicating the latent feature spatial representation and / or the mask of the perceptual domain audio signal is received in a quantized form, and inverse quantization is performed before step (b).

47. The method according to any one of claims 43 to 46, wherein the generation of the perceptual domain audio signal by the neural network is performed in the time domain.

48. The method according to any one of claims 43 to 47, wherein the conversion of the perceptual domain audio signal to a signal domain to which the mask is not applied is performed in the frequency domain.

49. An apparatus for processing audio signals using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of obtaining a perceptual domain audio signal; (b) The step of inputting the perceptual domain audio signal into the neural network for processing the perceptual domain audio signal; (c) The step of obtaining a processed perceptual domain audio signal as the output from the neural network; (d) The step of converting the processed perceptual region audio signal to a signal region to which the mask is not applied, based on a mask indicating a masking threshold derived from a psychoacoustic model, Device.

50. An apparatus for acquiring and transmitting a latent feature space representation of a perceptual domain audio signal using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of obtaining a perceptual domain audio signal by applying a mask indicating a masking threshold derived from a psychoacoustic model to the audio signal; (b) The step of inputting the perceptual domain audio signal into a neural network for mapping the perceptual domain audio signal to a latent feature space representation; (c) The step of obtaining the latent feature space representation of the perceptual domain audio signal as the output from the neural network; (d) The step of outputting the latent feature space representation of the perceptual domain audio signal as a bitstream, Device.

51. An apparatus for obtaining an audio signal from a latent feature space representation of a perceptual domain audio signal using a neural network, the apparatus comprising a neural network and one or more processors configured to perform a method, the method being: (a) The step of receiving the latent feature space representation of the perceptual domain audio signal as a bitstream; (b) The step of inputting the latent feature space representation into a neural network for generating the perceptual domain audio signal; (c) The step of obtaining the perceptual domain audio signal as the output from the neural network; (d) The step of converting the perceptual domain audio signal to a signal domain to which the mask is not applied, based on a mask indicating a masking threshold derived from a psychoacoustic model, Device.

52. An apparatus configured to perform the method described in any one of claims 1 to 48.

53. A computer program having instructions adapted to cause a device having processing capabilities to perform the method described in any one of claims 1 to 48 when executed by such a device.

54. A computer-readable storage medium having instructions adapted to cause a device having processing capabilities to perform the method described in any one of claims 1 to 48, when executed by such a device.