Neural Network Loss Conditional Training and Use for Audio Processing Using Neural Networks

The loss conditional training method for neural networks addresses the complexity of handling diverse audio signals by using a single network adjusted with coefficient vectors, enhancing audio quality across different categories and conditions efficiently.

JP2025520152AActive Publication Date: 2025-07-01DOLBY INTERNATIONAL AB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024570926
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-08
Filing Date
2023-06-07
Publication Date
2025-07-01
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

Existing deep learning models for audio processing are complex and computationally expensive due to the need for separate networks to handle different signal categories, bitrates, and codecs, leading to inefficiencies in training and inference.

Method used

A method for loss conditional training of a neural network involves randomly sampling a coefficient vector from a distribution to adjust the network based on weight coefficients, allowing a single neural network to cover a wide range of conditions, using Feature-wise Linear Modulation (FiLM) and training in a generative adversarial network setting.

Benefits of technology

This approach enables a single neural network to generate improved audio outputs across various signal categories and conditions, reducing the need for multiple networks and improving computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520152000001_ABST
    Figure 2025520152000001_ABST
Patent Text Reader

Abstract

A computer implementation method for loss conditional training of a neural network for outputting an upsampled audio signal, the method comprising: randomly sampling a coefficient vector from a coefficient distribution, wherein elements of the coefficient vector represent weight coefficients corresponding to loss terms of a loss function; adjusting the neural network based on the coefficient vector; and training the adjusted neural network based on an audio training signal, wherein the training involves calculating the loss function of the audio training signal after processing by the adjusted neural network using the weight coefficients indicated by the coefficient vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Application No. 63 / 350,099, filed on June 08, 2022, and European Patent Application No. 22177849.1, filed on June 08, 2022, all of which are hereby incorporated by reference herein.

[0002] Technique The present disclosure generally relates to methods for loss - conditional training of neural networks. In particular, coefficient vectors are randomly sampled from a distribution of coefficients, and the neural network is adjusted based on the coefficient vectors. The present disclosure further relates to computer - implemented methods for processing audio signals using loss - conditionally trained neural networks. The present disclosure further relates to respective apparatuses and respective computer program products.

[0003] Some embodiments are described herein with particular reference to their disclosure, but it will be understood that the present disclosure is not limited to such fields of use and is applicable in a broader context.

Background Art

[0004] Any discussion of background art throughout this disclosure should not be construed as an admission that such art is widely known or forms part of common general knowledge in the field.

[0005] Audio quality perceived by humans is a core performance metric in many audio devices. An audio codec is a computer program designed to encode and decode digital audio streams. More precisely, it compresses digital audio data into a compressed format and decompresses it from the compressed format with the help of codec algorithms. The audio codec is intended to reduce memory space and bandwidth while maintaining a high fidelity of the transmitted signal. However, lossy compression methods introduce encoding artifacts that may degrade the quality of the audio.

[0006] Deep learning approaches are becoming increasingly attractive in various application fields, including audio improvement. Most of the previous deep learning approaches have been related to speech noise removal.

[0007] Regarding noise removal in general, intuitively, one might think that encoding artifact reduction and noise removal are highly related. However, removing encoding artifacts / noises that are highly correlated with the desired sound often seems to be more complex than removing other noise types that are less correlated (in noise removal applications). The characteristics of encoding artifacts depend on the codec, the encoding tools used, and the selected bitrate. Additionally, modeling audio signals that include tonal content such as speech and music is further complicated due to the periodic functions naturally contained in this type of signal.

[0008] However, deep convolutional models used to reduce encoding artifacts and encoding noise are very complex in terms of model parameters and / or memory usage, thus introducing a high computational load by themselves. Further, when it is necessary to cover different signal categories such as speech, music, mixtures of speech and music, applause, etc., as well as various bitrates and codecs, typically separate models are trained, each model giving the best possible performance for each task.

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0009] In view of the above, there is a need to improve a single model for more arbitrary inputs covering different categories and conditions.

MEANS FOR SOLVING THE PROBLEMS

[0010] According to a first aspect of the present disclosure, a method (e.g., a computer-implemented method) of loss conditional training of a neural network for outputting an improved audio signal is provided. The method may include randomly sampling a coefficient vector from a coefficient distribution. Elements of the coefficient vector may represent weight coefficients of a loss function. The weight coefficients may correspond to loss terms of the loss function. The method may further include conditioning the neural network based on the coefficient vector. The method may also include training the conditioned neural network based on an audio training signal, the training involving calculating a loss function for the audio training signal after processing by the conditioned neural network using the weight coefficients indicated by the coefficient vector.

[0011] In some embodiments, the loss function may be a multi-objective loss function.

[0012] In some embodiments, the distribution of the coefficients may be a uniform distribution within a predetermined range.

[0013] In some embodiments, adjusting the neural network may include Feature-wise Linear Modulation (FiLM).

[0014] In some embodiments, randomly sampling the coefficient vector, adjusting the neural network, and training the adjusted neural network may form at least a part of an epoch, and the method may further include performing two or more epochs for each of a set of audio content types.

[0015] In some embodiments, training the adjusted neural network may be performed in perceptually weighted regions.

[0016] In some embodiments, the neural network may implement a deep learning-based generator, the generator including an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage being mapped to a latent feature space.

[0017] In some embodiments, the generator may be trained in a generative adversarial network (GAN) setting including a generator and a discriminator.

[0018] In some embodiments, adjusting the neural network may involve adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space.

[0019] In some embodiments, training the adjusted neural network is: Inputting an audio training signal into the adjusted generator; Generating a processed audio training signal based on the audio training signal by the adjusted generator; Inputting the processed audio training signal and the corresponding original audio signal from which the audio training signal is derived, one by one, into the discriminator; Determining by the discriminator whether the input audio signal is the processed audio training signal or the original audio signal; Sequentially and iteratively tuning the parameters of the generator until the discriminator can no longer distinguish the processed audio training signal from the original audio signal may be included.

[0020] In some embodiments, further, a random noise vector z may be applied to the latent feature space to modify the audio.

[0021] According to a second aspect of the present disclosure, a computer-implemented method of processing an audio signal using a loss-conditioned trained neural network is provided. The method may include adjusting the neural network based on adjustment information including a coefficient vector. Elements of the coefficient vector may indicate weight coefficients of a loss function. The weight coefficients may correspond to loss terms of the loss function. The method may further include inputting the audio signal into the adjusted neural network to process the audio signal. The method may further include processing the audio signal based on the adjustment information by the adjusted neural network. Also, the method may include obtaining an improved audio signal as an output from the adjusted neural network.

[0022] In some embodiments, the loss function may be a multi-objective loss function.

[0023] In some embodiments, the adjustment information may be based on the content type and / or bitrate of the audio signal.

[0024] In some embodiments, adjusting the neural network may include Feature-wise Linear Modulation (FiLM).

[0025] In some embodiments, the neural network may implement a deep learning-based generator, which includes an encoder stage and a decoder stage, each including a plurality of layers with one or more filters in each layer, and the last layer of the encoder stage is mapped to a latent feature space.

[0026] In some embodiments, adjusting the neural network may involve adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space.

[0027] In some embodiments, a random noise vector z may be applied to the latent feature space to modify the audio.

[0028] In some embodiments, the method may further include receiving an audio bitstream including the audio signal and adjustment information.

[0029] In some embodiments, the method may further include core-decoding the audio bitstream to obtain the audio signal.

[0030] In some embodiments, the method may further include extracting adjustment information from the received bitstream.

[0031] In some embodiments, the method may further include analyzing the audio signal and determining adjustment information based on the result of the analysis.

[0032] In some embodiments, the method may be performed in a perceptually weighted region, and an improved audio signal in the perceptually weighted region may be obtained as an output from the adjusted neural network.

[0033] In some embodiments, the method may further include converting the improved audio signal from the perceptually weighted region to the original signal region.

[0034] In some embodiments, the neural network may be trained in a perceptually weighted region.

[0035] According to a third aspect of the present disclosure, there is provided an apparatus for processing an audio signal using a loss-conditioned trained neural network. The apparatus may include one or more processors configured to execute a method that: adjusting a neural network based on adjustment information including a coefficient vector, wherein elements of the coefficient vector may indicate weight coefficients of a loss function, and the weight coefficients may correspond to loss terms of the loss function; inputting the audio signal into the adjusted neural network to process the audio signal; processing the audio signal based on the adjustment information by the adjusted neural network; and obtaining an improved audio signal as an output from the adjusted neural network.

[0036] According to a fourth aspect of the present disclosure, there is provided a computer program including instructions that, when executed by a computing device, cause the computing device to execute the method of loss-conditioned training of the neural network described herein.

[0037] According to a fifth aspect of the present disclosure, there is provided a computer-readable storage medium storing the computer program.

[0038] According to a sixth aspect of the present disclosure, there is provided a computer program including instructions that, when executed by a computing device, cause the computing device to execute a method of processing an audio signal using the loss-conditioned trained neural network described herein.

[0039] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium storing the computer program.

Brief Description of the Drawings

[0040] Exemplary embodiments of the present disclosure will now be described by way of example only, with reference to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0041] Overview In deep learning-based approaches for improving (symbolized) audio, the performance of a neural network (model) generally depends not on a single characteristic but on several characteristics. An approach for training a neural network is to balance these various characteristics by minimizing a loss function, which is a weighted sum of terms that measure these various characteristics. Depending on these weight coefficients, training using this loss function results in a model that is most suitable for a certain content type, bitrate, or codec.

[0042] However, if it is desired to cover different signal categories, such as speech, music, a mixture of speech and music, applause, as well as different bitrates and codecs, typically several separate neural networks with different weighting coefficients in the loss function are trained to obtain the best possible performance for each category. This requires maintaining multiple neural networks both during training and during inference, and is computationally expensive both during training and during inference.

[0043] The methods and apparatus described herein propose a loss conditional training and inference strategy that allows training and inference of a single neural network for tasks that would normally require a large set of separately trained neural networks. This is based on the idea that if all of these separate neural networks are solving highly related problems, some information can be shared between them. The loss function has coefficients to be tuned, and the described methods allow training and using a single neural network that covers a wide range of these coefficients. This provides a simple way to avoid inefficiencies and cover all trade-offs using a single neural network when normally a set of neural networks optimized for different losses would be required. That is, by varying the conditioning values, improved outputs can be generated using a single neural network.

[0044] Method for loss-conditioned training of a neural network Referring to the example of FIG. 1, a method of loss conditional training of a neural network is shown. Loss conditional training may include conditioning a neural network (model) with respect to a loss function having a weight coefficient λ corresponding to a loss term. For example, the loss function may consist of four terms, three reconstruction terms, and an adversarial term responsible for the generative characteristics of the neural network. The reconstruction terms control how similar the enhanced signal generated by the neural network should be to the original signal, while the adversarial loss defines the amount of generative features that should be carried over to the enhanced signal. Thus, the weight coefficient λ is sometimes said to balance / control the ratio of terms in the loss function. That is, for a single task / condition to be performed, there may be a single set (ordered set) of weight coefficients λ that determines an optimized loss function for that task / condition. Similarly, there may also be a set of weight coefficients λ that determines an optimized loss function that functions for multiple tasks / conditions. To select / find the set of weight coefficients λ for which the neural network functions best for multiple tasks / conditions, the training may involve covering a wide range of different sets of weight coefficients. This can be achieved by sampling respective vectors from respective distributions, as detailed below. The training may further involve covering different loss functions (determined by different weight coefficients) and / or different conditions such as different content types / bitrates / codecs, etc.

[0045] To train a single neural network to cover a wide range of coefficients, in step S101, the coefficient vector is randomly sampled from the distribution of the coefficients. The elements of the coefficient vector can indicate the weight coefficients of the loss function. It can be said that the coefficient vector represents an ordered set of the weight coefficients of the loss function. The number of weight coefficients in the set is determined by the number of terms and / or the weighting within each loss function. Thus, it can be said that the distribution of the coefficients represents the distribution of the loss function. This makes it possible to train a single neural network over a family of loss functions. The term loss function as used herein is also referred to as the generator loss function.

[0046] In some embodiments, the generator loss function may be a multi-objective loss function. For example, the generator loss function may include a multi-resolution STFT loss function as given by the following equation (2). The results indicate that the multi-resolution STFT-based generator loss function provides quality improvement for processing various signal categories. In other words, when a single neural network is trained on various signal categories, such as those that can be realized by training the neural network on different audio training signals, quality improvement can be achieved if the coefficient vector is randomly sampled from the distribution of the weight coefficients for the multi-objective loss function.

[0047] In some embodiments, the distribution of the coefficients may be a uniform distribution within a predetermined range. That is, each element may be sampled from a (1D) distribution within the range [0, 100], for example. In this case, since normalization can be considered part of the weighting, subsequent normalization of the vector is not required.

[0048] Referring again to the example of FIG. 1, once the coefficient vector is sampled, in step S102, the neural network is adjusted based on the coefficient vector. The adjustment may be performed via a conditioning network. That is, the calculations performed by the neural network can be adjusted or modulated by the coefficient vector. In some embodiments, adjusting the neural network may include Feature-wise Linear Modulation (FiLM). That is, FiLM layers may be introduced into the architecture of the neural network, and these layers are parameterized by an adjustment based on the coefficient vector. For example, a randomly sampled adjustment vector λ may be supplied to two multi-layer perceptron (MLP) networks, creating vectors σ(λ) and μ(λ) of the same dimension as the number of feature maps in the output of the convolutional / transposed layer to be modulated / adjusted. Each feature map is first scaled by σ(λ). Then, the scaled feature map is shifted by μ(λ).

[0049] Once the neural network is adjusted, in step S103, the adjusted neural network is then trained based on the audio training signal. The training may involve calculating a loss function for the audio training signal after processing by the adjusted neural network, using the weight coefficients indicated by the coefficient vector. In some embodiments, training the adjusted neural network may be performed in perceptually weighted regions.

[0050] The method described above enables the neural network to learn to model the entire family of loss functions. The architecture and training of the neural network will be described in more detail below.

[0051] Architecture of a neural network The above method can be implemented using any neural network, and thus it should be noted that the architecture of the neural network is not limited. However, in some embodiments, the neural network may implement a deep learning-based generator, which includes an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage maps to a latent feature space.

[0052] A non-limiting example of a simple architecture of the generator is schematically shown in the example of FIG. 3. The generator 1000 includes an encoder stage 1001 and a decoder stage 1002. The encoder stage 1001 and the decoder stage 1002 of the generator 1000 may be fully convolutional. The decoder stage 1002 may mirror the encoder stage 1001. The encoder stage 1001 and the decoder stage 1002 may each include a plurality of layers 1001a, 1001b, 1001c, 1002a, 1002b, 1002c having a plurality of filters in each layer, and the plurality of filters in each layer of the decoder stage can perform a filtering operation to generate a plurality of feature maps, and the last layer of the encoder stage 1001 can map to a latent feature space representation c* 1003.

[0053] That is, the encoder stage 1001 and the decoder stage 1002 may each include several L layers, each layer L having several N filters. L may be a natural number ≥ 1, and N may be a natural number ≥ 1. The size of the N filters (also known as the kernel size) is not limited, but the filter size may be the same for each of the L layers. For example, the filter size may be 31. In each layer, the number of filters may increase. Each filter may operate, for example, on the audio signal input to each layer of the generator with a stride of 2. Thus, learnable downsampling by a factor of 2 may be performed in the encoder layers, and learnable upsampling by a factor of 2 may be performed in the decoder layers. In other words, the encoder stage 1001 of the generator may include a plurality of 1D convolutional layers with a stride of 2 (without a bias term), and the decoder stage 1002 of the generator may include a plurality of 1D transposed convolutional layers with a stride of 2 (without a bias term).

[0054] In some embodiments, adjusting the neural network may involve adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space. In some embodiments, this adjustment may be a FiLM adjustment as described above.

[0055] In at least one layer of the encoder stage 1001, a non-linear operation may be further performed as an activation including one or more of a parametric rectified linear unit (PReLU), a rectified linear unit (ReLU), a leaky rectified linear unit (LReLU), an exponential linear unit (eLU), and a scaled exponential linear unit (SeLU). For example, the non-linear operation may be based on PReLU.

[0056] In some embodiments, the generator 1000 may further include a non-strided (stride = 1) transposed convolutional layer as an output layer following the last layer 1002a of the decoder stage 1002. The output layer may include, for example, N = 1 filter for a mono audio signal and N = 2 filters for a stereo audio signal as an example of a multi-channel audio signal. The filter size may be 31. In the output layer, since the audio signal output from the decoder stage 1002 needs to be constrained to +1 and -1, the activation may be based on the tanh operation tanh(x) activation.

[0057] As shown in the example of FIG. 3, in some embodiments, one or more skip connections 1005 may exist between respective like layers of the encoder stage 1001 and the decoder stage 1002 of the generator 1000. Here, the latent feature space representation c* 1003 is bypassed to prevent loss of information. The skip connection 1005 may be implemented using one or more of concatenation and signal addition. For the implementation of the skip connection 1005, the number of filter outputs may be "substantially" doubled.

[0058] In one embodiment, the random noise vector z(1004) can be further applied to the latent feature space representation c*(1003) to modify the audio.

[0059] Architecture of a discriminator As described above, in some embodiments, the generator may be a generator trained in a adversarial generative network setting (GAN setting). The GAN setting generally includes a generator G and a discriminator D that are trained by a sequential iterative process. The architecture of the discriminator may have the same structure as the encoder stage of the generator. In other words, the discriminator architecture may mirror the structure of the encoder stage of the generator, but without adjustment. That is, the discriminator may include multiple layers each having multiple filters in each layer. That is, the discriminator may include some L layers, each having some N filters in layer L. L may be a natural number ≥ 1, and N may be a natural number ≥ 1. The size of the N filters (also known as the kernel size) is not limited, but the filter size may be the same in each of the L layers. For example, the filter size may be 31. In each layer, the number of filters may increase. Each of the filters may operate on the audio signal input to each layer of the discriminator, for example, with a stride of 2. In other words, the discriminator may include multiple 1D convolutional layers with a stride of 2 (without a bias term). The non-linear operation performed in at least one layer of the discriminator may include LReLU. Prepending, the discriminator may include an input layer. The input layer may be a non-stride convolutional layer (stride = 1 means non-stride). The discriminator may further include an output layer. The output layer may have an N = 1 filter with a filter size of 1 (the discriminator makes a single true / false judgment). Here, the filter size of the output layer may be different from the filter size of the discriminator layer. Thus, the output layer may be a one-dimensional (1D) convolutional layer that does not downsample the hidden activation. This means that the filter in the output layer operates with a stride of 1, while all previous layers of the discriminator may use a stride of 2. The activation in the output layer may be different from the activation in the at least one of the discriminator layers. The activation may be sigmoid.However, when the least squares training method is used, sigmoid activation may not be required and is thus optional.

[0060] Method for loss-conditioned training of a generator in a generative adversarial network (GAN) setting In the following, the loss conditional training of the generator in the adversarial generation network setting (GAN setting) will be described. Generally, during training in the adversarial generation network setting, the generator G maps to a latent feature space representation using an encoder stage and upsamples the latent feature space representation using a decoder stage to generate a processed audio training signal x*. The audio training signal can be derived from the original audio signal x that has been encoded and decoded respectively. A random noise vector can be applied to the latent feature space representation. However, the random noise vector may be set to z = 0. Setting the random noise vector to z = 0 may yield the best results for reducing encoding artifacts. Alternatively, training may also be performed without the input of the random noise vector z.

[0061] During training, the generator G attempts to output a processed audio training signal x* that is indistinguishable from the original audio signal x. The discriminator D is fed the generated processed audio training signal x* and the original audio signal x one at a time, and determines in a fake / real fashion whether the input signal is the processed audio training signal x* or the original audio signal x. Here, the discriminator D attempts to distinguish the original audio signal x from the processed audio training signal x*. During the sequential iterative process, the generator G tunes its parameters to generate an increasingly better processed audio training signal x* compared to the original audio signal x, and the discriminator D learns to better distinguish between the processed audio training signal x* and the original audio signal x. This adversarial learning process can be described by the following equation (1).

Number

[0062] Note that discriminator D may initially be trained to train generator G in the final step. Training and updating discriminator D may involve maximizing the probability of assigning a high score to the original audio signal x and a low score to the processed audio training signal x*. The goal in training discriminator D may be for the original audio signal (unencoded) to be recognized as genuine while the processed audio training signal x* (generated) is recognized as fake. The parameters of generator G may be left fixed while discriminator D is being trained and updated.

[0063] Next, training and updating generator G may involve minimizing the difference between the original audio signal x and the generated processed audio training signal x*. The goal when training generator G may be to achieve that discriminator D recognizes the generated processed audio training signal x* as genuine.

[0064] Referring to the example of FIG. 2, conditional training of the loss of the generator in an adversarial generation network (GAN) setting including a generator and a discriminator is shown. Conditional training of the loss of generator G(100) may include the following.

[0065] The coefficient vector λ is randomly sampled from the coefficient distribution p(t). The distribution may be uniform in the range [0, 100]. It can be said that the coefficient vector represents an ordered set of weight coefficients λ for each loss function. The number of weight coefficients in the set is determined by the number of terms and / or weighting within each loss function. Thus, it can be said that the coefficient distribution represents the distribution of the loss functions. In one example, the elements of the coefficient vector can be independently sampled from a predetermined range (e.g., the range [0, 100]).

[0066] Once the coefficient vector is sampled, the generator G100 is then adjusted based on the coefficient vector 108. The adjustment may include Feature-wise Linear Modulation (FiLM). In some embodiments, adjusting the neural network may involve adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space. That is, for example, only the one or more layers of the encoder stage may be modulated by FiLM adjustment as described above.

[0067] The audio training signal (tilde-x) 103 and optionally the random noise vector z 104 can be input to the (FiLM) adjusted generator G100. In one embodiment, the random noise vector z may be set to z = 0. Alternatively, the training may be performed without the input of the random noise vector z.

[0068] The audio training signal (tilde-x) 103 can be obtained by encoding and decoding the original audio signal x 102. Then, based on the input, the adjusted generator G100 processes the input using the encoder stage to map it to a latent feature space representation and the decoder stage to upsample the latent feature space representation, thereby generating the processed audio training signal x* 105.

[0069] The original audio signal x 102 from which the audio training signal (tilde x) 103 is derived, one by one at a time, and the generated processed audio training signal x* 105 are input to the discriminator D 101. As additional information, the audio training signal (tilde x) 103 may also be input to the discriminator D 101 each time. Next, the discriminator D 101 determines (106) whether the input data is the processed audio training signal x* 105 (fake) or the original audio signal x 102 (real).

[0070] In the next step, the parameters of the adjusted generator G 100 are adjusted until the discriminator D 101 can no longer distinguish the processed audio training signal x* 105 from the original audio signal x 102. This may be done in an iterative process 107.

[0071] The determination 101 by the discriminator D may be based on one or more perceptually motivated objective functions according to the following equation (2). [Equation]

[0072] As can be seen from the first term in Equation (2), the adjusted adversarial generation network setting is applied by inputting the audio training signal (tilde x) to the discriminator as additional information.

[0073] The two terms in the above Equation (2) (of the generator loss function) including the coefficients λ2 and λ3 are sometimes called the multi-resolution STFT loss terms. The multi-resolution STFT loss can be said to be the sum of different STFT-based loss functions using different STFT parameters. L sc m (Spectral convergence loss) and L mag m(Logarithmic scale STFT magnitude loss) can apply STFT-based losses at M different resolutions, each with the number of FFT bins ∈ {512, 1024, 2048}, hop size ∈ {50, 120, 240}, and finally window length ∈ {240, 600, 1200}. The results showed that multiple-resolution STFT loss terms provide quality improvement for handling general audio (i.e., any content type).

[0074] The term containing coefficient λ1 in Equation (2) is the 1-norm distance scaled by factor lambda λ1. The value of this lambda can be selected from 1 to 100 depending on the application and / or the signal length input to the generator. For example, λ1 may be selected such that λ1 = 100. Further, the scaling for the multiple-resolution STFT loss terms may be set to the same value as lambda.

[0075] The term

Number

[0076] Alternatively, the generator loss function of Equation (2) may be applied in the following form.

Number

[0077] In the above equation, the term

Number

[0078] In some embodiments, training the adjusted neural network (adjusted generator) can be performed in the perceptually weighted region. The perceptually weighted audio training signal (tilde-x) can then be input to the adjusted generator G100. The perceptually weighted audio training signal (tilde-x) may be obtained by encoding and decoding the original perceptually weighted audio signal x. The original perceptually weighted audio signal x is derived by applying a mask or masking curve P to the original audio signal, and the mask or masking curve indicates a masking threshold derived from a psychoacoustic model. Then, based on the input, the adjusted generator G100 generates a processed perceptually weighted audio training signal x*. In this case, the original perceptually weighted audio signal x from which the perceptually weighted audio training signal (tilde-x) is derived and the generated processed perceptually weighted audio training signal x* are input to the discriminator D101.

[0079] In some embodiments, randomly sampling a coefficient vector, adjusting a neural network, and training the adjusted neural network may form at least a portion of an epoch, and the method may further include performing two or more epochs for each of a set of audio content types. The set of audio content types may include, for example, one or more of speech, music, speech and music, and applause. When performing two or more epochs for each of the set of audio content types, the neural network may be trained under multiple conditions. During training under multiple conditions, the neural network "looks at" the distribution of weight coefficients and "selects" the best coefficient vector that functions best across the multiple conditions. Note that in addition to the audio content type, the multiple conditions may also include one or more of a set of bitrates and a set of codecs.

[0080] For example, in each epoch, the original audio signal is taken from a set of audio samples related to the content type. Next, the coefficient vector is randomly sampled and used to adjust the neural network (generator). The adjusted neural network then processes each audio training signal, and a loss function is calculated for the processed audio training signal. Here, the weight coefficients of the loss function are given by the coefficient vector λ. Training is then continued until the value of the calculated loss function is minimized as described above. That is, in each epoch, for each pair of an audio training signal derived from the original audio signal taken from the set of audio samples and a randomly sampled coefficient vector, the parameters of the adjusted generator are tuned until the discriminator can no longer distinguish the processed audio training signal from the original audio signal. When performing two or more epochs for each set of audio content types, a single model can be trained to cover different signal categories. However, for each signal category, the training yields the best possible coefficient vector. Similarly, one or more epochs can be performed for a set of bitrates and / or codecs. For example, the neural network may first be trained with a set of audio content types as described above. The same neural network may then be further trained with a set of bitrates and / or a set of codecs, or vice versa. This training method yields a single coefficient vector that functions best for multiple conditions experienced during training. However, this training method also allows a single neural network during inference to "remember" which coefficient vector yielded the best results during training for one particular condition (bitrate / codec / content type).Then, during inference, this vector may be selected / included in the conditioning information based on each bitrate / content type / codec.

[0081] Referring again to the example of FIG. 2, generally, the training of the discriminator D 101 can follow the same general process as described above for the training of the generator G 100. In this case, the parameters of the generator G 100 may be fixed, while the parameters of the discriminator D 101 may be varied. The training of the discriminator D 101 can be described by the following equation (3) that enables the discriminator D 101 to discriminate the processed audio training signal x* 105 as false.

Equation

[0082] In the above case, the least squares approach (LS) and the adjusted adversarial generation network setting are also applied by inputting the audio training signal (x with a tilde) to the discriminator D 101 as additional information.

[0083] In addition to the least squares method, in the adversarial generation network setting, other training methods can also be used to train the adjusted generator and discriminator. The present disclosure is not limited to a specific training method. Alternatively or additionally, a so-called Wasserstein approach may be used. In this case, instead of the least squares distance, the Earth Mover's distance, also known as the Wasserstein distance, may be used. Generally, different training methods make the training of the adjusted generator and discriminator more stable. However, the type of training method applied does not affect the architecture of the (adjusted) generator.

[0084] Method for processing an audio signal using a loss-conditioned trained neural network Referring to the example of FIG. 4, a computer-implemented method of processing an audio signal using a loss-conditioned trained neural network is shown. In step S201, the neural network is adjusted based on adjustment information including a coefficient vector. In some embodiments, as described above, the elements of the coefficient vector may represent the weight coefficients of the loss function. In some embodiments, the loss function may be a multi-objective loss function such as, for example, the multi-resolution STFT-based generator loss function of Equation (2).

[0085] In some embodiments, the adjustment information may be (may be determined by) based on the content type and / or bit rate of the audio signal to be processed. The adjustment information may further be selected depending on the training results, for example, based on which coefficient vector functioned best. As described above, the loss-conditioned trained neural network may be trained by performing two or more epochs for each set of audio content types so that the neural network can cover a wide variety of different categories and conditions. Similarly, this may be applicable to bit rate and / or codec. That is, the adjustment information may alternatively or additionally be based on the codec.

[0086] In some embodiments, adjusting the neural network may include feature-wise linear modulation (FiLM) as described above. In some embodiments, the neural network may implement a deep learning-based generator. The generator may include an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage maps to a latent feature space. Adjusting the neural network may then involve adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space.

[0087] Furthermore, in some embodiments, a random noise vector z may be applied to the latent feature space to modify the audio.

[0088] Referring to the example of FIG. 5, in some embodiments, the method may further include receiving (200) an audio bitstream. The audio bitstream (200) includes an audio signal and adjustment information. In this case, the method may further include, for example, core-decoding the audio bitstream by an audio decoder 201 to obtain an audio signal 203. The method may also include extracting adjustment information 204 from the received bitstream 200. In an exemplary embodiment, the audio bitstream may include metadata, the adjustment information may be included in the metadata, or the metadata may indicate the adjustment information. The trained neural network 202 may then be adjusted based on the extracted adjustment information 204. Next, the audio signal 203, and optionally the random noise vector z 205, may be input to the trained and adjusted neural network 202 to process the audio signal 203. As an output from the trained and adjusted neural network 202, a processed audio signal 206 is obtained.

[0089] Alternatively, in some embodiments, the method may further include analyzing (203) the audio signal and determining (204) adjustment information based on the result of the analysis.

[0090] Referring again to the example of FIG. 4, in step S202, the audio signal is input to an adjusted neural network to process the audio signal. In step S203, the adjusted neural network processes the audio signal based on the adjustment information. Then, as an output from the adjusted neural network, in step S204, a processed audio signal is obtained.

[0091] In some embodiments, the method may be performed in a perceptually weighted domain. Then, the processed audio signal in the perceptually weighted domain can be obtained as an output from the adjusted neural network. In this case, the method may further include converting the processed audio signal from the perceptually weighted domain back to the original signal domain. In some embodiments, the neural network may be trained in the perceptually weighted domain.

[0092] Referring to the example of FIG. 6, the method described above can be implemented by an apparatus for processing an audio signal using a loss-conditioned trained neural network. Apparatus 300 may include one or more processors 301 configured to perform the methods described above.

[0093] Alternatively or additionally, the methods described herein may also be implemented by a computer program including instructions that, when executed by a computing device, cause the computing device to perform the methods described herein. The computer program may be provided on a computer-readable storage medium that stores the computer program.

[0094] The results showed that a neural network trained as described herein can generate outputs (processed audio signals) for different content types / bitrates / codecs that are as good as the outputs generated by individual neural networks trained for each specific task. That is, the methods and apparatus described herein allow for training and inference of a single neural network for tasks that would normally require a large set of neural networks trained separately. This one-to-many approach provides a simple way to improve the efficiency of training and inference. Specifically, this approach reduces memory consumption on the decoder side because only one neural network needs to be stored.

[0095] Interpretation Unless otherwise specified, as will be apparent from the following description, throughout the description of the present disclosure, discussions using terms such as "processing," "calculating," "determining," "analyzing," etc. refer to actions and / or processes of a computer or computing system, or similar electronic computing devices that manipulate and / or transform data represented as physical quantities, such as electronic quantities, into other data represented as physical quantities.

[0096] Similarly, the term "processor" can refer to any device or part of a device that processes electronic data from, for example, registers and / or memory and converts that electronic data into other electronic data that can be stored, for example, in registers and / or memory. A "computer" or "computing machine" or "computing platform" can include one or more processors.

[0097] The method described herein is executable by one or more processors that accept computer-readable (also called machine-readable) code including a set of instructions that, when executed by one or more of the processors, execute at least one of the methods described herein. Any processor capable of executing a set of instructions (sequentially or otherwise) that specify the actions to be taken is included. Thus, one example is a typical processing system that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system may further include a memory subsystem that includes main RAM and / or static RAM, and / or ROM. A bus subsystem may be included for communication between components. The processing system may further be a distributed processing system having processors coupled by a network. If the processing system requires a display, such a display may include, for example, a liquid crystal display (LCD) or a cathode ray tube (CRT) display. If manual data input is required, the processing system also includes input devices such as one or more of an alphanumeric input unit such as a keyboard, a pointing control device such as a mouse, etc. The processing system may include a storage system such as a disk drive unit. In some configurations, the processing system may include an audio output device and a network interface device. Thus, the memory subsystem includes a computer-readable carrier medium carrying computer-readable code (e.g., software) including a set of instructions for causing one or more of the methods described herein to be executed when executed by one or more processors. Note that when a method includes several elements, e.g., several steps, the order of such elements is not implied unless specifically stated. The software may be present on a hard disk or may be present, fully or at least partially, in RAM and / or in the processor during its execution by the computer system.Thus, the memory and the processor also constitute a computer-readable carrier medium carrying computer-readable code. Further, the computer-readable carrier medium may form or be included in a computer program product.

[0098] In an alternative exemplary embodiment, one or more processors may operate as a stand-alone device or be connected to other processors in a networked deployment, for example, network-connected. One or more processors may operate as a server or user machine in a server-user network environment or as a peer machine in a peer-to-peer or distributed network environment. One or more processors may form any machine capable of executing a set (sequential or otherwise) of instructions that specify actions to be taken by a personal computer (PC), tablet PC, personal digital assistant (PDA), cellular phone, web appliance, network router, switch or bridge, or that machine.

[0099] Note that the term "machine" is also to be construed as including any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods described herein.

[0100] Thus, each one exemplary embodiment of the methods described herein is in the form of a set of instructions, such as a computer program, for execution on one or more processors, such as one or more processors that are part of a web server configuration. Thus, as will be understood by those skilled in the art, the exemplary embodiments of the present disclosure may be embodied as a method, an apparatus such as a dedicated apparatus, an apparatus such as a data processing system, or a computer-readable carrier medium, such as a computer program product. The computer-readable carrier medium carries computer-readable code including a set of instructions that, when executed on one or more processors, cause the one or more processors to implement the method. Thus, aspects of the present disclosure may take the form of a method, an exemplary embodiment that is entirely hardware, an exemplary embodiment that is entirely software, or an exemplary embodiment that combines software aspects and hardware aspects. Further, the present disclosure may take the form of a carrier medium (e.g., a computer program product on a computer-readable storage medium) carrying computer-readable program code embodied in the medium.

[0101] The software may be further transmitted or received through a network via a network interface device. The carrier medium is a single medium in an exemplary embodiment, but the term "carrier medium" should be construed to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache and server) that store one or more sets of instructions. The term "carrier medium" should also be construed to include any medium that can store, encode, or carry a set of instructions for execution by one or more processors and that can cause one or more processors to execute any one or more of the methods of the present disclosure. The carrier medium can take many forms including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical disks, magnetic disks, and magneto-optical disks. Volatile media includes dynamic memory such as main memory. Transmission media includes coaxial cables, copper wire, and fiber optics, including wire in a bus subsystem. Transmission media can also take the form of acoustic or light waves such as those generated during radio wave and infrared data communications. Thus, for example, the term "carrier medium" should be construed to include computer products embodied in solid memory, optical and magnetic media; media that carry a propagated signal detectable by at least one processor or one or more processors and that represents a set of instructions for performing a method when executed; and transmission media in a network that carry a propagated signal detectable by at least one of one or more processors and that represents a set of instructions, but is not limited thereto.

[0102] The steps of the method discussed, in one exemplary embodiment, will be understood to be performed by a suitable processor (or processors) of a processing (e.g., computer) system that executes instructions (computer-readable code) stored in a memory device. It should be understood that the present disclosure is not limited to any particular implementation or programming technique, and that the present disclosure may be implemented using any suitable technique for implementing the functions described herein. The present disclosure is not limited to any particular programming language or operating system.

[0103] Throughout this disclosure, references to "an embodiment", "some embodiments", or "exemplary embodiments" mean that the particular features, structures, or characteristics described in connection with those embodiments are included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases "in an embodiment", "in some embodiments", or "in an exemplary embodiment" throughout this disclosure are not necessarily all referring to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more exemplary embodiments, as will be apparent to those skilled in the art from the present disclosure.

[0104] As used herein, unless otherwise specified, the use of ordinal adjectives such as "first", "second", "third", etc. to describe a common object merely indicates that different instances of similar objects are being referred to, and is not intended to imply that the objects so described must be in a given order, whether temporally, spatially, in ranking, or in any other manner.

[0105] In the following claims and the description of this specification, any of the terms having, consisting of, or comprising are open terms meaning including at least the elements / features that follow it, but not excluding others. Thus, when used in the claims, the term comprising should not be construed as being limited to the recited means or elements or steps. For example, the scope of the expression a device comprising A and B should not be limited to a device consisting only of elements A and B. Any of the terms comprising or including or incorporating used in this specification are also open terms meaning including at least the elements / features that follow that term, but not excluding others. Thus, comprising is synonymous with having and means having.

[0106] In the foregoing description of the exemplary embodiments of the present disclosure, it should be understood that various features of the present disclosure may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of enhancing the flow of the present disclosure and assisting in the understanding of one or more of the various aspects of the invention. However, this method of disclosure should not be construed as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as reflected in the following claims, aspects of the invention lie in less than all of the features of a single foregoing disclosed exemplary embodiment. Thus, the claims that follow this specification are expressly incorporated herein, and each claim stands on its own as a separate exemplary embodiment of the present disclosure.

[0107] Furthermore, some of the exemplary embodiments described herein include some features that are included in other exemplary embodiments, but do not include other features, and combinations of features of different exemplary embodiments are intended to be within the scope of the present disclosure and form different exemplary embodiments. This will be understood by those skilled in the art. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.

[0108] In the description provided herein, numerous specific details are set forth. However, it is understood that exemplary embodiments of the disclosure may be practiced without these specific details. On the other hand, well-known methods, structures, and techniques have not been shown in detail in order not to obscure the understanding of this document.

[0109] Thus, while what is considered to be the best mode of the disclosure has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the disclosure. It is intended to claim all such changes and modifications that fall within the scope of the disclosure. For example, any of the formulas given above are merely representative examples of procedures that may be used. Functions may be added to or removed from the block diagrams, and operations may be exchanged between functional blocks. Steps may be added to or removed from the described methods within the scope of the disclosure.

[0110] Various aspects and implementations of the disclosure may also be understood from the following numbered example embodiments (EEE) that are not the claims.

[0111] 〔EEE1〕 A method for loss conditional training of a neural network, the method comprising: randomly sampling a coefficient vector from a coefficient distribution, wherein elements of the coefficient vector represent weight coefficients of a loss function; adjusting the neural network based on the coefficient vector; training the adjusted neural network based on an audio training signal, the training including calculating the loss function for the audio training signal after processing by the adjusted neural network using the weight coefficients indicated by the coefficient vector; and including. 〔EEE2〕 The method according to EEE1, wherein the loss function is a multi-objective loss function. 〔EEE3〕 The method according to EEE1 or 2, wherein the distribution of the coefficients is a uniform distribution within a predetermined range. 〔EEE4〕 The method according to any one of EEE1 to 3, wherein adjusting the neural network includes Feature-wise Linear Modulation (FiLM). 〔EEE5〕 The method according to any one of EEE1 to 4, wherein randomly sampling the coefficient vector, adjusting the neural network, and training the adjusted neural network form at least a part of an epoch, and the method further includes executing two or more epochs for each set of audio content types. 〔EEE6〕 The method according to any one of EEE1 to 5, wherein training the adjusted neural network is performed in a perceptually weighted region. 〔EEE7〕 The method according to any one of EEE1 to 6, wherein the neural network implements a deep learning-based generator, the generator includes an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage is mapped to a latent feature space. 〔EEE8〕 The method according to EEE6, wherein the generator is trained in an adversarial generative network (GAN) setting including the generator and a discriminator. 〔EEE9〕 The method according to EEE6, wherein adjusting the neural network includes adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space. 〔EEE10〕 Training the adjusted neural network comprises: inputting an audio training signal into the adjusted generator; generating, by the adjusted generator, a processed audio training signal based on the audio training signal; inputting the processed audio training signal and the corresponding original audio signal from which the audio training signal is derived, one by one, into the discriminator; determining, by the discriminator, whether the input audio signal is the processed audio training signal or the original audio signal; sequentially and iteratively tuning the parameters of the generator until the discriminator can no longer distinguish the processed audio training signal from the original audio signal, The method according to EEE8 or 9. 〔EEE11〕 The method according to EEE10, wherein a random noise vector z is applied to the latent feature space to modify the audio. 〔EEE12〕 A computer-implemented method of processing an audio signal using a loss-conditioned trained neural network, the method comprising: adjusting the neural network based on adjustment information including a coefficient vector; inputting the audio signal into the adjusted neural network to process the audio signal; processing the audio signal by the adjusted neural network based on the adjustment information; obtaining a processed audio signal as an output from the adjusted neural network. Method. 〔EEE13〕 The method according to EEE12, wherein the coefficient vector represents a weight coefficient of a loss function. 〔EEE14〕 The loss function is the method described in EEE13, which is a multi-objective loss function. 〔EEE15〕 The adjustment information is the method described in any one of EEE12 to 14 based on the content type and / or bit rate of the audio signal. 〔EEE16〕 Adjusting the neural network is the method described in any one of EEE12 to 15, including feature-wise linear modulation (FiLM). 〔EEE17〕 The neural network implements a deep learning-based generator, and the generator includes an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage is mapped to a latent feature space, which is the method described in any one of EEE12 to 16. 〔EEE18〕 Adjusting the neural network includes adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space, which is the method described in EEE17. 〔EEE19〕 A random noise vector z is applied to the latent feature space to modify the audio, which is the method described in EEE18 or 19. 〔EEE20〕 The method further includes receiving an audio bitstream including the audio signal and the adjustment information, which is the method described in any one of EEE12 to 19. 〔EEE21〕 The method further includes core-decoding the audio bitstream to obtain the audio signal, which is the method described in EEE20. 〔EEE22〕 The method further includes extracting the adjustment information from the received bitstream, which is the method described in EEE20 or 21. 〔EEE23〕 The method according to any one of EEE12 to 21, further comprising analyzing the audio signal and determining the adjustment information based on the result of the analysis. 〔EEE24〕 The method according to any one of EEE12 to 23, wherein the method is performed in a perceptually weighted region, and the processed audio signal in the perceptually weighted region is obtained as an output from the adjusted neural network. 〔EEE25〕 The method according to EEE24, further comprising converting the processed audio signal from the perceptually weighted region to the original signal region. 〔EEE26〕 The method according to any one of EEE12 to 25, wherein the neural network is trained in the perceptually weighted region. 〔EEE27〕 An apparatus for processing an audio signal using a neural network trained with loss conditioning, the apparatus including one or more processors configured to execute a method, the method comprising: adjusting the neural network based on adjustment information including a coefficient vector; inputting the audio signal to the adjusted neural network to process the audio signal; processing the audio signal based on the adjustment information by the adjusted neural network; obtaining a processed audio signal as an output from the adjusted neural network. Apparatus. 〔EEE28〕 A computer program including instructions that, when executed by a computing device, cause the computing device to execute the method according to any one of EEE1 to 11. 〔EEE29〕 A computer-readable storage medium storing the computer program described in EEE28. 〔EEE30〕 A computer program that, when executed by a computing device, includes instructions for causing the computing device to execute the method according to any one of EEE12 to 26. 〔EEE31〕 A computer-readable storage medium storing the computer program described in EEE30.

Claims

1. A method for loss conditional training of a neural network implemented on a computer to output an improved audio signal, the method comprising: randomly sampling a coefficient vector from a distribution of coefficients, wherein elements of the coefficient vector represent weight coefficients corresponding to loss terms of a loss function; adjusting the neural network based on the coefficient vector; training the adjusted neural network based on an audio training signal, the training including calculating the loss function for the audio training signal after processing by the adjusted neural network using the weight coefficients indicated by the coefficient vector; The method includes the above steps.

2. The method according to claim 1, wherein the loss function is a multi-objective loss function.

3. The method according to claim 1, wherein the distribution of the coefficients is a uniform distribution within a predetermined range.

4. The method according to claim 1, wherein adjusting the neural network includes Feature-wise Linear Modulation (FiLM).

5. Randomly sampling the coefficient vector, adjusting the neural network, and training the adjusted neural network form at least a part of an epoch, and the method further includes performing two or more epochs for each set of audio content types.

6. The method according to claim 1, wherein training the adjusted neural network is performed in a perceptually weighted region.

7. The neural network implements a deep learning-based generator, the generator comprising an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage being mapped to a latent feature space.

8. The method according to claim 7, wherein adjusting the neural network includes adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space. Claim 9 The method according to claim 7, wherein the generator is trained in an adversarial generation network (GAN) setting including the generator and a discriminator. Claim 10 Training the adjusted neural network comprises: inputting an audio training signal into the adjusted generator; generating, by the adjusted generator, a processed audio training signal based on the audio training signal; inputting, one by one, the processed audio training signal and the corresponding original audio signal from which the audio training signal is derived into the discriminator; determining, by the discriminator, whether the input audio signal is the processed audio training signal or the original audio signal; sequentially and iteratively tuning the parameters of the generator until the discriminator can no longer distinguish the processed audio training signal from the original audio signal, the method according to claim 9. The method according to claim 9 Claim 11 The method according to claim 10, wherein a random noise vector z is applied to the latent feature space to modify the audio. Claim 12 A computer-implemented method of processing an audio signal using a loss-conditioned trained neural network, the method comprising: adjusting the neural network based on adjustment information including a coefficient vector, wherein elements of the coefficient vector indicate weight coefficients corresponding to loss terms of a loss function; inputting the audio signal into the adjusted neural network to process the audio signal; processing the audio signal by the adjusted neural network based on the adjustment information; obtaining an improved audio signal as an output from the adjusted neural network. Method Claim 13 The method according to claim 12, wherein the loss function is a multi-objective loss function. Claim 14 The method according to claim 12, wherein the adjustment information is based on a content type and / or bit rate of the audio signal. Claim 15 Adjusting the neural network is the method according to claim 12, including feature-wise linear modulation (FiLM).

16. The neural network implements a deep learning-based generator, the generator including an encoder stage and a decoder stage, each including a plurality of layers having one or more filters in each layer, and the last layer of the encoder stage being mapped to a latent feature space, the method according to claim 12.

17. Adjusting the neural network includes adjusting one or more layers of the encoder stage of the generator adjacent to the latent feature space, the method according to claim 16.

18. A random noise vector z is applied to the latent feature space to modify the audio, the method according to claim 16.

19. The method further includes receiving an audio bitstream including the audio signal and the adjustment information, the method according to claim 12.

20. The method further includes core-decoding the audio bitstream to obtain the audio signal, the method according to claim 19.

21. The method further includes extracting the adjustment information from the received bitstream, the method according to claim 19.

22. The method further includes analyzing the audio signal and determining the adjustment information based on the result of the analysis, the method according to claim 12.

23. The method is executed in a perceptually weighted region, and an improved audio signal in the perceptually weighted region is obtained as an output from the adjusted neural network, the method according to claim 12.

24. The method further includes converting the improved audio signal from the perceptually weighted region to the original signal region, the method according to claim 23.

25. The neural network is trained in the perceptually weighted region, the method according to claim 12.

26. An apparatus for processing an audio signal using a loss-conditioned trained neural network, the apparatus including one or more processors configured to execute a method, the method comprising: Adjusting the neural network based on adjustment information including a coefficient vector, wherein elements of the coefficient vector indicate weight coefficients corresponding to loss terms of a loss function; Inputting the audio signal into the adjusted neural network to process the audio signal; Processing the audio signal based on the adjustment information by the adjusted neural network; Obtaining an improved audio signal as an output from the adjusted neural network, Device.

27. A computer program including instructions that, when executed by a computing device, cause the computing device to execute the method according to any one of claims 1 to 11.

28. A computer-readable storage medium storing the computer program according to claim 27.

29. A computer program including instructions that, when executed by a computing device, cause the computing device to execute the method according to any one of claims 12 to 25.

30. A computer-readable storage medium storing the computer program according to claim 29.

Citation Information

Patent Citations

  • Disturbance component suppressing device, computer program, and speech recognition system

    JP2006243290A

  • Signal processor, signal processing method and program

    JP2020148909A

  • Voice recognition system and method

    JP2022079397A

  • Neural tuning code for multilingual style-dependent spoken language processing

    JP2022512233A

  • Method And System For Speech Enhancement

    US20210241780A1