Apparatus for providing a processed audio signal, method for providing a processed audio signal, apparatus for providing neural network parameters and method for providing neural network parameters
Patent Information
- Application Number
- CN202180085895.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-20
- Filing Date
- 2021-05-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2041-05-06
Smart Images

Figure CN116648747B_ABST
Abstract
Description
Technical Field
[0001] According to embodiments of the present invention, there is an apparatus for providing processed audio signals.
[0002] Other embodiments of the invention relate to a method for providing processed audio signals.
[0003] Other embodiments of the present invention relate to an apparatus for providing neural network parameters.
[0004] Other embodiments of the present invention relate to a method for providing neural network parameters.
[0005] Embodiments of this application relate to audio signal processing using neural networks, particularly to audio signal enhancement, and especially to speech enhancement.
[0006] According to one aspect, embodiments of the invention can be applied to provide direct enhancement of noisy speech via neural networks. Background Technology
[0007] Several audio enhancement methods are known, particularly speech enhancement that distinguishes the target speech signal from intrusive background. The goal of speech enhancement is to emphasize the target speech signal from interfering backgrounds to ensure clearer spoken content. Speech enhancement is important for a wide range of applications, including, for example, hearing aids or automatic speech recognition.
[0008] In recent years, various generation methods for speech enhancement have been increasingly used, such as variational autoencoders, generative adversarial networks (GANs), and autoregressive models.
[0009] In view of the above, it is desirable to create a concept for audio signal enhancement that offers an improved trade-off between computational complexity and achievable audio quality.
[0010] This objective is achieved through the subject matter of the pending independent claims.
[0011] Other advantages are the subject of the dependent claims. Summary of the Invention
[0012] An apparatus is created according to embodiments of the present invention for providing a processed audio signal, such as a processed speech signal, such as an enhanced audio signal, such as an enhanced speech signal, or such as an enhanced general audio signal, based on an input audio signal, such as a speech signal, such as a distorted audio signal, such as a noisy speech signal y, such as a clean signal x extracted from the noisy speech signal y. ), where, for example, y = x + n, and n is noise, such as a noisy background. The device is configured to use one or more process blocks, such as eight process blocks, such as using a process block system, such as including affine coupling layers, such as including reversible convolution, such as using affine scaling or a series of affine scaling operations, to process a noise signal, such as z, or a signal derived from a noise signal, in order to obtain a processed audio signal, such as an enhanced audio signal, for example... The device is configured to adapt a process performed using one or more flow blocks to an input audio signal, such as a distorted audio signal, or a noisy speech signal y; for example, a noisy time-domain speech sample. For example, the neural network is based on the distorted audio signal and preferably also provides one or more processing parameters to the flow blocks, such as parameters for affine processing, like scaling factors and shift values, based on at least a portion of the noisy signal or its processed version.
[0013] This embodiment is based on the following findings: for example, for speech enhancement purposes, the processing of audio signals can be directly performed using block processing, which can, for example, model the generation process. It has been found that block processing allows processing, in a manner conditioned on the input audio signal, such as a noisy frequency signal y, noise signals, such as noise signals z, generated by the device or stored in the device. The noise signal z represents (or includes) a given (e.g., simple or complex) probability distribution, preferably a Gaussian probability distribution. It has been found that when processing noise signals conditioned on distorted audio signals, as a result of the processing, an enhanced clean portion of the input audio signal is provided without introducing that clean portion, for example, in the absence of a noisy background, as input to the device.
[0014] The proposed apparatus provides an efficient and easily implemented audio signal processing method, particularly direct audio signal processing, such as direct enhancement of speech samples. Simultaneously, the proposed apparatus offers high performance, such as improved speech enhancement, or improved quality of the processed audio signal.
[0015] In summary, the concepts described in this paper offer an improved trade-off between computational complexity and achievable audio quality.
[0016] According to an embodiment, the input audio signal is represented by a set of time-domain audio samples, such as noisy time-domain audio, such as speech, or samples, such as time-domain speech utterances. For example, time-domain audio samples of the input audio signal, or time-domain audio samples derived therefrom, are input into a neural network, wherein, for example, the time-domain audio samples of the input audio signal, or time-domain audio samples derived therefrom, are processed in the neural network in the form of a time-domain representation without applying a transformation to the domain representation, such as a spectral domain representation.
[0017] Performing block processing directly in the speech domain (or time domain) allows audio signal processing without any predefined features or time-frequency (TF) transforms. Since both the noise signal and the input audio signal have the same dimension, no upsampling layer is needed during generation. Furthermore, it has been recognized that processing time-domain samples allows for efficient modification of signal statistics within a series of blocks performing invertible affine processing, and also allows for the derivation of audio signals from noise signals within such a series of blocks. It has been found that processing time-domain samples within a series of blocks allows for adaptation of signal characteristics in a way that results in reconstructed audio signals with a good auditory impression. Furthermore, it has been recognized that performing processing in the time domain avoids resource-intensive transformation operations between different signal representation domains. Additionally, it has been recognized that performing block processing directly in the speech domain (or time domain) reduces the number of parameters required for processing blocks using neural networks. Therefore, a computationally less computationally intensive audio signal processing method is provided.
[0018] According to an embodiment, a neural network associated with a given process block in one or more process blocks, such as a given stage of affine processing, is configured to determine one or more processing parameters of the given process block, such as scaling factors, such as S, and shift values, such as T, based on a noise signal, such as z or a signal derived from the noise signal and based on an input audio signal, such as y.
[0019] Using a neural network that also receives and processes time-domain samples of the input audio signal, one or more processing parameters for the affine processing are determined, allowing control over the synthesis of the processed audio signal based on the noise signal from the input audio signal. Therefore, the neural network can be trained in such a way that it provides appropriate processing parameters for the affine processing based on the input audio signal (and typically also based on a portion of the noise signal or a portion of the processed noise signal). Furthermore, it has been recognized that a neural network can be trained with reasonable effort using a training structure that includes the affine processing, which is the opposite of the affine processing used to obtain the processed audio signal.
[0020] According to an embodiment, a neural network associated with a given process block, such as a given stage of affine processing, is configured to provide one or more parameters of the affine processing, such as scaling factors, such as S, and shift values, such as T, which are applied during processing to a noise signal, or a processed version of a noise signal, or a portion of a noise signal, or a portion of a processed version of a noise signal, such as z, for example, in an affine coupling layer.
[0021] By using a neural network to provide one or more parameters for affine processing, and by applying the affine processing to, for example, a noise signal or a processed version of a noise signal, the processing applied to the noise signal is reversible. Therefore, feeding the complete noise signal through the neural network can be avoided, which would typically lead to irreversible operation. However, by using a neural network to control reversible (affine) processing, the training of the neural network can be significantly facilitated, resulting in a more complex and manipulable processing capability.
[0022] According to one embodiment, a neural network associated with a given process block (e.g., a given stage of affine processing) is configured to determine one or more parameters of the affine processing, such as a scaling factor, e.g., S, and a shift value, e.g., T, based on a first portion, such as z1, of the process block input signal, e.g., z, or based on a first portion, of the preprocessed process block input signal, e.g., z', and based on an input audio signal, e.g., y. The affine processing associated with the given process block, e.g., a given stage of affine processing, is configured to apply the determined parameters, such as the scaling factor, e.g., S, and the shift value, e.g., T, to a second portion, such as z2, of the process block input signal, e.g., z, or to a second portion, of the preprocessed process block input signal, e.g., z', to obtain an affine processing signal. The input signals of the process block, such as the first part of z (e.g., z1) or the input signals of the preprocessed process block that have not been modified by affine processing (e.g., the first part of z') and the signals that have undergone affine processing, are as follows: Formation, for example, composition, of a given process block, such as the output signal of a given stage of affine processing, for example, z. new For example, the stage output signal. Given the affine processing of a flow block, such as an affine coupling layer, it ensures that the processed audio signal is generated by reversing the flow block processing used during the training of the neural network.
[0023] According to an embodiment, the neural network associated with a given process block includes depthwise separable convolutions in the affine processing associated with the given process block. For example, the neural network may include depthwise separable convolutions instead of any standard convolutions conventionally used in neural networks. Applying depthwise separable convolutions, for example, instead of any other standard convolutions, can reduce the number of parameters used to process the process block using the neural network. For example, applying depthwise separable convolutions in the neural network and performing process block processing directly in the speech domain (or time domain) can reduce the number of neural network parameters, for example, from 80 million to 20 to 50 million, such as 25 million. Therefore, computationally less demanding audio signal processing is provided.
[0024] According to an embodiment, the device is configured to apply a reversible convolution, such as a 1×1 reversible convolution, to the output signal of a given process block, such as z, at a given stage of an affine process.new For example, stage output signals, to obtain the processed process block output signal z' new For example, a processed version of the output signal of a process block, or a convolutional version of the output signal of a process block, where the output signal of the process block can be the input signal for subsequent stages or other subsequent stages after the first stage. Reversible convolution can help ensure that different samples are processed by affine processing in different process blocks (or processing stages). Additionally, reversible convolution can help ensure that different samples are fed into neural networks of different (subsequent) process blocks. Therefore, the synthesis of processed audio signals based on noise signals can be improved by efficiently altering the statistical properties of a series of time-domain samples.
[0025] According to an embodiment, the device is configured to apply nonlinear compression (e.g., μ-law transformation) to the input audio signal (e.g., y) before processing the noise signal (e.g., z) according to the input audio signal (e.g., y). Regarding the advantages of this function, refer to the following discussion of devices for providing neural network parameters, and in particular, the discussion of the nonlinear compression algorithm used in devices for providing neural network parameters.
[0026] According to an embodiment, the device is configured to apply a μ-law transform (e.g., a μ-law function) as a nonlinear compression to an input audio signal (e.g., y). Regarding the advantages of this function, refer to the following discussion of devices for providing neural network parameters, particularly the use of applying a μ-law transform as a nonlinear compression algorithm in devices for providing neural network parameters.
[0027] According to an embodiment, the device is configured to, based on The transformation is applied to the input audio signal, for example, where sgn() is the sign function and μ is a parameter defining the compression level. For the advantages of this function, refer to the following discussion of devices used to provide neural network parameters, particularly the discussion of applying the same transformation as a nonlinear compression algorithm in devices used to provide neural network parameters.
[0028] According to an embodiment, the device is configured to apply a nonlinear extension (e.g., an inverse μ-law transform, such as a reverting μ-law transform) to a processed (e.g., enhanced) audio signal. This provides efficient post-processing tools, such as efficient post-processing techniques for density estimation, which enhance the results and improve the performance of audio signal processing. As a result, an enhanced signal with minimized high-frequency additives is provided as the output of the device.
[0029] According to an embodiment, the device is configured to apply an inverse μ-law transform (e.g., an inverse μ-law function; e.g., via a reverting μ-law transform) as a nonlinear extension to a processed (e.g., enhanced) audio signal. Using the inverse μ-law transform provides improved modeling of the generation process from a noisy input signal to an enhanced output signal, resulting in improved enhancement performance. This provides efficient post-processing tools, such as efficient post-processing techniques for density estimation, which enhance the enhancement results and improve the performance of audio signal processing.
[0030] According to an embodiment, the device is configured to, based on The transformation is applied to processed (e.g., enhanced) audio signals, such as... Where sgn() is the sign function; μ is a parameter defining the expansion level. This provides enhanced results and improved performance in audio signal processing. It offers efficient post-processing tools, such as efficient post-processing techniques for density estimation, which further enhance the results and improve the performance of audio signal processing.
[0031] According to an embodiment, in one or more training process blocks, processing of a training audio signal or a processed version thereof is used to obtain, for example, predetermined neural network parameters for a neural network used to process noise signals or signals derived from noise signals, stored in a device or a remote server, to obtain a training result signal. This is achieved by adapting the processing of the training audio signal or its processed version using one or more training process blocks according to a distorted version of the training audio signal. The neural network parameters are determined such that the characteristics (e.g., probability distribution) of the training result audio signal approximate or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution). The one or more neural networks used to provide the processed audio signal are identical to the one or more neural networks used to provide the training result signal; wherein the training process blocks perform affine processing opposite to that performed in providing the processed audio signal.
[0032] Therefore, an efficient training tool for neural networks associated with flow blocks is provided, which provides the parameters of the neural network to be used in flow block processing within the device. This results in improved audio signal processing, particularly improved signal enhancement within the device. For example, obtaining neural network parameters in this way allows for efficient training. Inverse processing methods (e.g., defined by inverse affine transformations) can be used in the training and inference of neural network parameters (the acquisition of the processed audio signal), which leads to highly efficient and well-predictable signal transformations. Thus, a good auditory impression can be achieved with feasible complexity.
[0033] According to an embodiment, the apparatus is configured to provide neural network parameters for a neural network used to process a noise signal or a signal derived from a noise signal. The apparatus is configured to process a training audio signal or a processed version thereof using one or more process blocks to obtain a training result signal. The apparatus is configured to adapt the processing of the training audio signal or its processed version performed using one or more process blocks, based on a distorted version of the training audio signal and using a neural network. The apparatus is configured to determine the neural network parameters, for example, using an evaluation of a cost function (e.g., an optimization function); for example, using a parameter optimization procedure, such that the characteristics (e.g., probability distribution) of the training result audio signal approximate or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution). Using a parameter optimization procedure can, for example, reduce the number of neural network parameters, for example, from 80 million to 20 to 50 million, e.g., to 25 million. The apparatus is configured to provide neural network parameters to a neural network associated with a process block used in processing the audio signal in the apparatus. Therefore, the apparatus provides an efficient training tool for the neural network associated with the process block without requiring external training tools.
[0034] According to an embodiment, the apparatus includes means for providing neural network parameters, wherein the means for providing neural network parameters is configured to provide neural network parameters for a neural network that processes a noise signal or a signal derived from a noise signal. The means for providing neural network parameters is configured to process a training audio signal or a processed version thereof using one or more training procedure blocks to obtain a training result signal. The means for providing neural network parameters is configured to adapt the processing of the training audio signal or its processed version performed using one or more procedure blocks based on a distorted version of the training audio signal and using a neural network. The means is configured to determine the neural network parameters, for example, using an evaluation of a cost function (e.g., an optimization function); for example, using a parameter optimization procedure, such that the characteristics (e.g., probability distribution) of the training result audio signal approximate or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution). Therefore, the apparatus includes an efficient training tool for a neural network associated with a procedure block, without requiring external training tools.
[0035] According to an embodiment, one or more process blocks are configured to synthesize a processed audio (e.g., speech) signal based on a noise signal, guided by an input audio (e.g., speech) signal. Therefore, the input audio signal can be used as input to a neural network to control the synthesis of the processed audio signal based on the noise signal. For example, the neural network can efficiently control affine processing to approximate the signal characteristics of the noise signal (or its processed version) to the (statistical) signal characteristics of the input audio signal, wherein the noise contribution of the input audio signal is at least partially reduced. Therefore, an improvement in the signal quality of the processed audio signal compared to the input audio signal can be achieved.
[0036] According to an embodiment, one or more process blocks are configured to synthesize a processed audio (e.g., speech) signal based on the noise signal, guided by an input audio (e.g., speech) signal, using affine processing of sample values of a noise signal or affine processing of sample values of a signal derived from the noise signal. A neural network is used to determine processing parameters for the affine processing, such as scaling factors (e.g., S) and shift values (e.g., T), based on, for example, time-domain sample values of the input audio signal. It has been found that, under reasonable processing loads, this processing yields good results in processing audio signal quality.
[0037] According to an embodiment, the device is configured to perform a normalization process to derive a processed audio signal from a noise signal, for example, guided by an input audio signal. It has been recognized that normalization process provides the ability to successfully generate high-quality samples of processed audio signals in audio enhancement applications.
[0038] An embodiment of the present invention provides a method for providing a processed audio signal based on an input audio signal. The method includes using one or more process blocks to process a noise signal or a signal derived from a noise signal to obtain a processed audio signal. The method includes adapting the processing performed using one or more process blocks to the input audio signal (e.g., a distorted audio signal) and a neural network.
[0039] The method according to this embodiment is based on the same considerations as the apparatus described above for providing processed audio signals. Furthermore, this disclosed embodiment may optionally be supplemented, individually and in combination, by any other features, functions, and details disclosed herein relating to the apparatus for providing processed audio signals.
[0040] An apparatus is created according to embodiments of the invention for providing neural network parameters based on a portion of a clean audio signal (e.g., x1) or a processed version thereof, and based on a distorted audio signal (e.g., y) in a training mode. This includes, for example, providing edge weights (e.g., θ) of the neural network with scaling factors (e.g., s) and shift values (e.g., t), which may correspond to edge weights of the neural network in an inference mode based on a portion of a noise signal (e.g., z), or a processed version thereof, and based on the input audio signal (e.g., y) for audio processing (e.g., speech processing). The apparatus is configured, for example, in multiple iterations, to process the training audio signal (e.g., speech signal, e.g., x) or a processed version thereof using one or more process blocks (e.g., eight process blocks, e.g., using a process block system (e.g., including affine coupling layers, e.g., including reversible convolutions)) to obtain a training result signal, which should be equal to, for example, the noise signal. The device is configured to adapt a distorted version of a training audio signal (e.g., y, for example, a distorted audio signal, such as a noisy speech signal y) and a neural network to a process performed using one or more flow blocks. For example, the neural network is based on the distorted version of the training audio signal and preferably also provides one or more processing parameters to the flow blocks based on at least a portion of the training audio signal or its processed version, such as parameters for affine processing, like scaling factors and shift values. The device is configured to determine the neural network parameters, for example, using an evaluation of a cost function (e.g., an optimization function), such as using a parameter optimization procedure performed by the neural network, such that the characteristics (e.g., probability distribution) of the training audio signal approximate or include predetermined characteristics (e.g., noise-like characteristics, such as a Gaussian distribution).
[0041] This embodiment is based on the finding that block processing can be applied to audio signal processing, particularly by learning a mapping from simple to more complex probability distributions conditioned on clean speech samples (e.g., x) and their noisy counterparts (e.g., y), such as learning the probability distribution of clean speech, to determine the neural network parameters to be used in audio signal processing. For example, it has been found that the parameters of a neural network associated with a series of blocks are useful in inference, i.e., when a processed audio signal is obtained based on a noisy signal. Furthermore, it has been found that inference blocks corresponding to training blocks can be easily designed, and the inference blocks can be controlled using a neural network employing the trained neural network parameters.
[0042] The proposed device provides an efficient training tool for neural networks associated with process blocks, which provides parameters for the neural network used in audio signal processing. This leads to the use of neural networks with the determined neural network parameters to improve audio signal processing, particularly improved audio signal enhancement, which provides high performance, such as improved speech enhancement, or, for example, improved quality of the processed audio signal.
[0043] According to an embodiment, the apparatus is configured to evaluate a cost function (e.g., a loss function) based on the characteristics of the obtained training result signal (e.g., based on the distribution of the obtained noise signal (e.g., a Gaussian distribution)), the variance δ² of the obtained noise signal, and processing parameters of the process blocks (e.g., scaling factors, such as s), which may depend on the input signals of each process block. The apparatus is configured to determine neural network parameters to reduce or minimize the cost defined by the cost function. The correlation between the modeling generation process and the generation process represented by the modeling is optimized. Furthermore, the cost function helps to adjust the neural network parameters in such a way that the processing in a series of process blocks controlled by the neural network transforms the training audio signal into a signal with desired statistical properties (e.g., a noise-like signal). The deviation between the desired statistical properties and the signal provided by the training process blocks can be efficiently represented by the cost function. Therefore, the neural network parameters can be trained or optimized in such a way that the processing in the training process blocks provides a signal whose statistical properties approximate the desired (e.g., noise-like) properties. In this training, the cost function can be a simple (and computationally efficient) training objective function, thus facilitating the adaptation of neural network parameters. The trained neural network parameters, trained in this way, can then be used for inference processing to synthesize processed audio signals based on noise signals.
[0044] According to an embodiment, the training audio signal (e.g., x) and / or a distorted version of the training audio signal (e.g., y) are represented by a set of time-domain audio samples (e.g., noisy time-domain audio (e.g., speech, samples, e.g., time-domain speech utterance)). For example, time-domain audio samples of the input audio signal or time-domain audio samples derived therefrom are input into a neural network. For example, time-domain audio samples of the training audio signal or time-domain audio samples derived therefrom are processed in the neural network in the form of a time-domain representation without applying a transform to the transform domain representation, such as a spectral domain representation. Performing block processing directly in the speech domain (or time domain) allows audio signal processing without any predefined features or time-frequency TF transforms. When both the training audio signal and the distorted version of the training audio signal have the same dimension, no upsampling layer is required in the processing. Furthermore, the advantages of time-domain processing discussed above are also referenced.
[0045] According to an embodiment, a neural network (e.g., a given stage of affine processing) associated with a given process block in one or more process blocks is configured to determine one or more processing parameters (e.g., scaling factors (e.g., s) and shift values (e.g., t)) for the given process block based on a training audio signal (e.g., x) or a signal derived from the training audio signal and based on a distorted version of the training audio signal (e.g., y). Regarding the advantages of this functionality, further reference is made to the above discussion of the means for providing the processed audio signal.
[0046] According to an embodiment, a neural network associated with a given process block (e.g., a given stage of affine processing) is configured to provide one or more parameters of the affine processing (e.g., scaling factors (e.g., s) and shift values (e.g., t) in an affine coupling layer), which are applied during processing to a training audio signal (e.g., x), or a processed version of the training audio signal, or a portion of the training audio signal, or a portion of a processed version of the training audio. Regarding the advantages of this functionality, further reference is made to the above discussion of the means for providing the processed audio signal.
[0047] According to an embodiment, a neural network associated with a given process block (e.g., a given stage of affine processing) is configured to determine one or more parameters of the affine processing (e.g., scaling factors (e.g., s) and shift values (e.g., t)) based on a first portion (e.g., x1) of the process block input signal (e.g., x) or based on a first portion of the preprocessed process block input signal (e.g., x') and a distorted version (e.g., y) of the training audio signal. The affine processing associated with the given process block (e.g., a given stage of affine processing) is configured to apply the determined parameters to a second portion (e.g., x2) of the process block input signal x or a second portion of the preprocessed process block input signal x' to obtain an affine processed signal, such as... For example, the first part (e.g., x1) of the flow block input signal (e.g., x) that has not been modified by affine processing, or the first part of the preprocessed flow block input signal (e.g., x') and the affine processing signal (e.g., x'). The process block output signal x forms (e.g., constitutes) a given process block (e.g., a given stage of affine processing). new(e.g., stage output signal). Affine processing of a given process block (e.g., affine coupling layer) ensures the reversibility of the process block processing and efficient computation of the characteristics of the training result audio signal, such as the Jacobian determinant used in determining the probability density function. Furthermore, reversibility is achieved by performing affine processing on only a portion of the process block input signal while keeping another portion of the process block input signal unchanged, while still allowing a portion of the process block input signal to be fed into the neural network. Since the portion of the process block input signal used as input to the neural network is unaffected by the affine processing, it is available before and after the affine processing, which, if the affine processing is reversible (which is usually the case), allows for reversal of the processing direction (when transitioning from the training phase to the inference phase). Therefore, the learning neural network coefficients learned during training are highly significant in the inference phase.
[0048] According to an embodiment, the neural network associated with a given process block includes depthwise separable convolutions in the affine processing associated with the given process block. For example, the neural network may include depthwise separable convolutions instead of any standard convolutions conventionally used in neural networks. Applying depthwise separable convolutions (e.g., instead of any other standard convolutions) can reduce the number of parameters used to process the process block using the neural network. For example, applying depthwise separable convolutions to the neural network and performing process block processing directly in the speech domain (or time domain) can reduce the number of neural network parameters, for example, from 80 million to 20-50 million, such as 25 million. Therefore, computationally less demanding audio signal processing is provided.
[0049] According to an embodiment, the apparatus is configured to apply a reversible convolution (e.g., a 1×1 reversible convolution) to a given process block input signal (e.g., x, e.g., a stage input signal for a given stage of affine processing). This process block input signal may be, for example, a training audio signal or a signal derived from the training audio signal of a first stage, and may be, for example, the output signal of a previous stage. For subsequent stages after the first stage, a preprocessed process block input signal, e.g., x', is obtained, a preprocessed version of the process block input signal, e.g., a convolved version of the process block input signal. Regarding the advantages of this functionality, refer to the above discussion of the apparatus for providing the processed audio signal.
[0050] According to an embodiment, the device is configured to apply nonlinear input compression (e.g., nonlinear compression, such as μ-law transformation) to the training audio signal (e.g., x) before processing it. The nonlinear compression algorithm is applied to map small amplitudes of the audio data samples to wider, larger amplitudes to smaller intervals. This addresses the problem that higher absolute amplitudes are underrepresented in clean data samples. This provides efficient preprocessing tools, such as efficient preprocessing techniques for density estimation, which use neural networks with defined neural network parameters to enhance the results and improve the performance of audio signal processing. The advantages of this functionality are also discussed above in the description of the device for providing the processed audio signal. For example, nonlinear input compression can be the opposite of the nonlinear expansion discussed above.
[0051] According to an embodiment, the device is configured to apply a μ-law transform (e.g., a μ-law function) as nonlinear input compression to a training audio signal (e.g., x). The distribution of the compressed signal is learned, rather than the distribution of the clean signal. Stream processing using the μ-law transform can capture finer-grained speech portions with less background leakage. Therefore, improved enhancement performance is provided when using a neural network with defined neural network parameters for audio signal processing. The advantages of this feature are also discussed above in the description of the device for providing the processed audio signal. This μ-law transform can, for example, be (at least approximately) the reverse of the transform discussed above for the device for providing the processed audio signal. This provides efficient preprocessing tools, such as efficient preprocessing techniques for density estimation, which use a neural network with defined neural network parameters to enhance the results and improve the performance of audio signal processing.
[0052] According to an embodiment, the device is configured to, based on The transformation is applied to the training audio signal (x), where sgn() is the sign function; μ is a parameter defining the compression level. This provides an enhancement to the audio signal processing results and improved performance using a neural network with defined neural network parameters. For further information on the advantages of this function, refer also to the discussion above regarding the apparatus for providing processed audio signals. For example, this transformation can be (at least approximately) the opposite of the transformation discussed above for the apparatus for providing processed audio signals. This provides efficient preprocessing tools, such as efficient preprocessing techniques for density estimation, which enhance the audio signal processing results and improved performance using a neural network with defined neural network parameters.
[0053] According to an embodiment, the device is configured to apply nonlinear input compression (e.g., μ-law transformation) to a distorted version of the training audio signal (e.g., y) before processing the training audio signal (e.g., x). For the advantages of this function, refer to the discussion above on nonlinear compression algorithms for processing training audio signals (e.g., x).
[0054] According to an embodiment, the device is configured to apply a μ-law transform (e.g., a μ-law function) as a nonlinear input compression to a distorted version (e.g., y) of a training audio signal. For the advantages of this functionality, refer to the discussion above regarding the application of the μ-law transform as a nonlinear compression algorithm to process a training audio signal (e.g., x).
[0055] According to an embodiment, the device is configured to, based on The transformation is applied to a distorted version of the training audio signal, for example, where sgn() is the sign function; μ is a parameter defining the compression level. For the advantages of this function, refer to the discussion above on applying the same transformation as a nonlinear compression algorithm to process training audio signals (e.g., x).
[0056] According to an embodiment, one or more process blocks are configured to convert training audio signals into training result signals that approximate noise signals or include noise-like characteristics. It has been found that neural networks associated with process blocks and trained to convert training audio signals into noise signals (or at least into noise-like signals) can be well used for speech enhancement (e.g., using "inverse" inference process blocks to perform functions substantially opposite to those of the training process blocks).
[0057] According to embodiments, for example, one or more process blocks are adjusted by appropriately determining neural network parameters to convert a training audio signal into a training result signal, guided by a distorted version of the training audio signal (e.g., a speech signal), using affine processing of sample values of the training audio signal or affine processing of sample values of a signal derived from the training audio signal. The neural network determines processing parameters for the affine processing, such as scaling factors (e.g., s) and shift values (e.g., t), based on time-domain sample values of, for example, a distorted version of the training audio signal. It has been found that neural networks for adjusting one or more process blocks (e.g., by providing scaling and / or shift values) are well-suited for audio enhancement in inference devices (e.g., devices for providing processed audio signals discussed herein).
[0058] According to an embodiment, the device is configured to perform a normalization process to derive a training result signal from the training audio signal, for example, guided by a distorted version of the training audio signal. The normalization process provides the ability to successfully generate high-quality samples of the training result signal. Furthermore, it has been found that the normalization process, using neural network parameters obtained through training, provides good results for speech enhancement.
[0059] An embodiment of the invention provides a method for providing neural network parameters based on a portion of a clean audio signal (e.g., x1) or a processed version thereof, and a distorted audio signal (e.g., y) in a training mode. This includes, for example, providing edge weights (e.g., θ) of the neural network with scaling factors (e.g., s) and shift values (e.g., t), which may correspond to edge weights of the neural network in an inference mode based on a portion of a noisy signal (e.g., z), or a processed version thereof, and based on the input audio signal (e.g., y) for audio processing (e.g., speech processing). The method includes, for example, processing the training audio signal (e.g., a speech signal, e.g., x) or a processed version thereof using one or more process blocks (e.g., using a process block system (e.g., including affine coupling layers, e.g., including reversible convolutions)) in multiple iterations to obtain a training result signal that should be equal to, for example, the noisy signal (e.g., z). The method also includes adapting the processing performed using one or more process blocks to a distorted version of the training audio signal (e.g., y, e.g., a distorted audio signal, e.g., based on a noisy speech signal y) and using a neural network. For example, the neural network provides one or more processing parameters, such as parameters for affine processing, like scaling factors and shift values, to the process block based on a distorted version of the training audio signal, and preferably also based on at least a portion of the training audio signal or its processed version. The method includes, for example, using an evaluation of a cost function, or using a parameter optimization procedure to determine the neural network parameters such that the characteristics (e.g., probability distribution) of the training audio signal approximate or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution).
[0060] The method according to this embodiment is based on the same considerations as the apparatus described above for providing neural network parameters. Furthermore, this disclosed embodiment may optionally be supplemented, individually and in combination, by any other features, functions, and details disclosed herein relating to the apparatus for providing neural network parameters.
[0061] According to embodiments of the present invention, a computer program with program code is created, which, when run on a computer, performs the method according to any of the above embodiments.
[0062] The apparatus for providing processed audio signals, the method for providing processed audio signals, the apparatus for providing neural network parameters, the method for providing neural network parameters, and the computer program for implementing these methods may optionally be supplemented individually and in combination by any features, functions, and details disclosed herein (throughout the document). Attached Figure Description
[0063] The preferred embodiments of this application are described below based on the accompanying drawings, wherein...
[0064] Figure 1 A schematic representation of an apparatus for providing a processed signal according to an embodiment is shown;
[0065] Figure 2 A schematic representation of an apparatus for providing a processed signal according to an embodiment is shown;
[0066] Figure 3 A schematic representation of an inference flow block for an apparatus for providing a processed signal according to an embodiment is shown;
[0067] Figure 4 A schematic representation of an apparatus for providing a processed signal according to an embodiment is shown;
[0068] Figure 5 A schematic representation of an apparatus for providing neural network parameters according to an embodiment is shown;
[0069] Figure 6 A schematic representation of an apparatus for providing neural network parameters according to an embodiment is shown;
[0070] Figure 7 A schematic representation of a training process block of an apparatus for providing neural network parameters according to an embodiment is shown;
[0071] Figure 8 A schematic representation of an apparatus for providing neural network parameters according to an embodiment is shown;
[0072] Figure 9 The illustration shows a nonlinear input compressive diffraction (compression and expansion) provided in an apparatus for providing processed signals according to an embodiment or in an apparatus for providing neural network parameters according to an embodiment;
[0073] Figure 10 A flow block system for audio signal processing according to an embodiment is shown;
[0074] Figure 11 A table is shown, illustrating a comparison between the apparatus and method according to the embodiments and conventional techniques;
[0075] Figure 12 shows a graphical representation of the performance of the apparatus and method according to the embodiments. Detailed Implementation
[0076] Figure 1 A schematic representation of an apparatus 100 for providing a processed audio signal according to an embodiment is shown.
[0077] The device 100 is configured to provide a processed (e.g., enhanced) audio signal 160 based on the input audio signals y, 130. For example, in N process blocks (e.g., inference process block 110) associated with a neural network (not shown). 1...N Processing is performed within (process block 110). 1...N It is configured to process incoming audio signals, such as speech signals.
[0078] The input audio signals y and 130 are introduced into device 100 for processing. For example, the input audio signal y is a noisy input signal, or for example, a distorted audio signal. For example, the input audio signal y and 130 can be defined as y = x + n, where x is the clean portion of the input signal and n is the noisy background. For example, the input audio signal y and 130 can be represented as a time-domain audio sample, such as a noisy time-domain speech sample.
[0079] The input audio signals y and 130 can optionally be preprocessed, such as compressed, for example... Figure 4 As shown, for example, through nonlinear compression, such as in the reference Figure 9 The nonlinear compression described.
[0080] The input audio signal y and its corresponding clean portion x can be optionally grouped into vector representations (or matrix representations).
[0081] Its noise signal z, 120 (or preprocessed version z (i=1)) is introduced into the first process block 1101 of the device 100 together with the input audio signal y, 130.
[0082] The noise signals z and 120 are generated, for example, at device 100, or generated externally and provided to device 100. The noise signals z and 120 may be stored in device 100, or provided to the device from external memory (e.g., a remote server). For example, the noise signals z and 120 are defined as samples from a normal distribution with zero mean and unit variance (e.g., z ~ N(z; 0; 1)). For example, the noise signals z and 120 are represented as noise samples, such as time-domain noise samples.
[0083] The signal z can be preprocessed into a noise signal z (i=1) before entering the device 100 or before entering the device 100.
[0084] For example, noise samples z of noise signal z or noise samples of preprocessed noise signal z (i=1) can be optionally grouped into sample groups, such as into a group of 8 samples, or into a vector representation (or matrix representation).
[0085] Figure 1 Optional preprocessing steps are not shown.
[0086] Noise signals z (i=1), 1401 (or, alternatively, noise signals z) are introduced into the first process block 1101 (e.g., inference process block) of device 100 along with input audio signals y, 130. (Refer to...) Figure 2 and Figure 3 Further description in the first process block 1101 and process block 110 1...N The subsequent process blocks handle the noise signal z(i=1), 1401, and the input audio signal y. The input signal z(i) is processed in process block 110. 1...N (Or, usually, 110) i The input audio signal y, 130 is adjusted, for example, in process block 110. 1...N In each process block.
[0087] After processing the noise signals z (i=1) and 1401 in the first process block 1101, the output signal z is... new (i=1), 1501 is output. Signal z new (i=1) and 1501 are the input signals z (i=2), 1402, and input audio signals y and 130 used for the second process block 1102 of device 100. new (i=2), 1502 are the input signals z (i=3) for the third process block, etc. The last N process blocks 110 N Having a signal z (i = N) as the input signal, 140 N And output signal z of the output signal 160 of the forming device 100. new (i = N), 150 N Signal z new (i = N), 150 N Forming processed audio signals 160, for example, an enhanced audio signal, the processed audio signal 160 represents, for example, the clean portion of the input audio signal y, 130.
[0088] The clean portion x of the input audio signals y, 130 is not separately introduced into device 100. Device 100 processes, for example, generated noise signals z, 120 based on the input audio signals y, 130 to receive, for example, generate, for example, an enhanced audio signal, which is, for example, an enhancement of the clean portion of the input audio signals y, 130.
[0089] Generally speaking, the device can be described as being configured to use one or more process blocks 1101 to 110. N The device 100 processes a noise signal (e.g., noise signal z) or a signal derived from a noise signal (e.g., a pre-processed noise signal z(i=1)) to obtain a processed (e.g., enhanced) audio signal 160. Generally, the device 100 is configured to adapt one or more process blocks 1101 to 110 to the process block based on an input audio signal (e.g., a distorted audio signal y) and using a neural network (e.g., the neural network may be based on the distorted audio signal, and preferably also on at least a portion of the noise signal or its processed version, providing one or more processing parameters, such as parameters of affine processing, like scaling factors and shift values) to the process block. N The execution process.
[0090] However, it should be noted that the device 100 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0091] Figure 2 A schematic representation of a device 200 for providing a processed signal according to an embodiment is shown.
[0092] In an embodiment, Figure 1 The features, functions and details of the device 100 shown may optionally be incorporated into the device 200 (both individually and in combination), and vice versa.
[0093] Device 200 is configured to provide a processed (e.g., enhanced) audio signal based on input audio signals y, 230. 260. In the N process blocks associated with the neural network (not shown) (e.g., inference process block 210) 1...N Processing is performed within (process block 210). 1...N It is configured to process incoming audio signals, such as speech signals.
[0094] The input audio signals y and 230 are introduced into the device 200 for processing. For example, the input audio signal y is a noisy input signal, or for example, a distorted audio signal. For example, the input audio signal y and 230 is defined as y = x + n, where x is the clean portion of the input signal and n is the noisy background. For example, the input audio signal y and 230 can be represented as a time-domain audio sample, such as a noisy time-domain speech sample.
[0095] The input audio signals y and 230 can optionally be preprocessed, such as compressed, for example... Figure 4 As shown, for example, through nonlinear compression, such as in the reference Figure 9 The nonlinear compression described.
[0096] The input audio signal y and its corresponding clean portion x can be optionally grouped into vector representations (or matrix representations).
[0097] The noise signal z, 220 (or a pre-processed version z(i=1)) is introduced together with the input audio signal y, 230 into the first process block 2101 of device 200. For example, the noise signal z, 220 may be generated at device 200, or generated externally and provided to device 200, for example. The noise signal z may be stored in device 200, or provided to the device from external memory (e.g., a remote server). For example, the noise signal z, 220 may be defined as being sampled from a normal distribution with zero mean and unit variance (e.g., z ~ N(z; 0; I)). For example, the noise signal z, 220 may be represented as a noise sample, such as a time-domain noise sample.
[0098] Signals z and 220 can be preprocessed before being introduced into the device 200. For example, noise samples of noise signals z and 220 can be optionally grouped into sample groups, such as into groups of 8 samples, or into vector representations (or matrix representations). Figure 2 Optional preprocessing steps are not shown.
[0099] The noise signals z(i=1) and 2401 are introduced together with the input audio signals y and 230 into the first process block 2101 (e.g., the inference process block) of the device 200. The noise signals z(i) and 2401 are in process block 110. 1...N The input audio signals y and 230 are adjusted, for example, by the input audio signals y and 230. The input audio signals y and 230 are introduced into process block 110. 1...N In each process block.
[0100] The processing in the first process block 2101 is performed in two steps, for example, in two blocks (or using two function blocks), for example in two operations: an affine coupling layer 2111 and an optional 1×1 reversible convolution 2121.
[0101] In the affine coupling layer block 2111, the noise signal z (i=1), 2401 is processed based on, for example, the input audio signal y, 230, and the input audio signal y, 230 is introduced into the affine coupling layer block 2111 of the first process block 2101. (Refer to...) Figure 3Further description of the affine coupling layer block 2111 in the first process block 2101 and in process block 210 1...N Affine coupling layer block 211 of the subsequent process blocks 1...N An example of processing noise signals z (i=1), 2401 and input audio signals y, 230. After processing in the affine coupling layer block 2111 of the first process block 2101, the output signal z new (i=1), 2501 is output.
[0102] In the reversible convolution block 2121, the mixed output signal z new (i=1), 2501 samples are used to receive the processed process block output signal z' new (i = 1). For example, the reversible convolution block 2121 inverts (or, typically, changes) the channel ordering at the output of the affine coupling layer block 2111. For example, the reversible convolution block 2121 can be performed using a weight matrix W (e.g., as a random (or pseudo-random but deterministic) rotation matrix or as a random (or pseudo-random but deterministic) permutation matrix). The first process block 2101 provides the output signal z. new (i=1) or the processed process block output signal z' new (i=1) serves as the output process block signal 2511, which corresponds to the input signals z (i=2), 2402, and the input audio signals y and 230 for the second process block 2102 of the device 200. new (i=2), 2502 are the input signals z (i=3) for the third process block, etc. The last N process blocks 210 N Having a signal z (i = N) as the input signal, 240 N And output signal z of the output signal 260 of the forming device 200. new (i = N), 250 N Signal z new (i=N), 250N forms the processed audio signal 260, for example, enhanced audio signals, processed audio signals 260 represents, for example, the enhanced clean portion of the input audio signal y, 230.
[0103] Process block 210 1...N The processing in all subsequent process blocks can be performed in two steps, for example, in two blocks, or in two operations: an affine coupling layer and a 1×1 reversible convolution. These two steps can be, for example, the same as the description associated with the first process block 2101 (where, for example, different neural network parameters can be used in different process blocks).
[0104] Process block 210 1...N Affine coupling block 211 in 1...N Associated with (or including) a corresponding neural network (not shown), which is related to process block 210 as described above. 1...N Related. For example, in reference Figures 5 to 8 During network training, the parameters of the described device (or function) are predetermined.
[0105] The clean portion x of the input audio signals y, 230 is not separately introduced into device 200. Device 200 processes, for example, generated noise signals z, 220 based on the input audio signals y, 230 to receive, for example, generate, for example, an enhanced audio signal, which is, for example, an enhancement of the clean portion of the input audio signals y, 230.
[0106] However, it should be noted that the device 200 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0107] Figure 3 A schematic representation of a process block 311 (e.g., an inference process block) according to an embodiment is shown.
[0108] Process block 311 can be, for example, by Figure 1 The device 100 shown or Figure 2 This is part of the processing of the device 200 shown. Figure 1 The process blocks of the apparatus 100 shown can have the same as Figure 3 The process block 311 shown has the same structure, or may include the functions (and / or structures) (e.g., additional functions) of the process block 311. Figure 2 The affine coupling layer of the process block of the device 200 shown can have a connection with... Figure 3 The process block 311 shown has the same structure as the process block 311, or may include the functions (and / or structure) of the process block 311 (e.g., together with additional functions).
[0109] For the sake of simplicity, in Figure 3 The process block index i is partially omitted in the following description.
[0110] Introduce input signal 340 into the process block. For example, as follows: Figure 1 As illustrated in the embodiments shown, the input signal 340 can represent a noise signal (or a processed version thereof) z(i). For example, the input signal 340 can be represented as a time-domain sample. The input signal 340 can optionally be grouped into a vector representation (or a matrix representation).
[0111] The input signal 340 is divided into two parts z1(i) and z2(i) (e.g., into two subsequent parts) in a random, pseudo-random but deterministic, or predetermined manner, for example.
[0112] The first part z1(i) (e.g., which may include a subset of time-domain samples of the input signal 340) is introduced into the neural network 380 (also designated NN(i)) associated with process block 311 (having process block index i). For example, neural network 380 may be associated with... Figure 1 Any process block 110 of the apparatus 100 shown 1...N The associated neural network. For example, neural network 380 could be associated with... Figure 2 Process block 210 of the apparatus 200 shown 1...N Any affine coupling layer associated with a neural network. For example, the parameters of neural network 380 can be, for example, determined by a reference during network training. Figures 5 to 8 The described apparatus is predetermined.
[0113] The first part z1(i) is introduced into the neural network 380 along with the input audio signal y, 330. For example, the input audio signal y, 330 is a noisy input signal, or for example, a distorted audio signal. For example, the input audio signal y, 330 is defined as y = x + n, where x is the clean part of the input audio signal y, 330 and n is the noisy background.
[0114] The input audio signals y and 330 can optionally be preprocessed, such as compressed, for example... Figure 4 As shown, for example, through nonlinear compression, such as in the reference Figure 9 The nonlinear compression described.
[0115] The input audio signal y and its corresponding clean portion x can be optionally grouped into vector representations (or matrix representations).
[0116] Neural network 380 processes the first part z1(i) and the input audio signal y, 330, for example, according to, for example, the input audio signal y, 330. Neural network 380 determines processing parameters, such as scaling factors (e.g., S) and shift values (e.g., T), which are the output (371) of neural network 380. The determined parameters S, T have, for example, vector representations. For example, different scaling and / or shift values can be associated with different samples of the second part z2(i). The determined parameters S, T are used to process (372) the second part z2(i) of the noise signal z (which may, for example, include a subset of the time-domain samples of the input signal 340). The processed (affine) second part (i) Defined by the following formula:
[0117]
[0118] In this equation, s can be equal to S (e.g., if the neural network provides only a single scaling factor value), or s can be an element of a vector of scaling factor values S (e.g., if the neural network provides a vector of scaling factor values). Similarly, t can be equal to T (e.g., if the neural network provides only a single shift value), or t can be an element of a vector of shift values T (e.g., if the neural network provides a vector of scaling factor values, each of which is associated with different sample values of z2(i)).
[0119] For example, the above is used for The equation can be applied element-wise to a single element or group of elements in the second part z2. However, if the neural network provides only a single value s and a single value t, then that single value s and that single value t can be applied to all elements of the second part z2 in the same way.
[0120] The unprocessed first part z1(i) of signal z and the processed part of signal z The signals are combined (373) to form signal z, which is processed at process block 311. new 350. The output signal z new Introduced into the next process, such as a subsequent or following process block, for example, into the second process block, or into process block i plus 1. If i = N, signal z new 350 is the output signal of the corresponding device, for example
[0121] However, it should be noted that process block 311 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0122] Furthermore, process block 311 may optionally be used in any of the embodiments disclosed herein.
[0123] Figure 4 A schematic representation of a device 400 for providing a processed signal according to an embodiment is shown.
[0124] In an embodiment, Figure 1 The device 100 shown or Figure 2 The features, functions and details of the device 200 shown may be optionally incorporated into the device 400 (either individually or in combination), or vice versa.
[0125] In an embodiment, Figure 3 The process block 311 shown can be used, for example, in device 400.
[0126] The device 400 is configured to provide a processed (e.g., enhanced) audio signal based on the input audio signals y, 430. This is achieved through N process blocks (e.g., inference process block 410) associated with a neural network (not shown). 1...N Processing is performed within (process block 410). 1...N It is configured to process incoming audio signals, such as speech signals.
[0127] The input audio signals y and 430 are introduced into the device 400 for processing. For example, the input audio signals y and 430 are noisy input signals, or, for example, distorted audio signals. For example, the input audio signal y is defined as y = x + n, where x is the clean portion of the input signal and n is the noisy background. For example, the input audio signal y and 430 can be represented as a time-domain audio sample, such as a noisy time-domain speech sample.
[0128] The input audio signals y and 430 can optionally be preprocessed, for example, compressed, such as by non-linear compression 490.
[0129] The nonlinear compression step 490 can optionally be applied to the input audio signals y and 430. For example... Figure 4 As shown, step 490 is optional. For example, nonlinear compression step 490 can be applied to compress the input audio signals y, 430. In an embodiment, nonlinear input compression step 490 is as described in reference... Figure 9 As described.
[0130] In an embodiment, the nonlinear compression 490 can be represented, for example, by μ-law compression of the input audio signal y, 430, or by, for example, μ-law transformation. For example:
[0131]
[0132] Where sgn() is a symbolic function;
[0133] μ is a parameter that defines the compression level.
[0134] For example, the parameter μ can be set to 255, a common value used in telecommunications. The input audio signal y and its corresponding clean portion x can optionally be grouped into a vector representation (or matrix representation).
[0135] For example, the noise signal z, 420 is an input signal of device 400, or may alternatively be generated by device 400. Before introducing the noise signal z, 420 into the first process block 4101 of device 400, the audio samples of the noise signal z are grouped (e.g., in grouping block 405) into sample groups, for example, into groups of 8 samples, or into vector representations (or matrix representations). Figure 4 As shown, grouping step 405 is an optional step.
[0136] The noise signals z(i), 4401 (optionally grouped) are introduced into the first process block 4101 of the device 400 together with the input audio signals y, 430 or together with the preprocessed (e.g., compressed) input audio signals y'. For example, the noise signals z, 420 are generated at (or by) the device 400, or generated externally and provided to the device 400, for example. The noise signals z may be stored in the device 400 or provided to the device from external memory (e.g., a remote server). For example, the noise signals z, 420 are defined as samples from a normal distribution (or Gaussian distribution) with zero mean and unit variance (e.g., z ~ N(z; 0; 1)). For example, the noise signals z, 420 are represented as noise samples, such as time-domain noise samples.
[0137] The noise signal z(i) 4401 (optionally grouped) is introduced into the first process block 4101 of the device 400 along with the input audio signals y and 430. The noise signal z(i) is introduced into process block 4101. 1...N The input audio signal y is adjusted (or further processed) based on, for example, the input audio signal y, 430. For example, in process block 410 1...N The input audio signal y, 430 is introduced into each process block.
[0138] The processing in the first process block 24101 is performed in two steps, for example in two blocks, for example in two operations: affine coupling layer 4111 and optional 1×1 reversible convolution 4121.
[0139] In affine coupling layer block 4111, noise signals z (i=1), 4401 are processed based on, for example, the input audio signals y, 430, and the input audio signals y, 430 are introduced into the affine coupling layer block 4111 of the first process block 4101. It should be noted that the affine coupling layer block may include, for example, a single affine coupling layer or multiple affine coupling layers. (See reference...) Figure 3 The affine coupling layer block 4111 described in the first process block 4101 and in process block 410 1...N Subsequent process block 410 2...N Affine coupling layer 411 2...N The processing of noise signals z (i=1), 4401 and input audio signals y, 430. After processing in the affine coupling layer block 4111 of the first process block 4101, the output signal z new (i=1), 4501 is output.
[0140] In the reversible convolution block 4121, the output signal z is mixed (e.g., reordered, or subjected to invertible matrix operations such as rotation). new(i=1), 4501 samples are used to receive the processed process block output signal z' new (i=1). For example, the reversible convolution block 4121 reverses the order of channels (or samples) at the output of the affine coupling layer block 4111. For example, the reversible convolution block 4121 can be performed using a weight matrix W (e.g., as a random (or pseudo-random but deterministic) rotation matrix or as a random (or pseudo-random but deterministic) permutation matrix).
[0141] The first process block 4101 provides the output signal z new (i=1) or the processed process block output signal z' new (i=1) serves as the output process block signal 4511, which corresponds to the input signals z (i=2), 4402, and input audio signals y and 430 for the second process block 4102 of device 400. The output signal z of the second process block 4102... new (i=2) or z' new (i=2), 4502 are the input signals z (i=3) for the third process block, etc. The last N process blocks are 410. N Having a signal z (i = N) as the input signal, 440 N And output signal z of the forming device 400 output signal 460. new (i = N) or z' new (i = N), 450 N Signal z new (i = N) or z' new (i = N), 450 N Forming processed audio signals 460, for example, enhanced audio signals, processed audio signals 460 represents, for example, the enhanced clean portion of the input audio signal y, 430. In an embodiment, the processed audio signal 460 is, for example, the output signal of device 400.
[0142] For example, process block 410 1...N The processing in all subsequent process blocks is performed in two steps, for example, in two blocks, or in two operations: an affine coupling layer and a 1×1 reversible convolution. These two steps are, for example (e.g., qualitatively), the same as those described in relation to the first process block 4101. However, different neural network coefficients used to determine the scaling and shift values can be used in different processing stages. Furthermore, the reversible convolution may also be different in different stages (but may also be equal in different stages).
[0143] Process block 410 1...NThe affine coupling layer block is associated with the corresponding neural network (not shown), and the neural network is associated with process block 410 as indicated. 1...N Related.
[0144] The nonlinear extension step 415 may optionally be applied to the processed audio signal. 460. (For example) Figure 4 As shown, step 415 is optional. The nonlinear extension step 415 can be applied, for example, to the processed audio signal. 460 is extended to a conventional signal. In an embodiment, the non-linear extension may be, for example, a processed audio signal. It can be represented by the inverse μ-law transform of 460. For example:
[0145]
[0146] Where sgn() is a symbolic function;
[0147] μ is a parameter that defines the compression level.
[0148] The parameter μ can be set to, for example, 255, which is a common value used in telecommunications. For example, when nonlinear compression is used with process block 410... 1...N During the preprocessing steps of training the associated neural network, a nonlinear extension step 415 can be applied.
[0149] It should be noted that the clean portion x of the input audio signal y, 430 is not separately introduced into the device 400. The device 400 processes, for example, the generated noise signal z, 420 based on the input audio signal y, 430 to receive, for example, generate, for example, an enhanced audio signal, which is, for example, an enhancement of the clean portion of the input audio signal y, 430.
[0150] However, it should be noted that the device 400 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0151] Figure 5 A schematic representation of a device 500 for providing neural network parameters according to an embodiment is shown.
[0152] Device 500 is configured to provide neural network parameters (e.g., for use with process block 110) based on a distorted version of training audio signals x, 505 (e.g., a clean audio signal) and training audio signals y, 530 (e.g., a distorted audio signal). 1…N 210 1…N 410 1…N The associated neural network 380, NN(i) is used. For example, in conjunction with neural network 580... 1...NN related process blocks (e.g., training process block 510) 1...N The processing is performed within (e.g., training block 510). 1...N It is configured to process incoming audio signals, such as speech signals.
[0153] A distorted version of the training audio signal y, 530 is introduced into device 500 for processing (or generated by device 500). The distorted audio signal y is, for example, a noisy input signal. For example, the distorted training audio signal y, 530 can be defined as y = x + n, where x is the clean portion of the input signal, such as the training input signal x, 505, and n is the noisy background. The distorted training audio signal y, 530 can be represented, for example, as a time-domain audio sample, such as a noisy time-domain speech sample.
[0154] The training audio signal x and the distorted version of the corresponding training audio signal y can be optionally grouped into vector representations (or matrix representations).
[0155] Device 500 is configured to use clean-noisy (xy) pairs as the basis for neural network 580 after training block 5101. 1...N (For example, the neural network 580) 1...N This can correspond to neural network 380, NN(i), or neural network 580. 1...N It can even be equal to the corresponding neural network in neural network 380, NN(i)) to provide neural network parameters to map to the distribution (e.g. Gaussian distribution) of the training result audio signal 520 (e.g., noise signal).
[0156] The training audio signals x and 505 are introduced together with the distorted training audio signals y and 530 into the first process block 5101 of the device 500. For example, the training audio signals x and 505 are represented as audio samples, such as time-domain samples.
[0157] The training audio signal x can (optionally) be preprocessed into a training audio signal x (i=1) before entering the device 500. For example, the audio samples x of the noise signal x can be grouped into sample groups, such as into groups of 8 samples, or into vector representations (or matrix representations). Figure 1 Optional preprocessing steps are not shown.
[0158] The training audio signal x (i=1), 5401, and the distorted training audio signal y, 530 are introduced together into the first process block 5101 (e.g., the training process block) of the device 500. Reference Figure 6 and Figure 7 Further description in the first process block 5101 and process block 510 1...NThe subsequent process blocks handle the processing of training audio signals x (i=1), 5401, and distorted training audio signals y and 530. The training audio signals x (i=1), 5401 are processed in process block 510. 1...N The signal is processed (or further processed) based on, for example, a distorted training audio signal y, 530, which is adjusted (e.g., continuously or progressively). For example, the distorted training audio signal y, 530 is introduced into process block 510. 1...N In each process block.
[0159] After processing the training audio signal x (i=1) in the first process block 5101 and 5401, the output signal x is... new (i=1), 5501 is output. Signal x new (i=1) and 5501 are the input signals x (i=2) and 5402, and the distortion training audio signals y and 530, respectively, used for the second process block 5102 of device 500. new (i=2), 5502 are the input signals x (i=3) for the third process block, etc. The last N process blocks are 510. N Having a signal x (i = N) as the input signal, 540 N and output signal x new (i=N)550 N The signal x new (i=N)550 N The output signal 520 of the forming apparatus 500 or the training result audio signal z, 520, which is, for example, a noise signal (or at least a noise-like signal with similar statistical properties to a noise signal). The training result audio signal z, 520 may optionally be grouped into a vector representation (or a matrix representation).
[0160] The processing of the training audio signal x is in flowchart block 510. 1...N The process is executed based on the distorted training audio signal y, 530, for example, iteratively.
[0161] For example, an estimation (or evaluation) of the training result audio signal z,520 can be performed after each iteration to determine or estimate whether the characteristics (e.g., distribution (e.g., distribution of signal values) of the training result audio signal z,520 are close to predetermined characteristics (e.g., Gaussian distribution). If the characteristics of the training result audio signal z,520 are not close to the predetermined characteristics (e.g., within the expected tolerance), the neural network parameters can be changed before subsequent iterations.
[0162] Therefore, the neural network 580 can be determined (e.g., iteratively). 1..N The neural network parameters make the neural network 5801…N Under the control of a series of process blocks 510 1…N The training result audio signal obtained by processing the training audio signal includes (or approximates) the expected statistical characteristics (e.g., the expected distribution of values) within (e.g., a predetermined) allowable tolerance.
[0163] For example, neural network 580 1..N The parameters of the neural network can be evaluated using a cost function (e.g., an optimization function); for example, a parameter optimization procedure can be used to determine such that the characteristics (e.g., probability distribution) of the audio signal from the training result are close to or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution).
[0164] In device 500, a clean signal x is introduced together with a corresponding distorted (e.g., noisy) audio signal y for training and training process block 510. 1...N 580 related neural networks 1...N Considering the training result audio signal 520, as the training result, device 500 determines (590) neural network 580. 1...N The neural network parameters, such as edge weights (θ).
[0165] The neural network parameters determined by device 500 can, for example, be derived from... Figure 1 , 2 The neural network associated with the flow block of the device shown in 4 uses (where it should be noted that, Figure 1 , 2 The process block of device 510 can be configured, for example, to perform an affine transformation, which is substantially the same as that performed by process block 510. 1…N The affine transformation performed is the opposite.
[0166] However, it should be noted that the device 500 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0167] Figure 6 A schematic representation of a device 600 for providing neural network parameters according to an embodiment is shown.
[0168] In embodiments, the features, functions, and details of device 600 may optionally be incorporated into Figure 5 In the apparatus 500 shown (individually and in combination), and vice versa.
[0169] Device 600 is configured to use a distorted version y of the training audio signal x, 605 (e.g., a clean audio signal) and the training audio signal y. input 630 (e.g., distorted audio signal) is used to provide neural network parameters. In conjunction with a neural network (not shown) (e.g., as... Figure 5580 neural networks 1...N N process blocks (e.g., training process block 610) associated with such a neural network 1...N Processing is performed within (process block 610). 1...N It is configured to process incoming audio signals, such as speech signals.
[0170] A distorted version of the training audio signal, y, 630, is introduced into device 600 for processing. The distorted audio signal y, 630 is, for example, a noisy input signal. The distorted training audio signal y, 630 is, for example, defined as y = x + n, where x is the clean portion of the input signal, such as the training input signal x, 605, and n is the noisy background. The distorted training audio signal y, 630 is, for example, represented as a time-domain audio sample, such as a noisy time-domain speech sample.
[0171] The distorted versions of the training audio signal x and the corresponding training audio signal y can be optionally grouped into vector representations (or matrix representations).
[0172] Device 600 is configured to be based on training process block 610 1...N The subsequent clean-noise (xy) pair provides neural network parameters for the neural network (not shown) to map to the distribution (e.g., Gaussian distribution) of the training result audio signal 620 (e.g., noise signal).
[0173] The training audio signals x and 605 are introduced together with the distorted training audio signals y and 630 into the first process block 6101 of the device 600. The training audio signals x and 605 can be represented, for example, as audio samples, such as time-domain samples.
[0174] The training audio signal x may optionally be preprocessed into the input audio signal x before entering or within the device 600. input (i=1), 606. like Figure 6 As shown, the audio samples x with, for example, 16,000 samples of the training audio signal x are grouped into sample groups, for example, 2,000 groups of 8 samples, or grouped into vector representations (or matrix representations).
[0175] Input audio signal x input (i=1), 640 and the distorted training audio signal y, y input Together with 630, the first process block 6101 of device 600 is introduced, for example, a training process block. Input audio signal x input (i=1), 640 in process block 610 1...N Based on, for example, distorted training audio signals y, y input The 630 is adjusted and processed (e.g., continuously or progressively) (or further processed). Distorted training audio signal y, yinput 630 is introduced into process block 610 1...N In each process block.
[0176] The processing in the first process block 6101 is performed in two steps, for example in two blocks, for example in two operations: 1×1 reversible convolution 6121 and affine coupling layer 6111.
[0177] Before the introduction of the affine coupling layer block 6111, the input audio signal x is mixed (e.g., reordered, or subjected to invertible matrix operations such as rotation) in the invertible convolution block 6121. input (i=1), 640 samples. For example, the reversible convolutional block 6121 reverses the channel ordering at the input of the affine coupling layer block 6111. For example, the reversible convolutional block 6121 can be performed using a weight matrix W (e.g., as a random rotation matrix, a pseudo-random but deterministic rotation matrix, or a permutation matrix). The input audio signal x is processed in the reversible convolutional block 6121. input (i=1), 640, to output preprocessed (e.g., convolved) input audio signal x' input (i=1), 641. For example, distorted training audio signals y, y input 630 is not introduced into the reversible convolution block 6121, but is only used as input to the affine coupling layer block 6111. In an embodiment, the reversible convolution block may optionally be absent.
[0178] In the affine coupling layer block 6111, the preprocessed input audio signal x' input (i=1), 641 is based on, for example, distorted training audio signals y, y input The distorted training audio signal y was processed by adjusting 630. input 630 is introduced into the affine coupling layer block 6111 of the first process block 6101. (Refer to...) Figure 7 Further description of the affine coupling layer block 6111 and process block 610 in the first process block 6101 1...N The preprocessed input audio signal x' in the affine coupling layer block of the subsequent process block input (i=1), 641 and distorted training audio signals y, y input The processing of 630.
[0179] Process block 610 1...NThe processing in all subsequent process blocks is performed in two steps, for example, in two blocks, for example, in two operations: a 1×1 reversible convolution and an affine coupling layer. These two steps are, for example (e.g., qualitatively) the same as the description associated with the first process block 6101 (where the neural network of different processing stages or process blocks may include different parameters, and where the reversible convolution may be different in different process blocks or stages).
[0180] Process block 610 1...N The affine coupling layer is associated with the corresponding neural network (not shown).
[0181] After processing in the affine coupling layer block 6111 of the first process block 6101, the output signal x is... new (i=1), 6501 is output. Signal x new (i=1), 6501 is the input signal x of the second process block 6102 of device 600. input (i=2), 6402 and distorted training audio signals y, y input 630. Output signal x of the second process block 6102 new (i=2), 6502 are the input signals x (i=3) for the third process block, etc. The last N process blocks are 610. N A signal x is used as an input signal input (i=N), 640 N And output signal x of the output signal 620 of the forming device 600. new (i=N), 650 N Signal x new (i=N), 650 N The training result audio signal z,620 is formed, for example, a noise signal. The training result audio signal z,620 can optionally be grouped into a vector representation (or matrix representation).
[0182] The processing of the training audio signal x is in flowchart block 610. 1...N The process is performed based on the distorted training audio signal y, 630, for example, iteratively. For instance, after each iteration, an estimation (or evaluation or assessment) of the training result audio signal z, 620 can be performed to estimate whether the characteristics (e.g., distribution, such as the distribution of signal values) of the training result audio signal z, 620 are close to predetermined characteristics, such as a Gaussian distribution (e.g., within a desired tolerance). If the characteristics of the training result signal z, 620 are not close to the predetermined characteristics, the neural network parameters can be changed before subsequent iterations.
[0183] Therefore, the neural network can be determined (e.g., iteratively), which may correspond to neural network 580. 1..NThe parameters of the neural network are such that in the neural network 580 1…N Under the control of a series of process blocks 610 1…N The training results obtained by processing the training audio signals are audio signals 620 and 650. N This includes (or approximates) expected statistical properties (e.g., expected distribution of values) within (e.g., predetermined) allowable tolerances.
[0184] For example, the neural network parameters can be evaluated using a cost function (e.g., an optimization function); or determined using a parameter optimization procedure such that the characteristics (e.g., probability distribution) of the training result audio signal are close to or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution).
[0185] In device 600, a clean signal x is introduced together with a corresponding distorted (e.g., noisy) audio signal y for training with training process block 610. 1...N An associated neural network (not shown). Considering (or evaluating) the training result audio signal 620, as a result of the training, the device 600 determines the neural network parameters, such as edge weights (θ).
[0186] The neural network parameters determined by device 600 can, for example, be derived from... Figure 1 , 2 The neural network associated with the process block of the device shown in 4 is used, for example, in inference processing after training.
[0187] However, it should be noted that the device 600 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0188] Figure 7 A schematic representation of a process block 711 (e.g., a training process block) according to an embodiment is shown.
[0189] This process block can be, for example, composed of... Figure 5 The device 500 shown or made of Figure 5 This is part of the process performed by the device 600 shown. For example, Figure 5 The process blocks of the apparatus 500 shown can have the same as Figure 7 The process block 711 shown has the same structure or function. For example, Figure 6 The affine coupling layer of the process block of the device 600 shown can have a connection with... Figure 7 It has the same structure or function as the process block 711 shown.
[0190] For example, process block 711 is Figure 3The opposite version of the corresponding process block 311 shown, or an affine process that can be performed (at least substantially) in the opposite direction to the affine process performed in process block 311, can be executed. For example, the addition of shift values t in training process block 711 can be the opposite of the subtraction of shift values in inference process block 311. Similarly, multiplying by a scaling value s in training process block 711 can be the opposite of dividing by the scaling value s in inference process block 311. However, for example, the neural network in training process block 711 can be the same as the neural network in the corresponding inference process block 311.
[0191] For the sake of simplicity, in Figure 7 The process block index i is partially omitted in the following description.
[0192] Input signal 740 is introduced into process block 711. Input signal 740 may represent the training audio signal x(i), or, for example, a processed version of the training audio signal output from the preceding process block, or, for example, a preprocessed, convolutional input audio signal x'. input (i=1).
[0193] The input signal 740 is divided into two parts x1(i) and x2(i) in a random or pseudo-random (but deterministic) manner, for example (770).
[0194] The first part x1(i) is introduced into the neural network 780 associated with process block 711. The neural network 780 can be, for example, connected to... Figure 5 Process block 510 of the apparatus 500 shown 1...N A neural network associated with any process block (or a given process block). The neural network 780 can be, for example, a neural network associated with... Figure 6 Process block 610 of the apparatus 600 shown 1...N The neural network associated with any affine coupling block (or a given affine coupling block).
[0195] The first part x1(i) is introduced into the neural network 780 along with the distorted training audio signal y, 730. The distorted training audio signal y, 730 is, for example, a noisy signal, or for example, a distorted audio signal. The distorted training audio signal y, 730 is defined, for example, as y = x + n, where x is the clean training audio signal, such as the input signal 740, such as the clean part of the distorted training audio signal y, 730, and n is the noisy background.
[0196] The distorted versions of the training audio signal x and the corresponding training audio signal y can be optionally grouped into vector representations (or matrix representations).
[0197] The neural network 780 processes a first portion x1(i) of the input signal 740 and a distorted training audio signal y, 730, for example, by adjusting the first portion x1(i) according to the distorted training audio signal y, 730. The neural network 780 determines processing parameters, such as scaling factors (e.g., S) and shift values (e.g., T), which are the output (771) of the neural network 780. The determined parameters S, T have, for example, a vector representation. The second portion x2(i) of the input signal 740 is processed (772) using the determined parameters S, T.
[0198] The second part of the processed part (i) Defined by the following equation:
[0199]
[0200] In this equation, s can be equal to S (e.g., if the neural network provides only a single scaling factor value), or s can be an element of a vector of scaling factor values S (e.g., if the neural network provides a vector of scaling factor values). Similarly, t can be equal to T (e.g., if the neural network provides only a single shift value), or t can be an element of a vector of shift values T (e.g., if the neural network provides a vector of scaling factor values, with terms associated with different sample values of x2(i)).
[0201] For example, the above is used for The equation can be applied element-wise to a single element or group of elements in the second part x2. However, if the neural network provides only a single value s and a single value t, then that single value s and that single value t can be applied to all elements of the second part x2 in the same way.
[0202] The unprocessed first part x1(i) of signal x and the processed part of signal x The signals are combined (773) to form signal x, which is processed at process block 711. new 750. The output signal x new Introduced into the next process block, such as in a subsequent process block, such as in the second process block, such as in process block (i + 1). If i = N, signal x new 750 is the output signal of the corresponding device, for example, z. The output signal z can optionally be grouped into a vector representation (or matrix representation).
[0203] When a preprocessed noise signal x'(i) is used as input signal 740, input signal 740 is, for example, premixed to avoid processing the same x(i) in process block 711. For example, preprocessing (e.g., using reversible convolution) may have the effect of affine processing of different samples (e.g., from different original sample locations) in different process blocks (i.e., avoiding affine processing of the same subset of samples in each process block), and the effect of using different samples (e.g., from different original sample locations) as input signals to neural networks associated with different process blocks or processing stages (i.e., avoiding inputting the same subset of samples into the neural network in each process block). However, it should be noted that process block 711 may optionally be supplemented individually and in combination by any features, functions, and details disclosed herein.
[0204] Figure 8 A schematic representation of an apparatus 800 for providing neural network parameters according to an embodiment is shown.
[0205] In an embodiment, device 800 may, for example, be with Figure 5 The device 500 shown is combined, or for example with, the device 500 shown. Figure 8 The apparatus 600 shown is combined. In addition, the features, functions and details of apparatus 800 may be optionally incorporated into apparatus 500 or apparatus 600 (either individually or in combination), or vice versa.
[0206] In an embodiment, for example, Figure 7 The process block 711 shown can be used in device 800.
[0207] The device 800 is configured to provide neural network parameters based on training audio signals x, 805 (e.g., clean audio signals) and distorted versions y, 830 of the training audio signals (e.g., distorted audio signals). For example, in N process blocks (e.g., training process block 810) associated with the neural network (only the neural network 8801 of the first process block 8101 is shown). 1...N Processing is performed within (process block 810). 1...N It is configured to process incoming audio signals, such as speech signals.
[0208] A distorted version of the training audio signal, y, 830, is introduced into device 800 for processing. The distorted audio signal y, 830 is, for example, a noisy input signal. The distorted training audio signal y, 830 is, for example, defined as y = x + n, where x is the clean portion of the input signal (e.g., training audio signal x, 805), and n is the noisy background. The distorted training audio signal y, 830 is, for example, represented as a time-domain audio sample, such as a noisy time-domain speech sample.
[0209] The training audio signal x and the corresponding distorted version y of the training audio signal can be optionally grouped into vector representations (or matrix representations).
[0210] The training audio signals x and 805 are introduced together with the distorted training audio signals y and 830 into the flow block 8101 of the device 800. The training audio signals x and 805 are represented, for example, as audio samples, such as time-domain samples.
[0211] Device 800 is configured to be based on training process block 810 1...N The subsequent clean-noisy (xy) pair provides neural network parameters for the neural network (not shown) to map to the distribution (e.g., Gaussian distribution) of the training result signal 820 (e.g., a noise signal).
[0212] The nonlinear input compression step 815 can optionally be applied to the training audio signals x and 805. For example... Figure 8 As shown, step 815 is optional. For example, the nonlinear input compression step 815 can be applied to compress the training audio signal x, 805, instead of in the training and flow block 810. 1...N The associated neural network learns the distribution of clear speech (e.g., clear audio signal x), and if an optional non-linear input compression step 815 is present, learns the distribution of compressed signal. In an embodiment, the non-linear input compression step 815 is as described in reference... Figure 9 As described.
[0213] In embodiments, for example, nonlinear input compression 815 can be represented by μ-law compression of the training audio signals x, 805, or, for example, by μ-law transformation. For example:
[0214]
[0215] Where sgn() is a symbolic function;
[0216] μ is a parameter that defines the compression level.
[0217] For example, the parameter μ can be set to 255, which is a common value used in telecommunications. For example, when it is desirable to ensure that all values of the noise signal z to be learned are uniformly distributed, nonlinear input compression step 815 can be applied.
[0218] Before introducing the training audio signals x, 805 into the first process block 8101 of the device 800, audio samples of the training audio signals x, 805 or compressed audio samples of the training input signal x' may optionally be grouped (816) into sample groups, for example, into groups of 8 samples, or into vector representations (or matrix representations). Figure 8 As shown, grouping step 816 is an optional step.
[0219] The training audio signals x (i=1), 8401 (optionally grouped) and the distorted training audio signals y, 830 are introduced into the flow block 8101 of the device 800.
[0220] The nonlinear input compression step 815 can also optionally be applied to the distorted training audio signals y and 830. For example... Figure 8 As shown, step 815 is optional. For example, the nonlinear input compression step 815 can be applied to compress distorted training audio signals y, 830. In an embodiment, the nonlinear input compression step 815 is as described in reference... Figure 9 As described.
[0221] In embodiments, for example, nonlinear input compression 815 can be represented by μ-law compression of the distorted training audio signals y, 830, or, for example, by μ-law transformation. For example:
[0222]
[0223] Where sgn() is a symbolic function;
[0224] μ is a parameter that defines the compression level.
[0225] For example, the parameter μ can be set to 255, which is a common value used in telecommunications.
[0226] The training audio signals x (i=1) and 8401 are introduced together with the distorted training audio signals y and 830, or a pre-processed (e.g., compressed) distorted training audio signal y', into the first process block 8101 (e.g., the training process block) of the device 800. The training audio signals x (i=1) and 8401 are introduced into the first process block 8101 (e.g., the training process block) of the device 800. 1...N The signal is processed based on, for example, a distorted training audio signal y, 830. The distorted training audio signal y, 830 is introduced into process block 810. 1...N In each process block.
[0227] The processing in the first process block 8101 is performed in two steps, for example, in two blocks, for example, in two operations: 1×1 reversible convolution 8121 and affine coupling layer 8111.
[0228] Before the introduction of the affine coupling layer block 8111, samples of the training audio signals x (i=1), 8401 are mixed (e.g., reordered, or subjected to invertible matrix operations, such as rotation matrices) in the reversible convolution block 8121. For example, the reversible convolution block 8121 inverts (or alters) the channel order at the input of the affine coupling layer block 8111. For example, the reversible convolution block 8121 can be performed using a weight matrix W (e.g., as a random rotation matrix, or a pseudo-random but deterministic rotation or permutation matrix). For example, the training audio signals x (i=1), 8401 are processed in the reversible convolution block 8121 to output preprocessed (e.g., convolved) training audio signals x' (i=1), 8411. For example, the distorted training audio signals y, 830 are not introduced into the reversible convolution block 8121 but are only used as input to the affine coupling layer block 8111. In embodiments, the reversible convolution block may optionally be absent.
[0229] In the affine coupling layer block 8111, the preprocessed training audio signal x' (i=1) 8411 is processed based on, for example, a distorted training audio signal y, 830, which is then introduced into the affine coupling layer block 8111 of the first process block 8101. For example, refer to Figure 7 To describe the affine coupling layer block 8111 and process block 810 in the first process block 8101 1...N The processing of preprocessed training audio signals x' (i=1) 8411 and distorted training audio signals y and 830 in the affine coupling layer block of the subsequent process block.
[0230] Process block 810 1...N The processing in all subsequent process blocks is performed in two steps, for example, in two blocks, for example, in two operations: a 1×1 reversible convolution and an affine coupling layer. These two steps are, for example (e.g., qualitatively) the same as the description associated with the first process block 8101 (where the neural networks of different processing stages or process blocks may include different parameters, and where the reversible convolution may be different in different process blocks or stages).
[0231] Process block 810 1...N The affine coupling layer blocks are associated with the corresponding networks (only the neural network 8801 of the first process block 8101 is shown).
[0232] After processing in the affine coupling layer block 8111 of the first process block 8101, the output signal x is... new (i=1), 8501 is output. Signal x new(i=1) and 8501 are the input signals x (i=2) and 8402, and the distortion training audio signals y and 830, used for the second process block 8102 of device 800. new (i=2), 8502 are the input signals x (i=3) for the third process block, etc. The last N process blocks are 810. N Having a signal x (i = N) as the input signal, 840 N And output signal x of the forming device 800 output signal 820. new (i=N), 850 N For example, the audio signal of the training result. Signal x new (i=N)850 N The training result audio signal z,820 is formed, for example, a noise signal. The training result audio signal z,820 can optionally be grouped into a vector representation (or matrix representation).
[0233] The processing of the training audio signal x is in process block 810. 1...N The process is performed iteratively, for example, based on the distorted training audio signal y, 830. For instance, an estimation (or evaluation or assessment) of the training result signal z, 820 can be performed after each iteration to estimate whether the characteristics (e.g., distribution, such as the distribution of signal values) of the training result signal z, 820 are close to predetermined characteristics, such as a Gaussian distribution (e.g., within a desired tolerance). If the characteristics of the training result audio signal z, 820 are not close to the predetermined characteristics, the neural network parameters can be changed before subsequent iterations.
[0234] Therefore, the neural network can be determined (e.g., iteratively), which may correspond to neural network 580. 1..N 780 1..N The neural network parameters make the neural network 880 1…N Under the control of a series of process blocks 810 1…N The training results audio signals 820 and 850 were obtained by processing the training audio signals. N This includes (or approximates) expected statistical properties (e.g., expected distribution of values) within (e.g., predetermined) allowable tolerances.
[0235] For example, the neural network parameters can be evaluated using a cost function (e.g., an optimization function); or determined using a parameter optimization procedure such that the characteristics (e.g., probability distribution) of the training result audio signal are close to or include predetermined characteristics (e.g., noise-like characteristics; e.g., a Gaussian distribution).
[0236] In device 800, a clean signal x is introduced together with a corresponding distorted (e.g., noisy) audio signal y for training and training process block 810.1...N The associated neural network. Consider (or evaluate) the training result signal 820 as a result of training, and determine the neural network parameters, such as edge weights (θ).
[0237] The neural network parameters determined by device 800 can, for example, be derived from... Figure 1 , 2 The neural network associated with the process block of the device shown in 4 is used, for example, in inference processing after training.
[0238] However, it should be noted that the device 800 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0239] In the following sections, some basic considerations according to embodiments of the present invention will be described. For example, problem formulas will be provided, the basic principles of the normalization process will be described, and the speech enhancement process will be discussed. The concepts described below can be used alone or in combination with the embodiments described herein.
[0240] exist Figure 1 , 2 The process block processing used in the devices 100, 200, and 400 shown in Figure 4, and in Figure 5 , 6 The process block processing used in the devices 500, 600 and 800 shown in Figure 8 can be described, for example, as transforming a simple probability distribution into a more complex probability distribution using invertible and differentiable maps, formally expressed as follows:
[0241] x = f(z), (7)
[0242] Where x∈R D and z∈R D It is a D-dimensional random variable, and
[0243] f is a mapping function from z to x.
[0244] (5) represents a differentiable and invertible transformation.
[0245] The reversibility of f guarantees that this step can be recovered from x to z:
[0246] z = f -1 (x). (8)
[0247] Furthermore, if the function f is invertible and differentiable, then the synthesis of a series of 1 to T transformations is also invertible and can be described by a neural network:
[0248] x=f1°f2°...°f T (z) (9)
[0249] Next, the log probability density function, such as the log-likelihood function, log p x (x) can be calculated by changing the variable, for example, by direct calculation:
[0250] log p x (x)=log p z (f -1 (x))+log|det(J(x))| (10)
[0251] in The Jacobian determinant consisting of all first-order partial derivatives is defined.
[0252] It is important to note, for example, the function f -1 (For example, from training block 510) 1...N The executed partial function components) can be executed in devices 500, 600, and 800, where function f (for example, is derived from inference process block 110) 1...N The components of the executed functions can be executed in devices 100, 200, and 400.
[0253] The function definition of f can be derived, for example, from f -1 The function definition is derived because the partial function executed by the training block is reversible. Therefore, by determining (during training) the rules (e.g., neural network parameters) for defining the function f executed by training devices 500, 600, and 800, the function f is also implicitly obtained. -1 Definition.
[0254] In other words, f -1 The function definition of f can be determined during training (e.g., by determining the parameters of the neural network such that the training audio signal x is transformed into a noise-like signal z), and the function definition of f can be derived from f -1 The function definition is exported.
[0255] In the following description, some further (optional) details regarding the speech enhancement process will be described, which may be optionally used in the apparatus and methods according to embodiments of the present invention (e.g., in apparatus 100, 200, 400, 500, 600, 800 or in process blocks 311, 711).
[0256] In the case of speech enhancement (or, generally, in the case of audio enhancement), a time-domain mixed signal y∈R of length N 1×N It can be composed of clean speech utterances x∈R 1×N And some additional background interference n∈R 1×N Composition, for example, noisy mixing is shown as the sum of clean speech and interfering background, therefore
[0257] y = x + n. (11)
[0258] Furthermore, z∈R 1×N It is defined as sampling from a normal distribution with zero mean and unit variance, for example, as a Gaussian sample, i.e.
[0259] z~N(z;0,I).(12)
[0260] exist Figure 5 , 6 The block-based models proposed in devices 500, 600, and 800 shown in Figures 8 and 9 are defined as DNNs and are designed to depict the probability distribution p formed by clean speech utterances x conditioned on noisy mixtures y. x (x|y), for example, learning the probability distribution function of x conditionally with respect to y. For example, minimizing the negative log-likelihood function of a previously defined probability distribution is considered the training objective (where, for example, the value of the following expression can be minimized by optimizing the neural network parameters):
[0261]
[0262] Where θ represents the neural network parameters.
[0263] In the enhancement step (e.g., in Figure 1 , 2 (among the devices 100, 200 and 400 shown in Figure 4), from p z A sample is taken from (z) and fed into the neural network along with a noisy sample, for example, sample z follows the reverse flow along with noisy input y. For example, time-domain sample values of a noisy signal with a predetermined distribution (e.g., Gaussian distribution) can be input into the (first) flow block (and thus, for example, in a pre-processed form, for example, along with a sample of audio signal y into the neural network). After the reverse flow (e.g., reverse flow block processing, for example, the opposite of the processing performed in the training flow block), the neural network (e.g., and affine processing 372) maps the random sample (or sample) back to the distribution of clean speech to create an enhanced audio signal, such as an enhanced speech signal. in Ideally, it should be close to the lower layer x, for example.
[0264] These correspondences 7 to 13 are for Figure 10 The system 1000 shown is also correct.
[0265] In the following, nonlinear input companding, such as compression and / or expansion, which may optionally be used in embodiments according to the invention will be described.
[0266] Figure 9 A diagram illustrating the nonlinear input compression steps used in the apparatus described herein is shown.
[0267] The nonlinear input compression step can be used, for example, corresponding to... Figure 5 , 6 The preprocessing block of any of the devices 500, 600 or 800 shown in Figure 8.
[0268] Nonlinear compression algorithms can be applied to map small amplitudes of audio data samples to a wider range and larger amplitudes to a smaller range.
[0269] Figure 9 An audio signal, such as a clean signal x, is shown at reference numeral 910, which is illustrated as varying over time, for example, in a time-domain representation. An example of a speech signal, such as x, is shown. Figure 5 , 6 The neural network associated with devices 500, 600, or 800 shown in 8 learns from this example. Since the neural network models the probability distribution of temporal audio (e.g., speech, utterances), it is important to examine the range of values from which the distribution is learned. For example, audio data is typically stored as a normalized 32-bit stream in the range [-1, 1]. Temporal audio (e.g., speech) samples approximately follow a Laplace distribution.
[0270] The histogram of the audio (e.g., speech) signal before (a) and after (b) the application of a compression algorithm (e.g., non-linear input compression, such as 815). Figure 1 This is illustrated below. For example, compression is understood as a kind of histogram equalization, or as histogram expansion for relatively low signal values and / or as histogram compression for relatively large signal values. For example, a comparison of the first histogram 920 before the compression algorithm is applied and the second histogram 920 after the compression algorithm is applied shows that the histogram becomes wider. For example, the x-axis 922 of histogram 920 shows the signal values before compression, and the y-axis 924 describes the probability distribution of each signal value. For example, the x-axis 932 of histogram 920 shows the signal values after compression, and the y-axis 934 describes the probability distribution of each signal value. Clearly, the compressed signal values include a wider (more uniform distribution, fewer peaks) probability distribution, which has been found to be advantageous for processing in the flow block.
[0271] like Figure 9As shown in (a) (e.g., at reference numeral 920 in the figure), most values of the approximate Laplace distribution lie in a small range near zero. It has been recognized that in clean speech samples (or clean speech signals), e.g., x, data samples (or signal values) with higher absolute amplitudes carry important information and often indicate deficiencies, such as... Figure 9 As shown in (a), the application of compression algorithms makes the values of time-domain speech samples more uniformly distributed.
[0272] In an embodiment, nonlinear input compression can be represented by, for example, μ-law compression of unquantized input data:
[0273]
[0274] Where sgn() is a symbolic function;
[0275] μ is a parameter that defines the compression level.
[0276] For example, the parameter μ can be set to 255, which is a common value used in telecommunications.
[0277] Regarding learning, such as training objectives, in Figure 5 , 6 In the process block processing of the apparatus 500, 600 or 800 shown in Figure 8, the distribution of compressed signals (e.g., preprocessed signals x) is learned, rather than the distribution of clean speech (e.g., unprocessed clean signals x).
[0278] Algorithms that are the opposite of the described nonlinear input compression algorithm include, for example, in Figure 1 , 2 In the apparatus 100, 200, or 400 shown in 4, it is used as a final processing step, for example Figure 4 The nonlinear extension 415 is shown. Figure 1 , 2 Enhanced samples of 4 and 5 It can be extended to regular signals, for example, by using inverse μ-law transforms, such as:
[0279]
[0280] Where sgn() is a symbolic function;
[0281] μ is a parameter that defines the compression level.
[0282] However, it should be noted that Figure 9 The nonlinear input compression shown may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0283] Figure 10A schematic representation of a block system 1000 for audio signal processing according to an embodiment is shown.
[0284] The process block system 1000 represents a combination (or can be used alone) of means 1100 for providing neural network parameters and means 1200 for providing processed signals. For example, means 1100 can be implemented as Figure 5 , 6 And any of the devices 500, 600, or 800 shown in Figure 8. For example, device 1200 can be implemented as Figure 1 , 2 Any of the devices 100, 200, or 400 shown in Figure 4.
[0285] In apparatus 1100, a clean-noisy pair (xy) is processed after a flow block (or input into the flow block processing) to map to a Gaussian distribution N(z; 0; 1) (or a distribution approximating a Gaussian distribution). In inference (apparatus 1200), a sample z (e.g., a block of sample values) is extracted from this distribution (or from a signal having a desired, e.g., a Gaussian signal value distribution) and processed together with another noisy utterance y after a reverse flow block to generate an enhanced signal.
[0286] The device 1100 is configured to provide neural network parameters based on a training audio signal 1105 (e.g., a clean x, such as x1) and a distorted version of the training audio signal 1130 (e.g., a noisy y, such as y1 = x1 + n1). N process blocks (e.g., training process block 1010) are associated with the neural network (not shown). 1...N Processing is performed within (process block 1010). 1...N It is configured to process incoming audio signals, such as speech signals.
[0287] Before introducing the training audio signal 1105 into the first process block 11101 of the device 1100, the audio samples of the training audio signal x are grouped (1116) into sample groups, for example, into groups of 8 samples, or into vectors.
[0288] The training audio signal x (i=1) may be optionally grouped, for example, x1 (i=1) may be introduced into the first process block 11101 of the device 1100 together with the distorted training audio signal y, 1130 (e.g. y1).
[0289] The distorted audio signal y, 1130 is, for example, a noisy input signal. The distorted training audio signal y, 1130 is, for example, defined as y = x + n, where x is the clean portion of the input signal, such as the training input signal x, 1105, and n is the noisy background, for example, y1 = x1 + n1, where x1 is the clean portion of the input signal, such as the training input signal x1, 1105, and n1 is the noisy background. The distorted training audio signal y, 1130 can be represented, for example, as a time-domain audio sample, such as a noisy time-domain speech sample.
[0290] The training audio signal x and the corresponding distorted version y of the training audio signal can be optionally grouped into vector representations (or matrix representations).
[0291] Device 1100 is configured to feed into a neural network (e.g., neural network 580) based on clean-noisy (xy) pairs (or multiple clean-noisy pairs). 1...N This provides the neural network parameters, which are specified in training block 1110. 1...N Then (or by training block 1110) 1...N The training audio signal 1120 is processed to map to a distribution (e.g., a Gaussian distribution) of the training result audio signal 1120 (e.g., a noise signal, e.g., z). The training result audio signal 1120 may optionally be grouped into a vector representation (or a matrix representation).
[0292] The training audio signals x, 1105 and the distorted training audio signals y, 1130 are introduced together into the first process block 11101 of the device 1100. The training audio signals x, 1105 are represented, for example, as audio samples, such as time-domain samples.
[0293] Process block 1110 1...N It can be implemented, for example, in Figure 5 , 6 And the process block 510 of the device 500, 600 or 800 shown in Figure 8 1...N 610 1...N Or 810 1...N .
[0294] Process block 1110 1...N It may include, for example, affine coupling blocks, such as, for example, Figure 6 , 7 And the process blocks 6111, 7111, or 8111 shown in Figure 8.
[0295] As process block 1110 1...N The output provides the training result audio signal z, 1120, which may be a noise signal (or an approximate noise signal). The noise signal z, 1120 is defined, for example, as z ~ N(z; 0; 1).
[0296] In device 1100, clean signals x and 1105 are introduced together with corresponding distorted (e.g., noisy) audio signals y and 1130 for training and training process block 1110. 1...N The associated neural network, and as a result of training, determines the neural network parameters, such as edge weights (θ).
[0297] The neural network parameters determined by device 1100 can, for example, be further used in the inference provided by device 1200.
[0298] The device 1200 is configured to provide a processed (e.g., enhanced) audio signal based on the input audio signals y, 1230. This is achieved through N process blocks (e.g., inference process block 1210) associated with a neural network (not shown). 1...N Processing is performed within (process block 1210). 1...N It is configured to process incoming audio signals, such as speech signals.
[0299] The input audio signals y, 1230 (e.g., a new noisy signal y2) are introduced into device 1200 for processing. The input audio signal y is, for example, a noisy input signal, or, for example, a distorted audio signal. The input audio signals y, 1230 are defined as y = x + n, where x is the clean portion of the input audio signal and n is the noisy background, for example, y2 = x2 + n2. The input audio signals y, 1230 can be represented, for example, as time-domain audio samples, such as noisy time-domain speech samples.
[0300] The input audio signal y and its corresponding clean portion x can be optionally grouped into vector representations (or matrix representations).
[0301] A noise signal z, 1220 is obtained (e.g., generated) and introduced into the first process block 12101 of the apparatus 100 along with the input audio signal y, 1230. The noise signal z, 1220 is defined, for example, as being sampled from a normal distribution with zero mean and unit variance (e.g., z ~ N(z; 0; 1)). The noise signal z, 1220 is represented, for example, as a noise sample, such as a time-domain noise sample.
[0302] Before introducing the noise signal 1220 into the first process block 12101 of the device 1200, the audio samples of the noise signal z may optionally be grouped (1216) into sample groups, for example, into groups of 8 samples, or into vectors (or into matrices). This grouping step may be optional, for example.
[0303] The noise signal z (i=1) (e.g., x1 (i=1)) to be optionally grouped is introduced into the first process block 12101 of device 1200 along with the input audio signals y, y1230 (e.g., y2). Process block 1210 1...N Flow block 1110 of device 1100 1...N The reversal (e.g., when compared with the corresponding process block of device 1100, process block 1210) 1...N Perform inverse affine processing, and optionally also perform deconvolution processing.
[0304] Process block 1210 1...N It can be implemented, for example, in Figure 1 , 2 Flow block 110 of the apparatus 100, 200 or 400 shown in Figure 4 1...N 210 1...N Or 410 1...N .
[0305] Process block 1210 1...N It may include, for example, affine coupling blocks, such as, for example, Figure 1 , 2 And the process blocks 2111, 3111, or 4111 shown in 4.
[0306] As process block 1210 1...N The output provides a processed (e.g., enhanced) audio signal. 1260. Enhanced audio signal 1260 represents, for example, the enhanced clean portion of the input audio signal y, 1230.
[0307] The clean portion x of the input audio signals y, 1230 is not separately introduced into device 1200. Device 1200 processes, for example, generated noise signals z, 1220 based on the input audio signals y, 1230 to receive, for example, generate, for example, an enhanced output audio signal, which is an enhancement of the clean portion of the input audio signals y, 1230.
[0308] However, it should be noted that System 1000 may optionally be supplemented individually and in combination by any features, functions and details disclosed herein.
[0309] Figure 11 Table 1 is shown, illustrating a comparison between the apparatus and method according to the embodiments and conventional techniques.
[0310] Table 1 shows the evaluation results using objective evaluation metrics. (See above reference.) Figures 1 to 3As described in 5 to 7 and 10, SE-process represents the process-based method proposed according to the embodiments, and SE-process-μ represents a method including μ-law transformation, for example as nonlinear companding, such as reference Figure 9 The compression described, such as nonlinear input compression or nonlinear expansion, serves as a corresponding preprocessing or postprocessing step, as described in the corresponding reference above. Figure 8 and Figure 4 The embodiments described herein.
[0311] As shown in the table, the model using μ-compressorization exhibits better results across all metrics compared to the two proposed process-based experiments. This demonstrates the efficiency of such simple pre- and post-processing techniques (e.g., nonlinear compressorization) for modeling time-domain signal distributions.
[0312] Figure 12 also shows an illustration of enhanced capabilities.
[0313] Figure 12 shows a graphical representation of the performance of the apparatus and method according to the embodiments.
[0314] Figure 12 shows example spectrograms to illustrate the performance of the proposed embodiment. In (a), noisy speech with a signal-to-noise ratio (SNR) of 2.5 dB is shown. (b) shows the corresponding clean speech. In (c) and (d), the results of the proposed process-based system according to an embodiment of the invention are shown, for example... Figure 10 As shown.
[0315] Other embodiments and aspects
[0316] Other aspects and embodiments of the invention will be described below, which may be used alone or in combination with any other embodiments disclosed herein.
[0317] Furthermore, the embodiments disclosed in this section may optionally be supplemented, individually and in combination, by any other features, functions, and details disclosed herein.
[0318] The following text will describe the concept of a process-based neural network for temporal speech enhancement.
[0319] The basic idea of embodiments of the present invention will be described below.
[0320] In the following description, some of the objectives and objects of the invention will be described, which may be achieved (at least in part) in some or all of the embodiments, and some aspects of the invention will be briefly outlined.
[0321] Speech enhancement involves distinguishing the target speech signal from intrusive background. While conventional generative methods using variational autoencoders or generative adversarial networks (GANs) have become increasingly prevalent in recent years, systems based on normalized process (NF) remain rare, despite their success in related fields. Therefore, below, according to embodiments, an NF framework is proposed to directly model the enhancement process through density estimation of clean speech utterances conditioned on their noisy counterparts. In the embodiments, conventional models inspired by speech synthesis are adapted to directly enhance noisy utterances in the temporal domain. Experimental evaluations on publicly available datasets, according to embodiments, show performance comparable to state-of-the-art GAN-based methods, while exceeding selected baselines using objective evaluation metrics.
[0322] Embodiments of the present invention can be used for speech enhancement. Embodiments of the present invention utilize normalization processes and / or deep learning and / or generative modeling.
[0323] A brief introduction will be provided below.
[0324] Traditionally, speech enhancement (SE) aims to emphasize the target speech signal from the background interference to ensure better understanding of spoken content [1]. It has been extensively studied in the past due to its importance for a wide range of applications, including hearing aids [2] or automatic speech recognition [3]. In this context, deep neural networks (DNNs) have largely replaced conventional techniques such as Wiener filtering [4], spectral subtraction [5], subspace methods [6], or minimum mean square error (MMSE) [7]. Typically, DNNs are used to estimate time-frequency (TF) masks that can separate speech and background from mixed signals [8]. However, systems based on time-domain inputs have been proposed in recent years, which have the advantage of avoiding the expensive TF transformation [9, 10, 11]. Recently, SE research has also focused on generative methods such as generative adversarial networks (GANs) [11, 12, 13], variational autoencoders (VAEs)
[14] , and autoregressive models
[10] . In particular, the use of GANs, in which the generator and discriminator are trained simultaneously in an adversarial manner, has been extensively studied in the past few years. For example, Pascual et al.
[11] proposed a GAN-based end-to-end system in which the generator directly enhances noisy speech samples at the waveform level. This approach is extended several times below, for example by leveraging the Wasserstein distance
[15] or by combining multiple generators to improve performance
[16] . Others have reported impressive SE results by cooperating with GANs to estimate clean TF spectrograms by implementing additional techniques such as mean squared error regularization
[12] or by directly optimizing the network for speech-specific evaluation metrics
[13] . While the above conventional approaches have gained popularity recently, systems based on normalized process (NF) remain rare in SE. Just recently, the work of Nugraha et al.
[17] proposed a process-based model combined with a VAE to learn deep latent representations that can be used as a prior for deep speech. However, their approach does not model the enhancement process itself and therefore relies on the SE algorithm combined with it. However, in fields such as computer vision
[18] or speech synthesis
[19] , it has been shown that NF is capable of successfully generating high-quality samples in their respective tasks. Therefore, this leads to the assumption that, in the basic embodiment of the invention, speech sample enhancement can be performed directly using a process-based system by modeling the generation process.
[0325] The idea behind embodiments of the invention is that NF can be successfully applied to SE by learning mappings from simple probability distributions to more complex probability distributions based on clean speech samples conditioned on their noisy counterparts. Therefore, in embodiments of the invention, conventional process-based DNN architectures are modified from speech synthesis to perform SE directly in the temporal domain without requiring any predefined features or TF transforms. Furthermore, in embodiments of the invention, simple preprocessing techniques, such as compression, e.g., nonlinear compression, are applied to the input signal as part of the companding process to improve the performance of the density estimation-based SE model. Experimental evaluations of the proposed methods and apparatus according to embodiments of the invention confirm these assumptions and show or improve performance compared to current state-of-the-art systems, while exceeding the results of other temporal GAN baselines.
[0326] Figure 10 An overview of the proposed system according to an embodiment is shown. A clean-to-noisy (xy) pair is mapped to a Gaussian distribution N(z; 0, 1) after the process step (blue solid line). In inference, a sample z is drawn from this distribution and, after reflow (red dashed line) and another noisy utterance y, to generate an enhanced signal. (It's best to view it by color, or consider the different shadows of the block outlines).
[0327] Problem Formulation and Implementation Examples
[0328] The following sections will provide formulations and explanations regarding aspects of the invention according to embodiments.
[0329] Normalization process basics
[0330] The normalization process can be described as using invertible and differentiable mappings
[20] to transform a simple probability distribution into a more complex one, formally represented by the following equation.
[0331] x = f(z), (16)
[0332] Where x∈R D and z∈R D Let z be a D-dimensional random variable, and f be a mapping function from z to x. The invertibility of f guarantees that this step can be recovered from z to x, that is,
[0333] z = f -1 (x). (17)
[0334] Furthermore, if the function f is invertible and differentiable, then the synthesis of a series of 1 to T transformations is also ensured to be invertible.
[0335] x=f1°f2°...°f T (z) (18)
[0336] Next, the log probability density function logp can be calculated by changing the variable
[21] . x (x):
[0337] logp x (x)=logp z (f -1 (x))+log|det(J(x))| (19)
[0338] in The Jacobian determinant consisting of all first-order partial derivatives is defined.
[0339] Speech enhancement process
[0340] In the case of speech enhancement according to an embodiment of the present invention, a time-domain mixed signal y∈R of length N is... 1×N Pronounced by clean speech x∈R 1×N And some additional background interference n∈R 1×N Composition, making
[0341] y = x + n. (20)
[0342] Furthermore, it is defined as z∈R sampled from a normal distribution with zero mean and unit variance. 1×N ,Right now,
[0343] z ~ N(z; 0, I). (21)
[0344] According to an embodiment of the invention, the proposed NF model is now defined as a DNN and aims to outline the probability distribution p formed by a clean speech utterance x conditioned on a noisy mixture y. x (x|y). As a training objective according to an embodiment of the invention, the negative log-likelihood function of a previously defined probability distribution of speech samples can now be simply minimized:
[0345]
[0346] Where θ represents the network parameters.
[0347] logp x (x|y;θ) is, for example, the probability distribution of the speech sample to be defined. For example, log|det(J(x))| is the likelihood function of the Gaussian function, which shows the level of change (e.g., how much we changed) of the Gaussian function used to create (e.g., generate) speech samples.
[0348] In the enhancement step, according to an embodiment of the invention, it can be from p zThe sample (z) can be delivered as input to the network along with the noisy sample. After the reverse process, the neural network maps the random sample back to the distribution of the clean utterance to create in Ideally, it should be close to the lower layer x, for example. This process also Figure 10 The diagram is shown in the image.
[0349] In practice, for example in modeling or neural networks, where the total number of neural network parameters is, for example, approximately 25 million, according to an embodiment of the present invention, during the training of the neural network, signals x and y are introduced into the neural network to be trained. The output of the neural network includes z (the processed signal after all process blocks), log|s| (from each affine coupling layer), and log|detW| (from each 1×1 invertible convolution).
[0350] Loss function to be optimized:
[0351]
[0352] It is the likelihood function of the Gaussian function (see above); (∑logs+∑log|detW|)–(det(J(x)) (see above).
[0353] The method proposed according to the embodiments of the present invention
[0354] Model Architecture
[0355] According to an embodiment of the present invention, the Bohui architecture
[19] was modified for speech synthesis to perform speech enhancement. Initially, the model combines the spoken utterance with the corresponding Mel-spectrum. Figure 1The input is used as input to several procedural steps to learn to generate realistic speech samples based on conditional spectrograms. A procedural block consists of a 1×1 invertible convolution
[22] and a so-called affine coupling layer
[23] , which ensures the exchange of information along the channel dimension and is used to ensure the invertibility and efficient computation of the Jacobian determinant. Thus, the input signal is split along the channel dimension, with one half being fed into a NN block similar to a wavenet, such as a wavenet line affine coupling layer, which defines the scaling and translation factors for the latter half. To create this multi-channel input, multiple audio samples are stacked in a set to simulate a multi-channel signal. The affine coupling layer is also where conditional information is contained. For more details on the procedure, see
[19] . The original wavenet is computationally very large (>87 million parameters), so some architectural modifications were made to make it feasible to train on a single GPU and enable augmented speech. According to an embodiment of the invention, instead of wavenet, a non-Mel-spectrum is used as the conditional input, but noisy temporal speech samples are used. Therefore, since the two signals have the same dimension, no upsampling layer is required. Furthermore, according to an embodiment of the invention, standard convolutions in the wave mesh block are replaced by depth-separable convolutions
[24] to reduce the number of parameters, as recommended in
[25] .
[0356] Nonlinear input companding
[0357] Figure 9 An example of the effect of nonlinear input companding (e.g., compression) according to an embodiment of the present invention is shown. Clean speech is shown at the top. (a) A histogram (n) of the clean speech is shown. bins =100). In (b), the effect of the companding (e.g., compression) algorithm on the value can be seen.
[0358] Since the network models the probability distribution of temporal speech utterances, it is important to examine the range of values from which this distribution is learned. Audio data is stored as normalized 32-bit floating-point numbers in the range [-1, 1]. Since the temporal speech samples approximately follow a Laplace distribution
[27] , it is easy to see that most values lie in a small range near zero (see
[27] ). Figure 9 (a)). However, especially in clean speech, data samples with higher absolute amplitudes carry important information and are not adequately represented in this context. To ensure that these values (e.g., learnable amplitude values) are spread more uniformly, a nonlinear companding (e.g., compression) algorithm according to an embodiment of the invention can be applied to map small amplitudes to wider amplitudes and larger amplitudes to smaller intervals. This is as follows: Figure 9The diagram shows a histogram of a speech sample and the values before and after applying a companding (e.g., compression) algorithm. In this sense, companding, such as compression, can be understood as a form of histogram equalization. Further experiments were then conducted using μ-law companding, for example, with unquantized input data compressed (ITU-T Recommendation G.711), formally defined as follows:
[0359]
[0360] Where sgn() is the sign function, and μ is a parameter defining the compression level. Here, according to an embodiment of the invention, μ is set to 255 throughout the experiment; 255 is also a common value used in telecommunications. Regarding the learning objective, it is not to learn the distribution of clean speech, but rather the distribution of the compressed signal. The enhanced samples can be expanded back to the normal signal after recovering the μ-law transform.
[0361] experiment
[0362] data
[0363] The dataset used in the experiment was published alongside the work of Valentini et al.
[28] and is a commonly used database for developing SE algorithms. It consists of 30 individual speakers from the Speech Library corpus
[29] , divided into a training set with 28 speakers and a test set with 2 speakers. Both sets are balanced according to male and female participants. The training samples are mixed with eight real noise samples from the Demands Database
[30] and two artificial (multipath overlap and speech shaping) samples according to signal-to-noise ratios (SNR) of 0, 5, 10, and 15 dB. In the test set, different noise samples are selected and mixed according to SNR values of 2.5, 7.5, 12.5, and 17.5 dB to ensure that the test set only includes unseen conditions. In addition, one male speaker and one female speaker are taken from the training set to form a validation set for model development.
[0364] Training strategies
[0365] In a small hyperparameter search, values were chosen for batch size ∈ [4, 8, 12], number of process blocks ∈ [8, 12, 16], and the number of samples grouped together as input ∈ [8, 12, 24]. Each individual model was trained for 150 epochs to select parameters based on the lowest validation loss. Some initial experiments were conducted with a higher batch size, but it was found that the model did not generalize well. The model selected based on a patient early stopping mechanism of 20 epochs was further trained until convergence. Fine-tuning steps were then performed using a reduced learning rate and the same early stopping criterion.
[0366] Model setting
[0367] As a result of the parameter search, a model was constructed using 16 process blocks with a batch size of 4 samples as input. In the initial training step, the learning rate was set to 3 × 10⁻⁶ using the Adam
[31] optimizer and weight normalization
[32] . -4 To fine-tune the settings, the learning rate was reduced to 3×10. -5 As training input, a 1-second block (sampling rate f) is randomly extracted from each audio file. s =16kHz). The standard deviation of the Gaussian distribution was set to σ = 1.0. Similar to other NF models
[33] , using a lower σ value in inference produces higher quality output, which is why it was set to σ = 0.9 in inference. Based on a block of similar wavenet with an original wavenet architecture of 8 layers of extended convolutions, 512 channels were used as residual connections in the affine coupling layers and 256 channels were used in the skip connections. In addition, after every 4 coupling layers, 2 channels were passed to the loss function to form a multi-scale architecture.
[0368] Evaluate
[0369] In order to compare the method according to the embodiments of the present invention with recent work in the field, the following evaluation metrics were used:
[0370] • (i) Perceived Voice Quality Assessment (PESQ) in the broadband version recommended in ITU-T P.862.2 (from -0.5 to 4.5).
[0371] • Three mean opinion score (from 1 to 5) measures
[34] : (ii) signal distortion prediction (CSIG), (iii) background intrusiveness prediction (CBAK) and (iv) overall speech quality prediction (COVL).
[0372] • SegSNR (segSNR)
[35] Improvement (from 0 to ∞).
[0373] As a baseline for the method proposed according to embodiments of the present invention, two generative temporal approaches, namely SEGAN
[11] and the improved deep SEGAN (DSEGAN)
[16] model, are defined because they are evaluated using the same database and metric. Furthermore, it is compared with two other state-of-the-art GAN-based systems, namely MMSE-GAN
[26] and metric-GAN
[13] , which are studying TF masks. It should be noted that several distinguishing approaches, such as [9, 36, 37], report higher performance on this dataset. However, this work focuses on generative models, which is why they are not included in the comparison.
[0374] Experimental results
[0375] The experimental results are presented in Figure 11 Table 1 shows the evaluation results using objective evaluation metrics. SE-process represents the proposed process-based method, SE-process-μ represents the method, and μ-law companding, such as including compression and expansion of the input data, is also included. Values for all comparison methods are taken from their respective papers.
[0376] As shown in the table, the model using μ-compressorization (e.g., including both compression and expansion) demonstrates better results across all metrics between the two proposed process-based experiments. This proves the efficiency of this simple preprocessing and corresponding post-processing technique for modeling time-domain signal distributions. The enhancement capability can also be illustrated in Figure 12.
[0377] Figure 12 shows example spectrograms to illustrate the performance of the proposed system. In (a), noisy speech is shown at 2.5 dB (signal-to-noise ratio, SNR). (b) shows the corresponding clean speech. In (c) and (d), the results of the proposed process-based system according to embodiments of the present invention are shown.
[0378] By comparing the spectrograms of the two proposed systems according to embodiments of the present invention, it can be seen that the SE-procedure can capture finer-grained speech components with less background leakage. It should be noted that the model according to embodiments of the present invention does not recover the breathing sounds at the end of the illustrated example, which emphasizes that the model proposed according to embodiments of the present invention focuses on realistic speech samples. Furthermore, it can be seen in the procedural-based example that when speech is active, there is more noise-like frequency content at higher frequencies compared to a clean signal. This can be explained by Gaussian sampling, which is not completely eliminated during inference.
[0379] Compared to the SEGAN baseline, the method and apparatus proposed according to embodiments of the present invention exhibit superior performance with a large margin across all metrics. It should be noted that only the segSNR performance of SEGAN is visible, as other methods were not evaluated. Looking at DSEGAN, it can be seen that the proposed SE-procedure achieves comparable performance in CSIG, while showing slightly lower values in other metrics. However, the SE-procedure-μ-based system, method, or apparatus according to embodiments of the present invention still performs better in all metrics except COVL. Therefore, in the time domain, the procedural-based model proposed according to embodiments of the present invention appears to better simulate the generation process from noisy signals to enhanced signals. Regarding MMSE-GAN, similar performance is observed, but slightly better than MMSE-GAN, although no additional regularization techniques are implemented here. However, compared to the proposed method, metric-GAN shows better results across all exhibited metrics. It is important to note, however, that this model according to embodiments of the present invention is directly optimized based on the PESQ metric, and therefore the good performance according to embodiments of the present invention is to be expected. Therefore, directly linking the optimization of training and evaluation metrics may also be a way to improve the efficiency of the system, method, or apparatus according to embodiments of the present invention.
[0380] in conclusion
[0381] This disclosure describes a speech enhancement method based on a normalization process according to embodiments of the present invention. The model according to embodiments of the present invention allows density estimation of clean speech samples given their noisy counterparts and allows signal enhancement through generative inference. Simple nonlinear companding (e.g., compression or expansion) techniques according to embodiments of the present invention have proven to be efficient (optional) preprocessing or, for example, post-processing tools to enhance the enhancement results. The systems, methods, and apparatuses proposed according to embodiments of the present invention outperform other GAN-based temporal baselines while approaching state-of-the-art TF techniques. Further exploration of different techniques in the coupling layer and the combination of temporal and frequency domain signals can be implemented according to embodiments of the present invention.
[0382] Furthermore, it should be noted that the embodiments and procedures can be used as described in this section (as well as in “Problem Formulation,” “Normalization Process Basis,” “Speech Enhancement Process,” “Proposed Method,” “Model Architecture,” “Nonlinear Input Companding,” “Experiments,” “Data,” “Training Strategy,” “Model Setting,” “Evaluation,” and “Experimental Results”), and may optionally be supplemented individually and in combination by any features, functions, and details disclosed herein (throughout the document).
[0383] However, the features, functions, and details described in any other section may also be optionally incorporated into embodiments according to the invention.
[0384] Furthermore, the embodiments described in the foregoing sections can be used alone and may be supplemented by any features, functions, and details in another section.
[0385] Additionally, it should be noted that the various aspects described herein can be used individually or in combination. Therefore, details can be added to each of the individual aspects without having to add details to the other aspect.
[0386] In particular, embodiments are also described in the claims. The embodiments described in the claims may optionally be supplemented, individually and in combination, by any features, functions, and details described herein.
[0387] Furthermore, the features and functions disclosed herein related to the methods can also be used in an apparatus (configured to perform such functions). Additionally, any features and functions disclosed herein regarding the apparatus can also be used in the corresponding methods. In other words, the methods disclosed herein can be supplemented by any features and functions described regarding the apparatus.
[0388] Additionally, any features and functions described herein may be implemented in hardware or software, or a combination of hardware and software, as described in the "Implementation Alternatives" section.
[0389] To further summarize, embodiments of the present invention have created a system based on normalized process (NF) for the field of speech enhancement (e.g., using normalized process to directly model the enhancement process to build a speech enhancement framework, which includes, for example, learning the probability distribution of clean speech).
[0390] To further summarize, an embodiment of the invention creates a concept in which a process-based system is applied to speech enhancement, specifically by performing audio signal enhancement directly using the process-based system and independently of other algorithms to be combined, without degrading the performance of the audio signal enhancement or the quality of the resulting signal.
[0391] Furthermore, embodiments of the present invention provide a trade-off between efficient modeling and audio signal enhancement capabilities using neural networks for process-based audio signal processing.
[0392] Implement alternative solutions
[0393] Although some aspects are described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of method steps also represent a description of a corresponding block, item, or feature of the corresponding apparatus. Some or all of the method steps can be performed by (or using) hardware devices, such as, for example, a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps can be performed by such a device.
[0394] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) that stores electronically readable control signals thereon, which cooperates (or is capable of cooperating with) a programmable computer system to execute the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0395] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0396] Typically, embodiments of the present invention can be implemented as a computer program product having program code that, when run on a computer, is operable to perform one of the methods. The program code may, for example, be stored on a machine-readable medium.
[0397] Other embodiments include a computer program stored on a machine-readable medium for performing one of the methods described herein.
[0398] In other words, therefore, an embodiment of the method of the present invention is a computer program having program code that, when run on a computer, performs one of the methods described herein.
[0399] Therefore, a further embodiment of the method of the present invention is a data carrier (or digital storage medium, or computer-readable medium) comprising a computer program recorded thereon for performing one of the methods described herein. The data carrier, digital storage medium, or recording medium is typically tangible and / or non-transitional.
[0400] Therefore, a further embodiment of the method of the present invention represents a data stream or signal sequence for performing one of the methods described herein. The data stream or signal sequence may, for example, be configured to be transmitted via a data communication connection, such as via the Internet.
[0401] Further embodiments include processing equipment, such as a computer or programmable logic device, configured or adapted to perform one of the methods described herein.
[0402] Further embodiments include a computer having a computer program installed thereon for performing one of the methods described herein.
[0403] Further embodiments of the invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. For example, the receiver may be a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0404] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, these methods are preferably performed by any hardware device.
[0405] The apparatus described herein can be implemented using hardware devices, a computer, or a combination of hardware devices and a computer.
[0406] The apparatus described herein, or any component thereof, may be implemented, at least in part, in hardware and / or software.
[0407] The methods described in this article can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0408] The methods or any component of the apparatus described herein may be performed, at least in part, by hardware and / or software.
[0409] The embodiments described herein are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be readily apparent to those skilled in the art. Therefore, the intent is to be limited only by the scope of the pending patent claims, and not by the specific details presented through the description and explanation of the embodiments herein.
[0410] References:
[0411] [1]P Loizou, Speech Enhancement: Theory and Practice, CRC Press, 2edition, 2013.
[0412] [2]K.Borisagar,D.Thanki,and B.Sedani,Speech Enhancement Techniquesfor Digital Hearing Aids,Springer International Publishing,2018.
[0413] [3]A.H.Moore,P.Peso Parada,and P.A.Naylor,“Speech enhancement forrobust automatic speech recognition:Evaluation using a baseline system andinstrumental measures,”Computer Speech&Language,vol.46,pp.574–584,2017.
[0414] [4]J.Lim and A.Oppenheim,“All-pole modeling of degraded speech,”IEEETransactions on Acoustics,Speech,and Signal Processing,vol.26,no.3,pp.197–210,1978.
[0415] [5]S.Boll,“Suppression of acoustic noise in speech using spectralsubtraction,”IEEE Transactions on Acoustics,Speech,and Signal Processing,vol.27,no.2,pp.113–120,1979.
[0416] [6]Y.Ephraim and H.L.Van Trees,“A signal subspace approach for speechenhancement,”IEEE Transactions on Speech and Audio Processing,vol.3,no.4,pp.251–266,1995.
[0417] [7]Y.Ephraim and D.Malah,“Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,”IEEE Transactions Acoustics,Speech and Signal Processing,vol.33,pp.443–445,05 1985.
[0418] [8]Y.Xu,J.Du,L.Dai,and C.Lee,“A regression approach to speechenhancement based on deep neural networks,”IEEE / ACM Transactions on Audio,Speech,and Language Processing,vol.23,no.1,pp.7–19,2015.
[0419] [9]F.Germain,Q.Chen,and V.Koltun,“Speech denoising with deep featurelosses,”in Proc.Interspeech Conf.,2019,pp.2723–2727.
[0420]
[10] K.Qian,Y.Zhang,S.Chang,X.Yang,D.Florencio,and
[0421] M.Hasegawa-Johnson,“Speech enhancement using bayesian wavenet,”inProc.Interspeech Conf.,2017,pp.2013–2017.
[0422]
[11] S.Pascual,A.Bonafonte,and J.Serra`,“Segan:Speech enhancementgenerative adversarial network,”in Proc.Inter-speech Conf.,2017,pp.3642–3646.
[0423]
[12] M.H.Soni,N.Shah,and H.A.Patil,“Time-frequency masking-basedspeech enhancement using generative adversarial network,”in Proc.IEEEIntl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),2018,pp.5039–5043.
[0424]
[13] S.-W.Fu,C.-F.Liao,Y.Tsao,and S.-D.Lin,“Metricgan:Generativeadversarial networks based black-box metric scores optimization for speechenhancement,”in Proc.Intl.Conf.Ma-chine Learning(ICML),2019,pp.2031–2041.
[0425]
[14] S.Leglaive,X.Alameda-Pineda,L.Girin,and R.Horaud,“A RecurrentVariational Autoencoder for Speech Enhancement,”in Proc.IEEE Intl.Conf.onAcoustics,Speech and Signal Processing(ICASSP),2020,pp.371–375.
[0426]
[15] N.Adiga,Y.Pantazis,V.Tsiaras,and Y.Stylianou,“Speech enhancementfor noise-robust speech synthesis using wasserstein gan,”in Proc.InterspeechConf.,2019,pp.1821–1825.
[0427]
[16] H.Phan,I.V.McLoughlin,L.Pham,O.Y.Chen,P.Koch,M.De Vos,andA.Mertins,“Improving gans for speech enhancement,”IEEE Signal ProcessingLetters,vol.27,pp.1700–1704,2020.
[0428]
[17] A.A.Nugraha,K.Sekiguchi,and K.Yoshii,“A flow-based deep latentvariable model for speech spectrogram modeling and enhancement,”IEEE / ACMTransactions on Audio,Speech,and Language Processing,vol.28,pp.1104–1117,2020.
[0429]
[18] J.Ho,X.Chen,A.Srinivas,Y.Duan,and P.Abbeel,“Flow++:Improvingflow-based generative models with variational dequantization and architecturedesign,”in Proc.of Machine Learning Research,2019,vol.97,pp.2722–2730.
[0430]
[19] R.Prenger,R.Valle,and B.Catanzaro,“Waveglow:A Flow-basedGenerative Network for Speech Synthesis,”in Proc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),2019,pp.3617–3621.
[0431]
[20] I.Kobyzev,S.Prince,and M.Brubaker,“Normalizing flows:Anintroduction and review of current methods,”IEEE Trans-actions on PatternAnalysis and Machine Intelligence,pp.1–1,2020.
[0432]
[21] G.Papamakarios,E.T.Nalisnick,D.J.Rezende,S.Mohamed,and BLakshminarayanan,“Normalizing flows for probabilistic modeling andinference,”in arXiv:1912.02762,2019.
[0433]
[22] D.P.Kingma and P.Dhariwal,“Glow:Generative flow with invertible1x1 convolutions,”in Advances in Neural Information Processing Systems 31,2018,pp.10215–10224.
[0434]
[23] L Dinh,J.Sohl-Dickstein,and S.Bengio,“Density estimation usingreal NVP,”in 5th Int.Conf.on Learning Representations,ICLR,2017.
[0435]
[24] F.Chollet,“Xception:Deep learning with depthwise separableconvolutions,”in IEEE Conf.on Computer Vision and Pattern Recognition(CVPR),2017,pp.1800–1807.
[0436]
[25] B.Zhai,T.Gao,F.Xue,D.Rothchild,B.Wu,J.Gonzalez,and K.Keutzer,“Squeezewave:Extremely lightweight vocoders for on-device speech synthesis,”in arXiv:2001.05685,2020.
[0437]
[26] D.Rethage,J.Pons,and X.Serra,“A wavenet for speech de-noising,”inProc.IEEE Intl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),2018,pp.5069–5073.
[0438]
[27] J.Jensen,I.Batina,R.C Hendriks,and R.Heusdens,“A study of thedistribution of time-domain speech samples and discrete fouriercoefficients,”in Proc.SPS-DARTS,2005,vol.1,pp.155–158.
[0439]
[28] C.Valentini Botinhao,X.Wang,S.Takaki,and J.Yamagishi,“Speechenhancement for a noise-robust text-to-speech synthesis system using deeprecurrent neural networks,”in Proc.In-terspeech Conf.,2016,pp.352–356.
[0440]
[29] C.Veaux,J.Yamagishi,and S.King,“The voice bank corpus:Design,collection and data analysis of a large regional accent speech database,”inInt.Conf.Oriental COCOSDA heldjointly with the Conf.on Asian Spoken LanguageResearch and Evaluation(O-COCOSDA / CASLRE),2013,pp.1–4.
[0441]
[30] J.Thiemann,N.Ito,and E.Vincent,“The diverse environments multi-channel acoustic noise database(demand):A database of multichannelenvironmental noise recordings,”Proc.of Meetings on Acoustics,vol.19,no.1,pp.035081,2013.
[0442]
[31] D.P.Kingma and J.Ba,“Adam:A method for stochastic optimization,”in 3rd Int.Conf.on Learning Representations,ICLR,2015.
[0443]
[32] T.Salimans and D.P.Kingma,“Weight normalization:A simplereparameterization to accelerate training of deep neural net-works,”inAdvances in Neural Information Processing Systems 29,2016,pp.901–909.
[0444]
[33] M.Pariente,A.Deleforge,and E.Vincent,“A statistically principledand computationally efficient approach to speech enhancement usingvariational autoencoders,”in Proc.Inter-speech Conf.,2019,pp.3158–3162.
[0445]
[34] Y.Hu and P.Loizou,“Evaluation of objective quality measures forspeech enhancement,”IEEE Transactions on Audio,Speech,and LanguageProcessing,vol.16,pp.229–238,02 2008.
[0446]
[35] J.Hansen and B.Pellom,“An effective quality evaluation protocolfor speech enhancement algorithms,”in ICSLP,1998.
[0447]
[36] R.Giri,U.Isik,and A.A.Krishnaswamy,“Attention wave-u-net forspeech enhancement,”in Proc.IEEE Workshop on Applications of SignalProcessing to Audio and Acoustics,2019,pp.249–253.
[0448]
[37] Y.Koizumi,K.Yatabe,M.Delcroix,Y.Masuyama,and D.Takeuchi,“Speechenhancement using self-adaptation and multi-head self-attention,”in Proc.IEEEIntl.Conf.on Acoustics,Speech and Signal Processing(ICASSP),2020,pp.181–185.
Claims
1. A method for providing a processed audio signal based on an input audio signal (y, 130, 230, 430). Devices (100, 200, 400) of 160, 260, 460. The apparatus (100, 200, 400) is configured to use one or more process blocks (110) 1...N 210 1...N 410 1...N To process the noise signal (z, 120, 220, 420) or the signal derived from the noise signal (z, 120, 220, 420) in order to obtain the processed audio signal ( , 160, 260, 460), The devices (100, 200, 400) are configured to adapt to the input audio signal (y, 130, 230, 430) and use a neural network to utilize one or more process blocks (110). 1...N 210 1...N 410 1...N The processing performed.
2. The apparatus (100, 200, 400) according to claim 1, wherein the input audio signal (y, 130, 230, 430) is represented by a set of time-domain audio samples.
3. The apparatus (100, 200, 400) according to claim 1, wherein the one or more process blocks (110) 1...N 210 1...N 410 1...N The neural network associated with a given process block in the process block is configured to determine one or more processing parameters for the given process block based on the noise signal (z, 120, 220, 420) or a signal derived from the noise signal (z, 120, 220, 420) and based on the input audio signal (y, 130, 230, 430).
4. The apparatus (100, 200, 400) of claim 1, wherein the neural network associated with a given process block is configured to provide one or more parameters for affine processing, said parameters being applied during said processing to the noise signal (z, 120, 220, 420), or to a processed version of the noise signal, or to a portion of the noise signal (z, 120, 220, 420), or to a portion of a processed version of the noise signal.
5. The apparatus (100, 200, 400) of claim 4, wherein the neural network associated with the given process block is configured to determine one or more parameters of the affine processing based on a first portion of the process block input signal and based on the input audio signal (y, 130, 230, 430), and The affine processing associated with the given process block is configured to apply the determined parameters to a second portion of the process block input signal to obtain an affine processing signal. );and The first portion of the process block input signal and the affine processing signal (wherein the process block input signal is the first portion and the affine processing ... second portion and the affine processing signal is the third portion. The process block output signal (z) forms the given process block. new ).
6. The apparatus (100, 200, 400) of claim 5, wherein the neural network associated with the given process block includes depth-direction separable convolutions in the affine processing associated with the given process block.
7. The apparatus (100, 200, 400) according to claim 5, wherein the apparatus (100, 200, 400) is configured to apply a reversible convolution to the block output signal (z) of the given block. new ), to obtain the processed process block output signal (z') new ).
8. The apparatus (100, 200, 400) according to claim 1, wherein the apparatus (100, 200, 400) is configured to apply nonlinear compression (490) to the input audio signal (y, 130, 230, 430) before processing the noise signal (z, 120, 220, 420) according to the input audio signal (y, 130, 230, 430).
9. The apparatus (100, 200, 400) of claim 8, wherein the apparatus (100, 200, 400) is configured to apply the µ-law transformation as the nonlinear compression (490) to the input audio signal (y, 130, 230, 430).
10. The apparatus (100, 200, 400) according to claim 8, wherein the apparatus (100, 200, 400) is configured to apply a transformation according to the following equation to the input audio signal (y, 130, 230, 430). ; Where sgn() is a symbolic function; µ is a parameter that defines the compression level.
11. The apparatus (100, 200, 400) according to claim 1, wherein the apparatus (100, 200, 400) is configured to apply a nonlinear extension (415) to the processed audio signal. ,160,260,460).
12. The apparatus (100, 200, 400) according to claim 11, wherein the apparatus (100, 200, 400) is configured to apply the inverse µ-law transform as the nonlinear extension (415) to the processed audio signal. ,160,260,460).
13. The apparatus (100, 200, 400) according to claim 11, wherein the apparatus (100, 200, 400) is configured to apply a transformation according to the following equation to the processed audio signal ( , 160, 260, 460), ; Where sgn() is a symbolic function; µ is a parameter that defines the extension level.
14. The apparatus (100, 200, 400) according to claim 1, wherein the neural network parameters of the neural network for processing the noise signal (z, 120, 220, 420) or the signal derived from the noise signal are obtained by means of... The training audio signal or its processed version is processed in one or more training procedure blocks to obtain a training result audio signal, wherein the processing of the training audio signal or its processed version is adapted using the one or more training procedure blocks based on a distorted version of the training audio signal and the neural network. The neural network parameters are determined such that the characteristics of the training result audio signal are close to or include predetermined characteristics.
15. The apparatus (100, 200, 400) of claim 1, wherein the apparatus (100, 200, 400) is configured to provide neural network parameters for the neural network used to process the noise signal or the signal derived from the noise signal. The apparatus (100, 200, 400) is configured to use the one or more process blocks to process the training audio signal or a processed version thereof to obtain a training result audio signal. The apparatus (100, 200, 400) is configured to adapt the processing of the training audio signal or its processed version performed using the neural network based on a distorted version of the training audio signal and the neural network. The devices (100, 200, 400) are configured to determine the neural network parameters such that the characteristics of the training result audio signal approximate or include predetermined characteristics.
16. The apparatus (100, 200, 400) according to claim 1, wherein the apparatus includes means for providing neural network parameters. The means (100, 200, 400) for providing neural network parameters are configured to provide neural network parameters for the neural network that processes the noise signal or the signal derived from the noise signal. The means (100, 200, 400) for providing neural network parameters are configured to process the training audio signal or a processed version thereof using one or more training procedure blocks to obtain a training result audio signal. The means (100, 200, 400) for providing neural network parameters are configured to adapt the processing of the training audio signal or its processed version performed using the one or more process blocks, based on a distorted version of the training audio signal and using the neural network. The devices (100, 200, 400) are configured to determine the neural network parameters such that the characteristics of the training result audio signal approximate or include predetermined characteristics.
17. The apparatus (100, 200, 400) according to claim 1, wherein the one or more process blocks (110) 1...N 210 1...N 410 1...N The input audio signal (y, 130, 230, 430) is configured to synthesize the processed audio signal based on the noise signal (z, 120, 220, 420) under the guidance of the noise signal (z, 120, 220, 420). ,160,260,460).
18. The apparatus (100, 200, 400) according to claim 1, wherein the one or more process blocks (110) 1...N 210 1...N 410 1...N The input audio signal (y, 130, 230, 430) is configured to synthesize the processed audio signal based on the noise signal (z, 120, 220, 420) using affine processing of sample values of the noise signal (z, 120, 220, 420) or a signal derived from the noise signal (y, 130, 230, 430). , 160, 260, 460), The neural network is used to determine the processing parameters of the affine processing based on sample values of the input audio signal (y, 130, 230, 430).
19. The apparatus (100, 200, 400) of claim 1, wherein the apparatus (100, 200, 400) is configured to perform a normalization process to derive the processed audio signal from the noise signal (z, 120, 220, 420).
20. A method for providing a processed audio signal based on an input audio signal (y, 130, 230, 430), The method described therein includes using one or more process blocks (110) 1...N 210 1...N 410 1...N ) to process the noise signal (z, 120, 220, 420) or the signal derived from the noise signal to obtain the processed audio signal ( , 160, 260, 460); The method includes adapting one or more process blocks based on the input audio signal (y, 130, 230, 430) and using a neural network. The processing performed by (160, 260, 460).
21. An apparatus (500, 600, 800) for providing neural network parameters for audio processing. The devices (500, 600, 800) are configured to use one or more process blocks (510). 1...N 610 1...N 810 1...N The training audio signal (x, 505, 605, 805) or its processed version is processed to obtain the training result audio signal (z, 520, 620, 820). The devices (500, 600, 800) are configured to adapt to a distorted version (y, 530, 630, 830) of the training audio signal and to use a neural network to adapt to the use of one or more process blocks (510). 1...N 610 1...N 810 1...N The processing performed; The devices (500, 600, 800) are configured to determine the neural network parameters of the neural network such that the characteristics of the training result audio signal (z, 520, 620, 820) are close to or include predetermined characteristics.
22. The apparatus (500, 600, 800) according to claim 21. The devices (500, 600, 800) are configured to evaluate a cost function based on the characteristics of the obtained training result audio signals (z, 520, 620, 820), and The devices (500, 600, 800) are configured to determine neural network parameters to reduce or minimize the cost defined by the cost function.
23. The apparatus (500, 600, 800) of claim 21, wherein the training audio signal (x, 505, 605, 805) and / or the distorted version (y, 530, 630, 830) of the training audio signal is represented by a set of time-domain audio samples.
24. The apparatus (500, 600, 800) according to claim 21, wherein the one or more process blocks (510) 1...N 610 1...N 810 1...N The neural network associated with a given process block in the process block is configured to determine one or more processing parameters for the given process block based on the training audio signal (x, 505, 605, 805) or a signal derived from the training audio signal and based on the distorted version of the training audio signal (y, 530, 630, 830).
25. The apparatus (500, 600, 800) of claim 21, wherein the neural network associated with a given process block is configured to provide one or more parameters for affine processing, said parameters being applied during said processing to the training audio signal (x, 505, 605, 805), or to a processed version of the training audio signal, or to a portion of the training audio signal (x, 505, 605, 805), or to a portion of a processed version of the training audio signal.
26. The apparatus (500, 600, 800) of claim 25, wherein the neural network associated with the given process block is configured to determine one or more parameters of the affine processing based on a first portion of the process block input signal or based on a first portion of the preprocessed process block input signal and based on the distorted version (y, 530, 630, 830) of the training audio signal, and The affine processing associated with the given process block is configured to apply the determined parameters to a second portion of the process block input signal or to a second portion of the preprocessed process block input signal to obtain an affine processing signal. ;and The first part of the process block input signal or the first part of the preprocessing process block input signal and the affine processing signal ( The process block output signal x forms the given process block. new .
27. The apparatus (500, 600, 800) of claim 26, wherein the neural network associated with the given process block includes depth-direction separable convolutions in the affine processing associated with the given process block.
28. The apparatus (500, 600, 800) of claim 26, wherein the apparatus (500, 600, 800) is configured to apply a reversible convolution to the process block input signal of the given process block to obtain the preprocessed process block input signal.
29. The apparatus (500, 600, 800) of claim 21, wherein the apparatus (500, 600, 800) is configured to apply nonlinear input compression (815) to the training audio signal (x, 505, 605, 805) before processing the training audio signal (x, 505, 605, 805).
30. The apparatus (500, 600, 800) of claim 29, wherein the apparatus (500, 600, 800) is configured to apply the µ-law transformation as the nonlinear input compression (815) to the training audio signal (x, 505, 605, 805).
31. The apparatus (500, 600, 800) of claim 29, wherein the apparatus (500, 600, 800) is configured to apply a transformation according to the following equation to the training audio signal (x, 505, 605, 805). ; Where sgn() is a symbolic function; µ is a parameter that defines the compression level.
32. The apparatus (500, 600, 800) of claim 21, wherein the apparatus (500, 600, 800) is configured to apply nonlinear input compression (815) to the distorted version (y, 530, 630, 830) of the training audio signal before processing the training audio signal (x, 505, 605, 805) according to the distorted version (y, 530, 630, 830) of the training audio signal.
33. The apparatus (500, 600, 800) of claim 32, wherein the apparatus (500, 600, 800) is configured to apply the µ-law transformation as the nonlinear input compression (815) to the distorted version (y, 530, 630, 830) of the training audio signal.
34. The apparatus (500, 600, 800) of claim 32, wherein the apparatus (500, 600, 800) is configured to apply a transformation according to the following equation to the distorted version (y, 530, 630, 830) of the training audio signal. ; Where sgn() is a symbolic function; µ is a parameter that defines the compression level.
35. The apparatus (500, 600, 800) according to claim 21, wherein the one or more process blocks (510) 1...N 610 1...N 810 1...N The system is configured to convert the training audio signal (x, 505, 605, 805) into the training result audio signal (z, 520, 620, 820).
36. The apparatus (500, 600, 800) according to claim 21, wherein the one or more process blocks (510) 1...N 610 1...N 810 1...N The training audio signal (x, 505, 605, 805) is adjusted to use the sample values of the training audio signal (x, 505, 605, 805) or the signal derived from the training audio signal (x, 505, 605, 805) through affine processing. Guided by the distorted version (y, 530, 630, 830) of the training audio signal, the training audio signal (x, 505, 605, 805) is converted into the training result audio signal (z, 520, 620, 820). The neural network is used to determine the processing parameters of the affine processing based on sample values of the distorted version (y, 530, 630, 830) of the training audio signal.
37. The apparatus (500, 600, 800) of claim 21, wherein the apparatus (500, 600, 800) is configured to perform a normalization process to derive the training result audio signal (z, 520, 620, 820) from the training audio signal (x, 505, 605, 805).
38. A method for providing neural network parameters for audio processing. The method described therein includes using one or more process blocks (510) 1...N 610 1...N 810 1...N The training audio signal (x, 505, 605, 805) or its processed version is processed to obtain the training result audio signal (z, 520, 620, 820). The method includes adapting the training audio signal to a distorted version (y, 530, 630, 830) using a neural network and employing one or more process blocks (510). 1...N 610 1...N 810 1...N The process performed, The method includes determining the neural network parameters such that the characteristics of the training result audio signal (z, 520, 620, 820) are close to or include predetermined characteristics.
39. A computer program product having program code, wherein when the computer program product is run on a computer, the program code performs the method according to any one of claim 20 or claim 38.