Neural-network based speech de-coloration
By employing neural networks with a frequency-domain and time-domain loss function, the method effectively addresses speech degradation due to reverberation and noise, enhancing speech intelligibility and recognition performance.
Patent Information
- Application Number
- PCT/US2024/058956
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2024-12-06
- Publication Date
- 2025-06-26
AI Technical Summary
Existing speech communication systems face challenges in effectively removing reverberation, background noise, and other interferences from captured speech signals, leading to speech degradation and decreased performance in automatic speech recognition systems.
The use of neural networks trained with a cascade of stages that receive input in the frequency-domain representation of speech signals, utilizing a loss function with both frequency-domain and time-domain contributions to achieve loudness-preserving de-coloration, de-reverberation, and de-noising.
This approach efficiently recovers the original direct path of speech signals, significantly improving speech intelligibility and the performance of speech recognition systems while preserving loudness.
Smart Images

Figure US2024058956_26062025_PF_FP_ABST
Abstract
Description
[0001] NEURAL-NETWORK BASED SPEECH DE-COLORATION
[0002] Cross-Reference To Related Applications
[0003] This application claims the benefit of priority from PCT Application No. PCT / CN2023 / 140696 filed on 21 December 2023, and U.S. Provisional Application No. 63 / 551,145, filed on 8 February 2024, each of which is incorporated by reference herein in its entirety
[0004] Technical Field
[0005] The present disclosure relates to techniques for neural-network based de-coloration, de- reverbing, and / or de-noising of speech (e.g.. de-coloration, de-reverbing, and de-noising of speech). In particular, the present disclosure relates to loudness preserving de-coloration, de- reverbing, and / or de-noising of speech.
[0006] Background
[0007] When a speech signal is acquired in an enclosed space, the sound wave created by the signal hits the surrounding walls and other objects, where it may be reflected. Therefore, the observed signal consists of a superposition of a large number of delayed and attenuated copies of the speech signal. These multiple reflections can number several thousands and give rise to the effect known as reverberation. The resulting signal consists of the direct path and of multiple attenuated reflections, arriving with different delays with respect to the direct signal. In practice, excessive reflections or resonances within the enclosed space may result in speech degradation. Deleterious effects ty pically increase as the distance between the talker and the microphone increases.
[0008] For speech communication systems, such as voice-controlled systems, hands-free mobile telephones, and hearing aids, the received microphone signals are degraded by room reverberation, background noise, and other interferences. This signal degradation may lead to total unintelligibility of the speech and may decrease the performance of automatic speech recognition systems. For audio reproduction systems in daily life, one often listens to recordings made in a specific reverberant environment (e.g., a recording room) which are reproduced in a different room (e.g., a playback room). For a scenario where the recording is eventually played back over loudspeakers, the room acoustics of the playback room typically will affect the final acoustic signal that arrives at the listener’s ear. Two rooms in the reproduction chain can be mathematically described with a convolution of two Room Impulse Responses (RIRs), and therefore strong resonances or reverberation could potentially be increased, which causes further speech degradation.
[0009] There is thus need for improved techniques for removing dereverberation and other undesired effects from captured speech signals, or in other words, estimating the original direct path of the speech signal, without reflections.
[0010] Summary
[0011] In view of this need, the present disclosure provides methods of training neural networks for de-coloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech) and methods of de-coloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech), as well as corresponding apparatus, computer programs, and computer-readable storage media, having the features of respective independent claims.
[0012] One aspect of the present disclosure relates to a method of training a neural network for decoloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech). The neural network may include a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal. The method may include, for each neural network stage, determining a loss function based on an output of the neural network stage. The method may further include, for each neural network stage, adjusting neural network parameters of the neural network stage based on the determined loss function using backpropagation. The loss function may include a first contribution and a second contribution. Therein, the first contribution to the loss function may relate to a frequency-domain loss function that is based on (e.g., depends on) a frequency -domain representation of the output of the neural network stage. Further, the second contribution to the loss function may relate to a time-domain loss function that is based on (e.g., depends on) a time-domain representation of the output of the neural network stage.
[0013] By this choice of loss function during training, loudness-preserving de-coloration (potentially alongside, de-reverberation and / or de-noising) using a neural network can be achieved. Moreover, training of the neural network can be efficiently performed. In some embodiments, the frequency-domain loss function may be a complex-valued loss function. On the other hand, the time-domain loss function may be a real-valued loss function.
[0014] In some embodiments, the output of the neural network stage may relate to a speech component of the speech signal. This speech component may relate to a de-noised speech component, a de-noised and de-reverbed speech component (i.e., colored speech component), or a de-noised, de-reverbed, and de-colored speech component (i.e., anechoic speech component), for example.
[0015] In some embodiments, the frequency-domain loss function may be based on a complex ideal ratio mask (CIRM) of the speech component in relation to the speech signal in the frequency domain.
[0016] In some embodiments, the frequency -domain loss function may be based on a range- compressed version of the CIRM. This range-compressed version of the CIRM may be obtained by applying a hyperbolic tangent function to the CIRM, for example. The frequency -domain loss function may be based on a mean squared error of the range- compressed version of the CIRM and a ground truth (e.g., target) thereof.
[0017] In some embodiments, the time-domain loss function may include a source-to-distortion ratio (SDR) based on the speech component.
[0018] In some embodiments, the time-domain loss function may be based on (e.g., involve) an energy ratio involving a ground truth for the speech component and the speech signal. For example, the time-domain loss function may be based on (e.g., involve) an energy ratio between the ground truth (e.g.. target) for the speech component and the speech signal.
[0019] In some embodiments, the time-domain loss function may be based on a scaling factor (i.e., scaler). This scaling factor may be derivable from a specific loudness and a target loudness for the speech component. For example, the scaling factor may be a scaling factor for changing the loudness between the speech component and the target loudness. It may be derived by making the loudness for the scaled speech component equal to the target loudness.
[0020] Using the scaling factor enables training the neural network for loudness-preserving decoloration of speech.
[0021] In some embodiments, the output of the neural network stage may relate to a CIRM of the speech component in relation to the speech signal in the frequency domain. A time-domain representation of the output of the neural network stage may be obtained by applying an inverse time-frequency transform to the speech component.
[0022] In some embodiments, the frequency -domain representation of the speech signal for input to the cascade of one or more neural network stages may relate to a time-frequency transform of the speech signal.
[0023] In some embodiments, the CIRM per time-frequency slot (t, f) may be given by where is a frequency-domain representation of the speech component and Y (t, f) is the frequency-domain representation of a speech signal.
[0024] In some embodiments, the range-compressed version of the CIRM per time- frequency slot (t, f) may be given by where index d ∈ {r, i} denotes real and imaginary parts, and C and Q are positive real constants. Here, [— Q, Q] may be the compressed range for the CIRM. C may relate to a steepness constraint.
[0025] In some embodiments, the frequency -domain loss function may depend on the range- compressed version of the CIRM and on a ground truth M'(t, f) for the range-compressed version of the CIRM For example, the frequency-domain loss function may relate to a mean squared error between M’(t, f) and
[0026] In some embodiments, the time domain loss function L2may be given by where a is a time-domain representation of the speech component, a is a ground truth for â, y is a time-domain representation of the speech signal, ∈ [0,1] and β are non-negative real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and ||- 1| indicates the vector 1-norm. Here, the time-domain representation a of the speech component may be obtained by inverse transformation via â(t) =
[0027] In some embodiments, constant g may be given by Constant μ. may thus relate to the aforementioned energy ratio involving the ground truth for the speech component and the speech signal. In some embodiments, constant β may be chosen such that Ψtarget= Ψ(A / β), where A is a frequency -domain representation of the speech component and Ψ(·) indicates specific loudness. Constant (β may thus correspond to the aforementioned scaling factor.
[0028] In some embodiments, the cascade of one or more neural network stages may relate to a cascade of a first neural network stage and a second neural network stage. That is, the neural network may comprise a cascade of first and second neural network stages. Then, the output of the first neural network stage may relate to a colored speech component of the speech signal. Further, the output of the second neural network stage may relate to an anechoic speech component of the speech signal.
[0029] In some embodiments, the first neural network stage may be trained first and the second neural network stage may be trained for fixed neural network parameters of the first neural network stage.
[0030] In some embodiments, the loss function Lnet1for the first neural network stage may be given by where is the CIRM between a frequency -domain representation of the colored speech component and a frequencydomain representation Y (t, f) of the speech signal and is given by is a range-compressed version of is a ground truth for is a time-domain representation of the colored speech component, x is a ground truth for is a non-negative real constant. is the frequency-domain loss function, L1and L2is the time-domain loss function given by where y is a time-domain representation of the speech signal, μ1∈ [0,1] and β1are non-negahve real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and || • || indicates the vector 1-norm.
[0031] In some embodiments, the loss function Lnet2for the second neural network stage may be given by where is the CIRM between a frequency -domain representation of the anechoic speech component and a frequency-domain representation Y (t, f) of the speech signal and is given by is a range-compressed version of is a ground truth for is a time-domain representation of the anechoic speech component, ,q(t) is a frequency-independent gain, s is a ground truth for is a non-negative real constant, and the time-domain loss function is given by where μ2∈ [0,1] and β2are non-negative real constants.
[0032] In some embodiments, the cascade of one or more neural network stages may relate to a single neural network stage. That is, the cascade (or neural network) may comprise a single neural network stage. Then, an output of the neural network stage may relate to an anechoic speech component of the speech signal.
[0033] In some embodiments, the loss function L for the neural network stage may be given by L = where is the CIRM between a frequency¬ domain representation of the anechoic speech component and a frequency -domain representation Y(t,f) of the speech signal and is given by is a range-compressed version of is a ground truth for is a time-domain representation of the anechoic speech component, s is a ground truth for is a non-negative real constant, is the frequency-domain loss function, and L2is the time-domain loss function given by where y is a time-domain representation of the speech signal, μ ∈ [0,1] and β are non-negative real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and ||- 1| indicates the vector 1-norm.
[0034] In some embodiments, the neural network may include one or more of a convolutional neural network, a recurrent neural network, and a deconvolutional neural network.
[0035] Another aspect of the disclosure relates to a method of de-coloration, de-reverbing, and / or denoising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech) using a neural network. The neural network may include a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal. Further, the neural network may have been trained using the method according to the preceding aspect or any of its embodiments.
[0036] According to another aspect, an apparatus is provided. The apparatus may include a processor and a memory coupled to the processor and storing instructions for the processor. The processor may be configured to perform the methods or method steps outlined throughout the present disclosure. Specifically, the processor may be adapted to train a neural network for de-coloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech), or to implement a neural network for de-coloration, de-reverbing. and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech). In the latter case, the neural network may include a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal. Further, the neural network may have been trained using the method according to the aforementioned aspect or any of its embodiments.
[0037] According to a further aspect, a computer program is described. The computer program may comprise executable instructions for performing the methods or method steps outlined throughout the present disclosure when executed by a computing device (e.g., processor).
[0038] According to another aspect, a computer-readable storage medium is described. The storage medium may store a computer program adapted for execution on a computing device (e.g., processor) and for performing the methods or method steps outlined throughout the present disclosure when carried out on the computing device.
[0039] It should be noted that the methods and apparatus including its preferred embodiments as outlined in the present disclosure may be used stand-alone or in combination with the other methods and apparatus disclosed in this document. Furthermore, all aspects of the methods and apparatus outlined in the present disclosure may be arbitrarily combined. In particular, the features of the claims may be combined with one another in an arbitrary manner.
[0040] It will be appreciated that apparatus features and method steps may be interchanged in manyways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus, and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) (and, e.g., their steps) are understood to likewise apply to the corresponding apparatus (and, e.g., their blocks, stages, units), and vice versa.
[0041] Brief Description of the Drawings
[0042] The invention is explained below in an exemplary manner with reference to the accompanying drawings, wherein
[0043] Fig. 1 schematically illustrates an overview of an example framework for speech decoloration to which techniques according to embodiments of the disclosure may be applied; Fig. 2 is a flowchart schematically illustrating an example of a method of training a neural network according to embodiments of the disclosure;
[0044] Fig. 3 schematically illustrates an example of a tw o-stage neural netw ork according to embodiments of the disclosure;
[0045] Fig. 4 schematically illustrates an example of a single-stage neural network according to embodiments of the disclosure;
[0046] Fig. 5 is a flowchart schematically illustrating an example of a method of using a trained neural network for speech processing according to embodiments of the disclosure;
[0047] Fig. 6 schematically illustrates an example use case for techniques according to embodiments of the disclosure; and
[0048] Fig. 7 schematically illustrates an example of an apparatus for implementing techniques according to embodiments of the disclosure.
[0049] Detailed Description
[0050] In the following, example embodiments of the disclosure will be described with reference to the appended figures. Identical elements in the figures may be indicated by identical reference numbers, and repeated description thereof may be omitted.
[0051] Introduction and Notation
[0052] The acoustic channel between a source and a microphone can be described by a Room Impulse Response (RIR). The RIR relates to the signal that is measured at a microphone in response to a source that produces a delta impulse of sound. The RIR can be divided into three segments, viz., the direct path, early reflections, and late reflections, whose convolution with the desired signal results in the direct sound, early reverberation, and late reverberation, respectively. The early reflections appear as separate delayed impulses in the RIR, while late reflections appear as a continuum of exponential decay. This decay is a w ell-known property of the RIR. which has motivated the notion of reverberation time. Reverberation time may be defined as the time that is necessary to reach a 60 dB decay of the sound energy after switching off a sound source, in which case it is denoted by T60. Intuitively, the reverberation time gives an indication of the severity of reverberation within a room and is affected by the volume of the enclosed space and the acoustic properties of the reflecting surfaces of the room. Both excessive early and late reflections could cause distortion: early reflections with a time delay shorter than for example 50ms will give so-called Box-Klangfarbe coloration, while late reflections with a time delay longer than for example 50ms will give a perceived temporal flutter. It has been shown that the late reverberation components are the major cause of the degradation of speech intelligibility. Moreover, excessive early reflections may cause coloration that lead to the speech sounding as if coming from inside a bucket or a small box (called Box-Klangfarbe), which results in an undesired auditory illusion.
[0053] Broadly speaking, the present disclosure aims to provide a speech de-coloration algorithm to recover the direct speech while preserving the loudness, based on a deep complex neural- network and a loudness preserving loss function.
[0054] Next, a number of definitions necessary or helpful for describing techniques according to the present disclosure will be given.
[0055] Let y(t), s(t), h(t), and n(t) denote the noisy-reverberant speech, clean speech, RIR, and background noise, respectively. The noisy -reverberant speech y(t) can be written as y(t) = s(t) * h(t) + n(t) = x(t) + r(t) + n(t)
[0056] (1) where * stands for the convolution operator, x(t) denotes the direct sound (anechoic speech) that is colored by early reverberation, and r(t) denotes the late reverberation.
[0057] The spectrum of the noisy-reverberant speech signal can be derived by using a timefrequency transform such as. for example, a Fast Fourier Transform (FFT) or Complexvalued Quadrature Mirror Filter (CQMF) transformation. The observed (e g., captured, recorded) sound can be modeled in the time-frequency (T-F) domain as where g(t) is the frequency -independent gain due to propagation loss and C(t, f) models the coloration due to early reflections as well as the frequency-dependent response due to propagation. g(t) • S(t, f) is the anechoic speech and R(t, f) and N(t, f) are the frequencydomain representations of late reverberation and background noise, respectively.
[0058] Fig. 1 illustrates a framework 100 for recovering the anechoic speech g(t) • S(t, f) based on the observed noisy-reverberant speech Y (t, f) in the T-F domain. At T-F transform block 120, an input waveform (observed / captured / recorded signal) 110 is transformed to the T-F domain. De-coloration (and optionally, de-reverberation and / or de-noising) is applied to the T-F domain representation of the sound signal at de-coloration block 130. A transform inverse to that of transform of block 120 is applied to the de-colored signal at inverse transform block 160, yielding output waveform 170.
[0059] According to embodiments of the disclosure, de-coloration (e.g., at block 130) may be implemented using a neural network, and in particular a complex neural network. Such complex neural network may output both real and imaginary components.
[0060] The (complex) neural network may comprise (e.g., relate to, be implemented by) one or more of a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN). and a deconvolutional neural network.
[0061] Moreover, as will be described in more detail below, the structure of the (complex) neural network for recovering the anechoic speech g(t) • S(t, f) from the observed speech signal Y(t, f) can be for example
[0062] (1) a 2-stage cascaded neural network, where the aim of the first stage is to suppress the noise and late reverberation, and the second stage aims to recover the direct sound from the colored sound, or
[0063] (2) an end-to-end single-stage neural network that completes up to three tasks: noise suppression, late reverberation suppression, and de-coloration.
[0064] Thus, the neural network may be said to comprise a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal. This frequency -domain representation of the speech signal may relate to a time-frequency transform of the speech signal.
[0065] It is further understood that the output of each stage of the neural network relates to a speech component of the speech signal. This speech component may relate for example to a denoised speech component, a de-noised and de-reverbed speech component (i.e., colored speech component), or a de-noised, de-reverbed, and de-colored speech component (i.e.. anechoic speech component), depending on the implementation of the respective neural network stage.
[0066] Loss Function
[0067] Next, examples of loss functions for training a neural network designed for one or more of de-coloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech) will be described.
[0068] In general, the loss function may have two contributions, with the first contribution relating to a frequency-domain loss function and the second contribution relating to a time-domain loss function. Further, the frequency-domain loss function may be a complex-valued loss function. On the other hand, the time-domain loss function may be a real-valued loss function.
[0069] Complex-Valued Loss Function
[0070] The aforementioned frequency-domain loss function may be based on a complex ideal ratio mask (CIRM). For example, the frequency-domain loss function for a given neural network stage may be based on the CIRM of the speech component of interest in relation to the (input) speech signal in the frequency domain.
[0071] The complex-valued CIRM of a target signal A(t, f) (e.g., output speech component, speech component of interest) and an observed signal B(t, f) (e.g.. input speech signal) is defined as the ratio between the T-F spectrum of the target signal A(t, f) and the observed signal B(t, f), where subscripts r and i indicate the real or imaginary components of a complex value, respectively, and 6 indicates the phase of a complex value.
[0072] To improve training of the neural network, the frequency-domain loss function in some implementations may be based on a range-compressed version of the CIRM.
[0073] For example, the CIRM can be compressed to a desired range [— Q, Q] using the hyperbolic tangent to limit the value range of Mrand Mh where d ∈ {r, i} denotes the real or imaginary components, Mrand Miare compressed to be within [-Q, Q], and C is a steepness constraint.
[0074] The complex-valued loss function can be based on the compressed-CIRM (e.g., a target or ground truth for the compressed-CIRM) and the estimated compressed-CIRM (e.g., predicted compressed-CIRM)
[0075] For example, the complex- valued loss function L1may be a measure of a difference or error between the compressed-CIRM and the estimated compressed-CIRM (e.g.. predicted compressed-CIRM). This measure may be mean squared error (MSE) for example, or any specific designed function. Accordingly, in one implementation the frequency -domain loss function may be based on a mean squared error of the range-compressed version of the CIRM and a ground truth (e.g., target) thereof.
[0076] With the above definitions, for a neural network or neural network stage that receives a frequency -domain representation of a speech signal Y(t, f) as input (or that receives as input the output of an upstream neural network stage which in turn receives the frequency-domain representation of the speech signal Y(t, f) as input) and that seeks to predict a frequencydomain representation of a speech component A(t, f) (for example by predicting the CIRM), the relevant CIRM per time-frequency slot (t, f) may be given by
[0077] The range-compressed version of the predicted CIRM per time-frequency slot (t, f) may then be given by in line with the above, where index d ∈ {r, i} again denotes real and imaginary parts, and C and Q are positive real constants. [— Q, Q] again specifies the compressed range for the CIRM and C relates to a steepness constraint.
[0078] Finally, the frequency -domain loss function L1for the case at hand may depend on the range- compressed version of the CIRM and on a ground truth M'(t, f) for the range-compressed version of the CIRM For example, the frequency-domain loss function may relate to a mean squared error between M'(t, f) and Loudness Preservation Loss Function
[0079] The aforementioned time-domain loss function may include a Source-to-Distortion Ratio (SDR) based on the speech component. This SDR may be evaluated in the time domain.
[0080] In some cases, the output of the neural network or neural netw ork stage at hand may relate to the (predicted) CIRM of the speech component in relation to the speech signal in the frequency domain. Then, a time-domain representation of the output of the neural network or neural network stage can be obtained by deriving the speech component from the CIRM and applying an inverse time-frequency transform to the derived speech component.
[0081] That is, the time domain estimated signal (e.g., time domain predicted signal) can be recovered by inverse transforming the multiplication between the estimated CIRM (e.g., predicted CIRM) and the observed signal via
[0082] The time domain SDR loss function can then be a measure such as
[0083] In general, the source-to-distortion ratio loss LSDRmay be defined as for waveform vectors v1, v2, where (•,•) indicates the inner product and ||· || indicates the vector 1-norm.
[0084] In addition to the time domain SDR, the time domain loss function may be based on an energy ratio involving the speech signal and a ground truth for the speech component. For example, the time-domain loss function may be based on an energy ratio between the ground truth (e.g., target) for the speech component and the speech signal.
[0085] In some embodiments, the loudness preservation loss function (i.e., the forementioned time domain loss function), based on the time domain SDR loss function LSDR, can be defined for example as where μ = ||a||2 / (||a||2+ ||b — a||2) is an example of the aforementioned energy ratio between the target signal a and the observed signal b.
[0086] For achieving loudness preservation of the time domain loss function, the time-domain loss function may be further based on a scaling factor (or scaler) that is derivable from a specific loudness and a target loudness for the speech component. For example, the scaling factor may be a scaling factor for changing the loudness between the speech component and the target loudness. One way to derive the scaling factor may be to make the loudness for the scaled speech component equal to the target loudness.
[0087] In the above example of Eq. (9). the scaling factor is implemented by scaler p that controls the loudness but does not affect the SDR loss.
[0088] For example, the scaler β may be derived by stipulating that ΨTarget = Ψ{A / β}
[0089] (10) where Ψ{·} indicates the specific loudness, and p scales the loudness of the anechoic speech to be equal to the set target loudness ΨTargetso that the estimated signal a could be scaled by 1 / β. Then, in the inferencing stage, the estimated signal a will be optimized to approach a / p to achieve and preserve the target loudness.
[0090] With the above definitions, for a neural network or neural network stage that receives a frequency -domain representation of a speech signal Y(t, f) as input (or that receives as input the output of an upstream neural network stage which in turn receives the frequency-domain representation of the speech signal Y(t, f) as input) and that seeks to predict a frequencydomain representation of a speech component A(t, f) (for example by predicting the CIRM), the time domain loss function L2may be given by L2= μLSDR(a, βâ) + (1 — μ)LSDR(y — where a is the time-domain representation of the predicted speech component is a ground truth for â, y is a time-domain representation of the speech signal, and μ ∈ [0,1] and β are non-negative real constants. Here, the time-domain representation â of the speech component may be obtained by inverse transformation via as indicated above. Further, the constant μ mav be given by and thus may relate to the aforementioned energy ratio involving the ground truth for the speech component.
[0091] Constant β may be chosen as described above as the aforementioned scaling factor.
[0092] Method of Training a Neural Network for Speech De-Coloration
[0093] Fig. 2 is a flowchart schematically illustrating an example of a method 200 of training a neural network according to embodiments of the disclosure. In line with the above, the neural network may be trained for de-coloration, de-reverbing, and / or de-noising of speech (e.g., decoloration, de-reverbing, and de-noising of speech). Moreover, the neural network is understood to comprise a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal. Method 200 comprises steps S210 through S230 that may be performed for each neural network stage of the cascade of one or more neural network stages.
[0094] At step S210. the frequency-domain representation of the speech signal is input to the respective neural network stage.
[0095] At step S220. a loss function is determined (e.g., evaluated) based on an output of the neural network stage. That is, for a given functional configuration of the loss function, selected in accordance with the considerations provided above, the loss function is evaluated by applying it to the output of the neural network stage during training, yielding one or more loss function values.
[0096] Notably, the loss function evaluated at this step comprises a first contribution and a second contribution. The first contribution to the loss function relates to (e.g., is) the frequencydomain loss function as defined above, which is based on a frequency-domain representation of the output of the neural network stage. The second contribution to the loss function relates to (e.g., is) the time-domain loss function defined above, which is based on a time-domain representation of the output of the neural network stage.
[0097] At step S230. neural network parameters of the neural network stage are adjusted based on the determined loss function using backpropagation. This may involve determining one or more derivatives of the determined loss function value with respect to respective neural network parameters. Example Neural Network Structures for Speech De-Coloration
[0098] Next, example neural network structures that may be used in the context of the disclosure will be described. As indicated above, neural networks for speech processing (e.g., de-coloration, de-reverbing, and / or de-noising) according to embodiments of the disclosure may comprise a cascade of one or more neural network stages. For example, the neural network may be a two-stage neural network, as shown in Fig. 3, or a single-stage neural network, as shown in
[0099] Fig. 4.
[0100] Two-Stage Structure
[0101] Fig. 3 illustrates an example of a two-stage neural network structure 300 with stages 1 and 2. Accordingly, the aforementioned cascade of one or more neural network stages relates to a cascade of a first neural network stage (stage 1) 330 and a second neural network stage (stage 2) 350.
[0102] For two-stage neural network structure 300. input waveform 310 is transformed to the T-F domain at transform block 320 and subsequently input to the first neural network stage 330. The output of the first neural network stage 330 relates to a colored speech component of the speech signal. For example, the first neural network stage 330 may estimate (e.g., predict) a first CIRM between the colored speech component and the observed (e.g., captured) speech signal that is input as the input waveform 310.
[0103] Letting net1(. ) represent the neural network in the first neural network stage 330, its output may be given by, for example, where, is the estimated (e.g., predicted) CIRM between colored speech (e.g., a frequency -domain representation of the colored speech component) and the observed
[0104] (e.g., captured, recorded) signal (e.g., a frequency-domain representation Y(t, f ) of the speech signal) defined as
[0105] In the training stage, the total loss function 336 for the first neural network stage 330 is a (weighted) combination of two contributions 332, 334 to the loss function, namely a frequency -domain contribution 332 and a time-domain contribution 334 as described above. Weighting of the two contributions 332, 334 may be done by a weight γ1within [0, 1], for example.
[0106] For example, the loss function Lnet1for the first neural network stage 330 may be given by where is the range-compressed version of (e.g., compressed as per Eq. (4)), is a ground truth for is a time-domain representation of the colored speech component, x is a ground truth for is the non-negative real weighting constant, and L1is the frequency-domain loss function (e.g., as per Eq. (5)). Further, L2is the timedomain loss function given by where y is a time-domain representation of the speech signal and μ1∈ [0,1] and β1, are nonnegative real constants, with the SDR loss LSDRdefined above in Eq. (8). For example. and β1may be determined in line with the aforementioned determination of μ and β. respectively.
[0107] For determining the above overall loss function 336, the CIRM may be the output of the first neural network stage 330. The time-domain representation of the colored speech component can be obtained by inverse transformation, for example at inverse transform block 342, of the frequency domain colored speech component which in turn can be obtained by multiplying the CIRM by the frequency -domain representation Y (t, f) of the speech signal.
[0108] Training the first neural network stage 330 involves finding neural network parameters NN1that minimize the overall loss function 336 (e.g., the loss function given in Eq. (13)),
[0109] This can be achieved by known methods in the field of deep neural networks, including for example back propagation of errors or the like.
[0110] At inference, the colored speech component is optimized in the first neural network stage 330, by suppressing, for example, noise and late reverberation. Actual de-coloration of the speech signal at inference in the framework at hand is performed by the second neural network stage 350. an output of which relates to the anechoic speech component of the speech signal.
[0111] The input to the second neural network stage 350 is derived from the output of the first neural network stage 330. For example, said input may relate to the frequency domain representation of the colored speech component which may be obtained from the CIRM or via time-frequency transformation of the time-domain representation of the colored speech component for example at transform block 344. Notably, subsequent application of the inverse transform block 342 and transform block 344 may amount to applying identity operation 340.
[0112] The first and second neural network stages 330, 350 may be trained separately. After training netl, the parameters NN1of net1(·) may be fixed for training the second neural network stage 350. That is, the first neural network stage 330 is trained first and the second neural network stage 350 is trained for fixed neural network parameters N N1of the first neural network stage 330.
[0113] Letting net2(. ) represent the neural network in the second neural network stage 350, its output may be given by, for example, where is the estimated CIRM (e.g., predicted CIRM) between anechoic speech (e.g., a frequency-domain representation of the anechoic speech component) and the observed signal (e.g., the frequency -domain representation Y(t,f) of the speech signal) defined for example as
[0114] In the training stage, the total loss function 356 for the second neural network stage 350 is a (weighted) combination of two contributions 352, 354 to the loss function, namely a frequency -domain contribution 352 and a time-domain contribution 354 as described above. Weighting of the two contributions 352, 354 may be done by a weight y2within [0, 1], for example. By way of example, the loss function Lnet2for the second neural network stage 350 may be given by where is a range-compressed version of (e.g., compressed as per Eq. (4)), is a ground truth for is a time-domain representation of the anechoic speech component, g(t) is a frequency-independent gain, s is a ground truth for is a non-negative real constant, and L1is the frequency-domain loss function (e.g., as per Eq. (5)). Further, the time-domain loss function may be given by where μ2∈ [0,1] and β2are non-negative real constants with the SDR loss LSDRdefined above in Eq. (8). For example. μ2and β2may be determined in line with the aforementioned determination of μ and β, respectively.
[0115] For determining the above overall loss function 356, the CIRM may be the output of the second neural network stage 350. The time-domain representation of the anechoic speech component can be obtained by inverse transformation, for example at inverse transform block 360. of the frequency domain anechoic speech component which in turn can be obtained by multiplying the CIRM by the frequency-domain representation Y(t,f) of the speech signal.
[0116] Training the second neural network stage 350 involves finding neural network parameters NN2that minimize the overall loss function 356 (e.g., the loss function given in Eq. (18)),
[0117] This can be achieved by known methods in the field of deep neural networks, including for example back propagation of errors or the like.
[0118] At inference, the anechoic speech component is recovered from coloration (e.g., from the colored speech component) in the second neural network stage 350. The anechoic speech component may be output as waveform output 370. for example. Single-Stage Structure
[0119] Fig. 4 illustrates an example of an end-to-end single stage neural network structure 400 with a single neural network stage 430 that may complete noise suppression, late reverberation suppression, and de-coloration tasks in the single neural network stage 430. Accordingly, the aforementioned cascade of one or more neural network stages relates to a single neural network stage 430.
[0120] For single-stage neural network structure 400, input waveform 410 is transformed to the T-F domain at transform block 420 and subsequently input to the neural network stage 430. The output of the neural network stage 430 relates to an anechoic speech component of the speech signal. For example, the neural network stage 430 may estimate (e.g., predict) a CIRM between the anechoic speech component and the observed (e.g., captured, recorded) speech signal that is input as the input waveform 410.
[0121] Letting net(. ) represent the neural network in the neural network stage 430, its output may be given by, for example, where, is the estimated (e.g., predicted) CIRM between anechoic speech (e.g., a frequency -domain representation of the anechoic speech component) and the observed (e.g., captured, recorded) signal (e.g., a frequency-domain representation Y(t,f) of the speech signal) defined as
[0122] In the training stage, the total loss function 436 for the neural network stage 430 is a (weighted) combination of two contributions 432. 434 to the loss function, namely a frequency -domain contribution 432 and a time-domain contribution 434 as described above. Weighting of the two contributions 432, 434 may be done by a weight γ within [0, 1], for example.
[0123] By way of example, the loss function Lnetfor the neural network stage 430 may be given by where is a range-compressed version of (e.g., compressed as per Eq. (4)). M' is a ground truth for is a time-domain representation of the anechoic speech component, g(t) is a frequency-independent gain, s is a ground truth for γ ∈ [0,1] is a non-negative real constant, and L1is the frequency-domain loss function (e.g., as per Eq. (5)). Further, the time-domain loss function may be given by where μ ∈ [0,1] and β are non-negative real constants with the SDR loss LSDRdefined above in Eq. (8). For example, μ and β may be determined in line with the aforementioned determination of μ and β, respectively.
[0124] For determining the above overall loss function 436, the CIRM may be the output of the neural netw ork stage 430. The time-domain representation of the anechoic speech component can be obtained by inverse transformation, for example at inverse transform block 460, of the frequency domain anechoic speech component which in turn can be obtained by multiplying the CIRM by the frequency-domain representation Y(t,f) of the speech signal.
[0125] Training the neural network stage 430 involves finding neural network parameters NN that minimize the overall loss function 436 (e.g., the loss function given in Eq. (23)),
[0126] This can be achieved by known methods in the field of deep neural networks, including for example back propagation of errors or the like.
[0127] Thereby, the single-stage neural network directly leams the mapping between noisy reverberant speech to anechoic speech. By the above choice of the loss function, the single- stage neural network leams at the same time to preserve loudness.
[0128] At inference, the anechoic speech component is recovered by the single-stage neural network. The anechoic speech component may be output as waveform output 470, for example.
[0129] Method of Neural Network Based Speech De-Coloration
[0130] Fig. 5 is a flow-chart schematically illustrating an example of a method 500 of using a trained neural network for speech processing according to embodiments of the disclosure. Speech processing in this sense may relate to de-coloration, de-reverbing, and / or de-noising of speech (e.g., de-coloration, de-reverbing, and de-noising of speech).
[0131] At step S510. a frequency -domain representation of a speech signal is input to a neural network. This neural network comprises a cascade of one or more neural network stages, as described above.
[0132] At step S520. de-coloration, de-reverbing, and / or denoising (e.g., de-coloration, de-reverbing, and denoising) of the speech signal is performed using the neural network, as described above. Notably, the neural network (or its neural network stages) used for the processing may have been trained using methods described throughout the present disclosure.
[0133] Application Examples
[0134] Fig. 6 illustrates a de-coloration application example in the Metaverse where a plurality of speakers are rendered in the same virtual space 640. Speech 610 of each speaker is recorded in different environment and by using different devices, potentially resulting in interferes caused by reverberation and noise. To render speaker sounds in Metaverse, the RIR representing the Metaverse virtual space will be applied to all speaker sounds. When the recorded speaker sounds are heavily disturbed by reverberation, there will be two RIRs applied to the speaker sounds after rendering: one from the actual recording environment and the other from the virtual space, which causes the speaker sounds to not match the virtual space 640. The recorded noise can degrade the sound quality, which causes other people in the Metaverse virtual space 640 to feel uncomfortable. To improve the rendering effects, speaker sounds can be enhanced first by using the proposed de-coloration method 620 to remove interferences by reverberation and noise, and then the clean anechoic sound can be used to be rendered 630 in the Metaverse virtual space 640 to achieve improved performance.
[0135] As another example, in content creation, when the dialog track is disturbed by reverberation or noise due to the non-studio recording environment, it will degrade effects and flexibility for artist re-creation. The proposed de-coloration method can also be used to recover the clean anechoic dialog from the dialog track before the artist’s re-creation.
[0136] Apparatus, Programs, and Recording Media
[0137] While methods and process flows have been described above, it is understood that the present disclosure likewise relates to apparatus (e.g., computer apparatus or apparatus having processing capability in general) for implementing these methods and neural networks (or techniques in general). An example of such apparatus 700 is schematically illustrated in Fig. 7 The apparatus 700 comprises a processor 710 and a memory 720 coupled to the processor 710. The memory 720 may store instructions for execution by the processor 710. The processor 710 may be adapted to implement the apparatus or neural networks described throughout the disclosure and / or to perform methods (e.g., methods of training neural networks, methods of speech processing using neural networks) described throughout the disclosure. The apparatus 700 may receive inputs (e.g., captured speech signals or captured speech signals with corresponding target (ground truth) signals) and generate outputs (e.g., processed speech, sets of neural network parameters after training) as described throughout the disclosure.
[0138] The present disclosure further relates to programs (e.g., computer programs) comprising instructions that, when executed by a processor, cause the processor to carry out any of the methods described throughout the disclosure, and to computer-readable storage media storing such programs.
[0139] Interpretation
[0140] Aspects of the systems described herein may be implemented in an appropriate computer- based sound processing network environment (e.g., server or cloud environment) for processing digital or digitized audio files. Portions of these systems may include one or more networks that comprise any desired number of individual machines, including one or more routers (not shown) that serve to buffer and route the data transmitted among the computers. Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.
[0141] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmw are, and / or as data and / or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.
[0142] Specifically, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits ("ASICs'’). As such, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, computer-implemented neural networks described herein can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0143] While one or more implementations have been described by way of example and in terms of the specific embodiments, it is to be understood that one or more implementations are not limited to the disclosed embodiments. To the contrary, it is intended to cover various modifications and similar arrangements as would be apparent to those skilled in the art. Therefore, the scope of the appended claims should be accorded the broadest interpretation so as to encompass all such modifications and similar arrangements.
[0144] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including.” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
[0145] Enumerated Example Embodiments
[0146] Various Aspects and implementations of the invention may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0147] EEE 1. A method of training a neural network for de-coloration, de-reverbing. and / or de-noising of speech, wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal, the method comprising, for each neural network stage: determining a loss function based on an output of the neural network stage; and adjusting neural network parameters of the neural network stage based on the determined loss function using backpropagation, wherein the loss function comprises a first contribution and a second contribution, wherein the first contribution to the loss function relates to a frequency-domain loss function that is based on a frequency-domain representation of the output of the neural network stage, and wherein the second contribution to the loss function relates to a timedomain loss function that is based on a time-domain representation of the output of the neural network stage.
[0148] EEE 2. The method according to EEE 1, wherein the frequency-domain loss function is a complex-valued loss function.
[0149] EEE 3. The method according to EEE 1 or 2, wherein the output of the neural network stage relates to a speech component of the speech signal.
[0150] EEE 4. The method according to EEE 3, wherein the frequency-domain loss function is based on a complex ideal ratio mask, CIRM, of the speech component in relation to the speech signal in the frequency domain.
[0151] EEE 5. The method according to EEE 4, wherein the frequency-domain loss function is based on a range-compressed version of the CIRM.
[0152] EEE 6. The method according to any one of EEEs 3 to 5, wherein the time-domain loss function includes a source-to-distortion ratio. SDR. based on the speech component.
[0153] EEE 7. The method according to any one of EEEs 3 to 6, wherein the time-domain loss function is based on an energy ratio involving a ground truth for the speech component and the speech signal.
[0154] EEE 8. The method according to any one of EEEs 3 to 7, wherein the time-domain loss function is based on a scaling factor, wherein the scaling factor is derivable from a specific loudness and a target loudness for the speech component.
[0155] EEE 9. The method according to any one of EEEs 3 to 8, wherein the output of the neural netw ork stage relates to a complex ideal ratio mask, CIRM, of the speech component in relation to the speech signal in the frequency domain; and wherein a time-domain representation of the output of the neural network stage is obtained by applying an inverse time-frequency transform to the speech component. EEE 10. The method according to any one of the preceding EEEs, wherein the frequency-domain representation of the speech signal for input to the cascade of one or more neural network stages relates to a time-frequency transform of the speech signal.
[0156] EEE 11. The method according to EEE 4 or any EEE dependent on EEE 4. wherein the CIRM per time-frequency slot (t, f) is given by where is a frequency -domain representation of the speech component and Y (t, f) is the frequency-domain representation of a speech signal.
[0157] EEE 12. The method according to EEE 11 when depending on EEE 5, wherein the range-compressed version of the CIRM per time-frequency slot (t,f) is given by where index d ∈ {r, i} denotes real and imaginary parts, and C and Q are positive real constants.
[0158] EEE 13. The method according to EEE 12, wherein the frequency-domain loss function depends on the range-compressed version of the CIRM and on a ground truth M'(t, f) for the range-compressed version of the CIRM
[0159] EEE 14. The method according to any one of EEEs 3 to 13. wherein the time domain loss function L2is given by
[0160] L2= μLSDR(a,βâ) + (1 - μ)LSDR(y - a, y - â), where a is a time-domain representation of the speech component, a is a ground truth for a, y is a time-domain representation of the speech signal, / i e [0,1] and β are non-negative real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and ||· || indicates the vector 1-norm.
[0161] EEE 15. The method according to EEE 14, wherein constant is given by
[0162] EEE 16. The method according to EEE 14 or 15, wherein constant β is chosen such that Ψtarget= Ψ(A / B), where A is a frequency-domain representation of the speech component and Ψ(·) indicates specific loudness.
[0163] EEE 17. The method according to any one of the preceding EEEs, wherein the cascade of one or more neural network stages relates to a cascade of a first neural network stage and a second neural network stage; wherein the output of the first neural network stage relates to a colored speech component of the speech signal; and wherein the output of the second neural network stage relates to an anechoic speech component of the speech signal.
[0164] EEE 18. The method according to EEE 17, wherein the first neural network stage is trained first and the second neural network stage is trained for fixed neural network parameters of the first neural network stage.
[0165] EEE 19. The method according to EEE 16 or 17, wherein the loss function Lnet1for the first neural network stage is given by where is the CIRM between a frequency-domain representation of the colored speech component and a frequency -domain representation Y (t, f) of the speech signal and is given by is a range-compressed version of is a ground truth for is a time-domain representation of the colored speech component, x is a ground truth for is a non-negative real constant, is the L1frequency -domain loss function, and L2is the time-domain loss function given by where y is a time-domain representation of the speech signal, μ1∈ [0.1] and β1are nonnegative real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and ||· || indicates the vector 1-norm.
[0166] EEE 20. The method according to EEE 19, wherein the loss function Lnet2for the second neural network stage is given by where is the CIRM between a frequency -domain representation of the anechoic speech component and a frequency -domain representation Y(t, f) of the speech signal and is given by is a range-compressed version of is a ground truth for is a time-domain representation of the anechoic speech component. g(t) is a frequency -independent gain, s is a ground truth for γ2∈ [0,1] is a non-negative real constant, and the time-domain loss function is given by where μ2∈ [0,1] and p2are non-negative real constants.
[0167] EEE 21. The method according to any one of EEEs 1 to 16, wherein the cascade of one or more neural network stages relates to a single neural network stage, wherein an output of the neural netw ork stage relates to an anechoic speech component of the speech signal.
[0168] EEE 22. The method according to EEE 21, wherein the loss function L for the neural network stage is given by where is the CIRM betw een a frequency -domain representation of the anechoic speech component and a frequency -domain representation Y (t, f) of the speech signal and is given by is a range-compressed version of M' is a ground truth for is a time-domain representation of the anechoic speech component, s is a ground truth for is a non-negative real constant, L1is the frequency-domain loss function, and L2is the time-domain loss function given by where y is a time-domain representation of the speech signal, μ ∈ [0,1] and p are non- negative real constants, and the source-to-distortion ratio loss LSDRis given by for waveform vectors v1, v2, where indicates the inner product and ||· || indicates the vector 1-norm. EEE 23. The method according to any one of the preceding EEEs, wherein the neural network comprises one or more of a convolutional neural network, a recurrent neural network, and a deconvolutional neural network.
[0169] EEE 24. A method of de-col oration, de-reverbing, and / or de-noising of speech using a neural network, wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal; and wherein the neural network has been trained using the method according to any one of EEEs 1 to 23.
[0170] EEE 25. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of EEEs 1 to 23.
[0171] EEE 26. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to implement a neural network for de-coloration, de-reverbing, and / or de-noising of speech using a neural network. wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal; and wherein the neural network has been trained using the method according to any one of EEEs 1 to 23.
[0172] EEE 27. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of EEEs 1 to 24.
[0173] EEE 28. A computer-readable storage medium storing the program of EEE 27.
Claims
Claims1. A method of training a neural network for de-coloration, de-reverbing, and / or denoising of speech, wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal, the method comprising, for each neural network stage: determining a loss function based on an output of the neural network stage; and adjusting neural network parameters of the neural network stage based on the determined loss function using backpropagation, wherein the loss function comprises a first contribution and a second contribution, wherein the first contribution to the loss function relates to a frequency-domain loss function that is based on a frequency-domain representation of the output of the neural network stage, and wherein the second contribution to the loss function relates to a timedomain loss function that is based on a time-domain representation of the output of the neural network stage.
2. The method according to claim 1, wherein the frequency -domain loss function is a complex-valued loss function.
3. The method according to claim 1 or 2, wherein the output of the neural network stage relates to a speech component of the speech signal.
4. The method according to claim 3. wherein the frequency -domain loss function is based on a complex ideal ratio mask, CIRM, of the speech component in relation to the speech signal in the frequency domain.
5. The method according to claim 4, wherein the frequency -domain loss function is based on a range-compressed version of the CIRM.
6. The method according to any one of claims 3 to 5, wherein the time-domain loss function includes a source-to-distortion ratio, SDR, based on the speech component.
7. The method according to any one of claims 3 to 6, wherein the time-domain loss function is based on an energy ratio involving a ground truth for the speech component and the speech signal.
8. The method according to any one of claims 3 to 7, wherein the time-domain loss function is based on a scaling factor, wherein the scaling factor is derivable from a specific loudness and a target loudness for the speech component.
9. The method according to any one of claims 3 to 8, wherein the output of the neural network stage relates to a complex ideal ratio mask, CIRM, of the speech component in relation to the speech signal in the frequency domain; and wherein a time-domain representation of the output of the neural network stage is obtained by applying an inverse time-frequency transform to the speech component.
10. The method according to any one of the preceding claims, wherein the frequencydomain representation of the speech signal for input to the cascade of one or more neural network stages relates to a time-frequency transform of the speech signal.
11. The method according to claim 4 or any claim dependent on claim 4, wherein theCIRMper time-frequency slot (t,f) is given bywhere is a frequency -domain representation of the speech component and Y (t, f) is thefrequency -domain representation of a speech signal.
12. The method according to claim 11 when depending on claim 5, wherein the range-compressed versionof the CIRMper time-frequency slot (t,f) is given by where index d ∈ {r, i} denotes real and imaginary parts, and Cand Q are positive real constants.
13. The method according to claim 12, wherein the frequency-domain loss function L1depends on the range-compressed versionof the CIRMand on a ground truth M'(t, f) for the range-compressed version of the CIRM14. The method according to any one of claims 3 to 13, wherein the time domain loss function L2is given bywhere â is a time-domain representation of the speech component, a is a ground truth for â. y is a time-domain representation of the speech signal, μ ∈ [0,1] and β are non-negative real constants, and the source-to-distortion ratio loss LSDRis given byfor waveform vectors v1, v2, where indicates the inner product and ||· || indicates thevector 1-norm.
15. The method according to claim 14, wherein constant μ is given by16. The method according to claim 14 or 15, wherein constant p is chosen such that Ψtarget= Ψ(A / β), where A is a frequency-domain representation of the speech component and Ψ(·) indicates specific loudness.
17. The method according to any one of the preceding claims, wherein the cascade of one or more neural network stages relates to a cascade of a first neural network stage and a second neural network stage; wherein the output of the first neural network stage relates to a colored speech component of the speech signal; and wherein the output of the second neural network stage relates to an anechoic speech component of the speech signal.
18. The method according to claim 17, wherein the first neural network stage is trained first and the second neural network stage is trained for fixed neural network parameters of the first neural network stage.
19. The method according to claim 16 or 17, wherein the loss function Lnet1for the first neural network stage is given bywhere is the CIRM between a frequency-domain representationof thecolored speech component and a frequency -domain representation Y (t, f) of the speech signal and is given byis a range-compressed version ofis a ground truth foris a time-domain representation of the colored speech component, x is a ground truth for is a non-negative real constant, L1is thefrequency -domain loss function, and L2is the time-domain loss function given bywhere y is a time-domain representation of the speech signal, μ1∈ [0,1] and are nonnegative real constants, and the source-to-distortion ratio loss LSDRis given byfor waveform vectors v1, v2, whereindicates the inner product and ||· || indicates the vector 1-norm.
20. The method according to claim 19, wherein the loss function Lnet2for the second neural network stage is given bywhere is the CIRM between a frequency -domain representation of theanechoic speech component and a frequency -domain representation T(t, f) of the speech signal and is given by is a range-compressedversion of is a ground truth foris a time-domain representation of theanechoic speech component. g(t) is a frequency -independent gain, s is a ground truth fory2∈ [0,1] is a non-negative real constant, and the time-domain loss functionis given bywhere μ2∈ [0,1] and β2are non-negative real constants.
21. The method according to any one of claims 1 to 16, wherein the cascade of one or more neural network stages relates to a single neural network stage, wherein an output of the neural network stage relates to an anechoic speech component of the speech signal.
22. The method according to claim 21, wherein the loss function L for the neural network stage is given bywhere is the CIRM between a frequency -domain representationof theanechoic speech component and a frequency -domain representation Y(t, f) of the speech signal and is given byis a range-compressed version of M' is a ground truth foris a time-domain representation of theanechoic speech component, s is a ground truth foris a non-negative real constant, L1is the frequency-domain loss function, and L2is the time-domain loss function given bywhere y is a time-domain representation of the speech signal, μ ∈ [0,1] and β are nonnegative real constants, and the source-to-distortion ratio loss LSDRis given byfor waveform vectors v1, v2, where indicates the inner product and ||· || indicates thevector 1-norm.
23. The method according to any one of the preceding claims, wherein the neural network comprises one or more of a convolutional neural network, a recurrent neural network, and a deconvolutional neural network.
24. A method of de-coloration, de-reverbing, and / or de-noising of speech using a neural network, wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal; and wherein the neural network has been trained using the method according to any one of claims 1 to 23.
25. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 23.
26. An apparatus comprising a processor and a memory coupled to the processor and storing instructions for the processor, wherein the processor is adapted to implement a neural network for de-coloration, de-reverbing, and / or de-noising of speech using a neural network, wherein the neural network comprises a cascade of one or more neural network stages for receiving input of a frequency-domain representation of a speech signal; and wherein the neural network has been trained using the method according to any one of claims 1 to 23.
27. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 24.
28. A computer-readable storage medium storing the program of claim 27.
Citation Information
Patent Citations
Speech enhancement method and device, equipment and storage medium
CN115588437A