Method and system for reverberation modeling of speech signals
The method uses deep learning and convolutional prediction to estimate and remove reverberation from speech signals, improving speech quality and ASR performance by accurately separating the direct path signal, addressing the challenges of noisy and reverberant environments.
Patent Information
- Application Number
- JP2025523226
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-15
- Filing Date
- 2023-06-02
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2043-06-02
AI Technical Summary
Existing methods struggle to effectively remove reverberation from speech signals, particularly in noisy and reverberant environments, leading to degraded speech quality and inaccurate Automatic Speech Recognition (ASR) performance due to the difficulty in identifying and distinguishing the direct path signal from attenuated and delayed copies.
A method and system utilizing deep learning techniques, specifically a convolutional prediction approach with a first deep neural network (DNN) to estimate the direct path signal and a filter that models the room impulse response (RIR), followed by a second DNN for improved estimation, to reduce reverberation by discriminating and removing delayed and attenuated copies.
The proposed method significantly reduces reverberation, enhancing speech quality and improving ASR performance by accurately separating the direct path signal from reverberant signals, even in complex environments with noise and multiple speakers.
Smart Images

Figure 2025523269000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to audio signal processing, and more particularly to methods and systems for reverberation modeling of speech signals.
Background Art
[0002] Generally, in an enclosed room, reverberation of an audio signal (e.g., speech) occurs during modern hands-free speech communication such as remote conferencing and interaction with smart devices such as microphones of smart speakers. In such an enclosed room, the speech signal propagates through the air and may be reflected by walls, floors, ceilings, and other objects in the room before being captured by the microphone. Reverberation is the multipath propagation of the speech signal from the source or speaker to the receiving end such as the microphone. Such speech reverberation occurs when sound reflects from the surfaces in the environment. A portion of the sound may be absorbed by those surfaces, resulting in multiple attenuation of the speech signal. The reflection and absorption of sound by those surfaces can generate multiple attenuation copies and delay copies of the speech signal. These multiple attenuation copies and delay copies can degrade the quality of the speech and may interfere with the performance of an Automatic Speech Recognition (ASR) system or any speech / audio processing system. For example, the ASR may generate inaccurate output due to degraded-quality speech input.
[0003] Speech reverberation can be reduced by removing the effects of reverberation from the sound. The removal of such reverberation effects is known as reverberation removal. Reverberation removal may include identifying a direct path signal and distinguishing the direct path signal from an attenuated copy and a delayed copy. The direct path signal corresponds to the signal that sound travels when the source and the microphone are in a line of sight. However, especially when the reverberation is large and there is noise from an unsteady source, it may be difficult to identify the direct path signal and distinguish the direct path signal from the copy. For example, an environment such as a sealed room with an unsteady source such as an air conditioning system may have significant indoor reverberation. Due to noise from the air conditioning system or any multi-source environmental noise, it may be highly difficult to reduce reverberation. Multi-source environmental noise can also apply to a scenario where multiple people are talking in the environment.
[0004] Therefore, such reverberation modeling in a sealed environment is advantageous for use in multiple applications such as speech reverberation removal, indoor impulse response modeling, acoustic modeling, etc.
[0005] Therefore, it is necessary to overcome the above problems. More specifically, it is necessary to develop a method and a system for reverberation modeling of a speech signal while overcoming the reverberation state and unsteady noise in a reverberant environment.
Summary of the Invention
[0006] The objective of some embodiments is to develop a method and a system for reverberation modeling of a speech signal. Another objective of some embodiments is to perform reverberation removal of a speech signal using deep learning techniques. Reverberation removal of a speech signal can be extended for tasks such as reverberation reduction, speech enhancement, speaker separation, etc.
[0007] Some embodiments are based on the understanding that clean speech exhibits a spectral-temporal pattern. Such spectral-temporal patterns are unique patterns shown in the time-frequency domain and can provide a useful cue for reverberation reduction. Some of these patterns are derived from the structure of the speech signal itself, but some patterns may also correspond to the linear filter structure of the reverberation (i.e., reflection of sound waves) specific to the space, including all objects, structures or entities present in the physical space where the recording is made, as well as the positions of the source speech signal and receivers such as microphones that record the signal. Using this linear filter structure, the signal generated from the source signal at the microphone position and the reflection of the signal from the walls and surfaces of objects in the space or from people can be explained, and the linear filter structure represents the effect of reverberation on the input signal as a linear convolution of the input signal and the room impulse response (RIR). 。The input signal is the original source signal, also known as the dry source signal. The indoor impulse response represents the influence of the space on the input signal and everything within that space. For example, by playing an impulse sound (e.g., a blank shot or a balloon burst), which is a short - term time - domain signal, at the source position in a room and recording the resulting signal at the receiver position, an estimated value of the RIR between the source position and the receiver position can be recorded in a physical space such as a room. The impulse excites the room to produce a reverberant impulse signal, which can be used for estimating the RIR. Next, the reverberation of the dry source sound signal that would be reproduced at the same source position and recorded at the same receiver position can be modeled by convolving the dry source signal and the estimated RIR. For that purpose, the objective in some embodiments is to estimate a basic filter for approximating or modeling the RIR. In some exemplary embodiments, the RIR can be estimated based on a linear regression problem solved for each frequency in the time - frequency domain. The filter estimate for modeling the RIR can be used for identifying the delayed and attenuated copies of the input signal for reverberation removal of speech signals.
[0008] Furthermore, such a linear filter can be utilized as regularization to improve the reverberation removal process. For example, the linear filter as regularization prevents over - fitting of the model of the reverberation removal process to the training data. Some embodiments are based on the recognition that a combination of linear prediction and deep learning can utilize a linear filter structure for single - channel and multi - channel reverberant speaker separation and reverberation removal tasks. For that purpose, deep learning techniques supported by convolutional prediction can be used for reverberation removal in an environment with noise signals, reverberation of speech signals, etc. Convolutional prediction is a linear prediction method for speech reverberation removal in a reverberant situation, and source estimation obtained by a Deep Neural Network (DNN) It relies on values and utilizes a linear filter structure between the source estimate value and the reverberant version of the source signal in the observed input signal.
[0009] To obtain the source estimate, the DNN is trained in the time-frequency domain or the time domain to predict the target speech from the reverberant speech. The target speech corresponds to the target direct-path signal between the source and a receiver such as a microphone. This approach can utilize prior knowledge of speech patterns.
[0010] Previous studies have also attempted to utilize some form of linear filter structure to perform reverberation removal. For example, Weighted Prediction Error (WPE) may be used for reverberation removal of speech signals. The WPE method calculates an inverse linear filter based on distributed normalized delayed linear prediction. The calculated linear filter is applied to past observations of the mixture input signal including reverberation and possibly noise, and for reverberation removal, the late reverberation of the target source signal in the mixture input signal is estimated from past observations of the reverberation. The estimated late reverberation is subtracted from the acoustic signal mixture received from various sources, and the target speech signal in the acoustic signal mixture is estimated. In some embodiments, the filter can also be estimated using the time-varying power spectral density (PSD) of the target speech signal. PSD is the distribution of the power of a signal over the frequency domain of the signal. Such a linear filter can be repeatedly estimated using WPE in an unsupervised manner. However, the iterative procedure of WPE for filter estimation may lead to sub-optimal results and may be computationally expensive.
[0011] To overcome the above-mentioned deficiencies of WPE, the iterative procedure for filter estimation can be replaced by a DNN-based WPE (DNN-WPE) approach. DNN-WPE uses the amplitude estimated by the DNN as the PSD of the target speech signal for filter estimation. However, DNN-WPE cannot reduce early reflections. This is because DNN-WPE requires a strict non-zero frame delay to avoid trivial solutions and cannot have a mechanism to utilize the phase estimated by the DNN for filter estimation. Also, DNN-WPE may not be robust to interference caused by noise signals. For example, DNN-WPE may estimate a filter that associates past observations containing noise with the current observation containing noise, thereby limiting the filter estimation accuracy. Further, DNN-WPE directly uses the linear prediction result as its output, and as a result, the reduction of reverberation may be partial or minimal.
[0012] For this purpose, another object of some embodiments is to remove both early reflections and late reverberations for reverberation removal. Early reflections and late reverberations can be removed using a convolution prediction approach. The convolution prediction approach utilizes both the amplitude and phase estimated by the DNN for filter estimation. Also, the convolution prediction approach provides closed-form solutions for linear filters (similar to the above DNN-WPE approach), and these closed-form solutions are suitable for online real-time processing applications and can be jointly trained with other DNN modules such as acoustic models.
[0013] In some embodiments, a first estimate of the target direct path signal is used to determine a filter using a convolutional prediction approach, where application of the filter to the target direct path estimate results in a weighted distance function between the acoustic signal mixture and the first estimate of the target direct path signal being as close as possible. The filter models the reverberation of the target direct path signal within the mixture. Alternatively, and functionally equivalent, the filter can be applied to the acoustic signal mixture such that, under a weighted distance function, application of the filter to the target direct path estimate is as close as possible, in which case the filter models the reverberation including the direct path contribution that results in the target direct path signal. Since each filter can be readily obtained from the other filters, these filters are considered equivalent. Further, the filter is applied to the first estimate of the target direct path signal in the time-frequency domain. When the filter is applied to the first estimate of the target direct path signal, a result is obtained that discriminates the delayed and attenuated copies of the estimated direct path signal from the acoustic signal mixture. These delayed and attenuated copies are, herein, the derived signals of the target direct path signal reflected in multiple paths due to reverberation. For example, the target direct path signal is reflected in various directions by various objects within an environment such as a room. Such identified delayed and attenuated copies can be removed from the acoustic signal mixture for reverberation removal. Removal of the delayed and attenuated copies produces a mixture with reduced reverberation.
[0014] When the filter is applied to the first estimate of the target direct-path signal, the result obtained, by virtue of the above configuration, is closest to the residual between the acoustic signal mixture and the first estimate of the target direct-path signal according to the distance function. The distance function is a weighted distance between the filtered direct-path signal and the residual obtained by subtracting the direct-path estimate from the mixture, and the weight at each time-frequency point in the time-frequency domain is determined by one or a combination of the received acoustic signal mixture and the first estimate of the target direct-path signal. In some embodiments, the distance function is based on the least-squares distance. Further, the result of applying the filter to the first estimate of the target direct-path signal is removed from the acoustic signal mixture to obtain a mixture with reduced reverberation of the target direct-path signal.
[0015] In some embodiments, this reverberation-reduced mixture is input into a second DNN. The second DNN outputs a second estimate of the target direct path signal, which can be an improved estimate of the target direct path signal as compared to a previous estimate of the target direct path signal. The second DNN can perform steps similar to those of the first DNN. However, in some embodiments, the second DNN can take as input a different set of signals, such as one or a combination of the acoustic signal mixture, the reverberation-reduced mixture, and an estimate of the direct path signal. By similarly using the second estimate of the target direct path signal, an improved filter can be obtained using a convolutional prediction approach, where the application of the improved filter to the improved estimate of the target direct path signal results in a weighted distance function between the acoustic signal mixture and the improved estimate of the target direct path signal that is as close as possible. Alternatively and functionally equivalently, the improved filter can be such that the application of the improved filter to the improved estimate of the target direct path signal results in a weighted distance function between the acoustic signal mixture that is as close as possible, in which case the improved filter models the reverberation including the direct path contribution that results in the target direct path signal.
[0016] In some embodiments, the downstream DNN uses one or more of the filter estimates output by the first estimate or the improved filter estimates output by the second DNN to perform a task that benefits from the RIR estimate. For example, the downstream DNN is used to perform a localization function.
[0017] Also, some embodiments are based on the understanding that each individual speaker or each of a plurality of speakers is convolved with a different RIR. Obtaining each RIR individually may be relevant for downstream applications or for performing data augmentation. The WPE method estimates a single filter to reduce the reverberation of all sources. However, calculating a single filter for reverberation removal of a mixture may not be possible if the noise or competing speakers are louder than the target source. The filter thus calculated is biased towards suppressing the reverberation of the higher energy source. Therefore, it is necessary to estimate a reverberation removal filter for each source, because each source is convolved with a different RIR. The DNN-WPE method can calculate different filters for each source, but can only calculate different filters by using the estimated PSD of each source as a weight in the distance function that DNN-WPE uses for the estimation of the linear prediction filter, which may limit the accuracy and types of these different filters.
[0018] In some embodiments, for reverberation removal of a speech signal, two DNNs, a first DNN and a second DNN, are trained based on a convolutional prediction approach. First, the first DNN of the two DNNs outputs a first estimate of the direct-path signal of a target source (hereinafter referred to as the speaker, the person speaking) from an input, i.e., an acoustic signal mixture including the speaker's utterance. The direct-path signal of the target source is hereinafter referred to as the target direct-path signal. The first estimate of the target direct-path signal is used for the determination of a filter using the convolutional prediction approach. The filter is such that, under some weighted distance function, the application of the filter to the target direct-path estimate results in a residual that is as close as possible to the target direct-path estimate subtracted from the mixture. Further, the filter is applied to the first estimate of the target direct-path signal in the time-frequency domain. When the filter is applied to the first estimate of the target direct-path signal, a result is obtained that discriminates a delayed copy and an attenuated copy of the estimated target direct-path signal from the acoustic signal mixture. These delayed copy and attenuated copy are, in this specification, derived signals of the target direct-path signal reflected in a plurality of paths due to reverberation. For example, the target direct-path signal is reflected in various directions by various objects in an environment such as a room. Such identified delayed copy and attenuated copy are removed from the acoustic signal mixture for reverberation removal. Removal of the delayed copy and the attenuated copy produces a mixture with reduced reverberation.
[0019] When the filter is applied to the first estimate of the target direct-path signal, the result obtained, due to the above configuration, will be closest to the residual between the acoustic signal mixture and the first estimate of the target direct-path signal according to the distance function. In some embodiments, the distance function is based on the least-squares distance. Further, the result of applying the filter to the first estimate of the target direct-path signal is removed from the acoustic signal mixture, obtaining a mixture with reduced reverberation of the target direct-path signal. In some embodiments, this mixture with reduced reverberation is input to the second DNN out of the two DNNs. The second DNN outputs a second estimate of the target direct-path signal, and this second estimate can be an improved estimate of the target direct-path signal compared to the first estimate of the target direct-path signal.
[0020] In some embodiments, the first DNN can be trained for the purpose of speaker separation. For that purpose, the first DNN generates a plurality of outputs corresponding to the first estimate of the target direct-path signal for a certain speaker out of a plurality of speakers. Further, the estimation of the filter and the acquisition of the mixture with reduced reverberation are repeated for each of the plurality of speakers, generating a corresponding filter and a corresponding mixture with reduced reverberation for each of the plurality of speakers. Next, the corresponding mixtures with reduced reverberation for each of the plurality of speakers are combined, and the combined mixtures with reduced reverberation for each of the plurality of speakers are input to the second DNN. Next, the second DNN generates a second estimate of the target direct-path signal for each of the plurality of speakers.
[0021] Additionally or alternatively, the reverberation-reduced mixture, i.e., the delayed copy and the attenuated copy, may be utilized as additional features for a second DNN to determine a second estimate of the target direct-path signal, which improves echo cancellation. Additionally or alternatively, features corresponding to the delayed copy and the attenuated copy may also be used for the speaker separation task. In some exemplary embodiments, the delayed copy and the attenuated copy may be identified based on a linear regression problem. In some embodiments, one or a combination of the acoustic signal mixture and the first estimate of the target direct-path signal is provided as an input to the second DNN to generate a second estimate of the target direct-path signal. In some embodiments, the acoustic signal mixture, the first estimate, and the reverberation-reduced mixture are provided as inputs to the second DNN to determine a second estimate of the target direct-path signal.
[0022] Some embodiments are based on the recognition that, when there are multiple speakers in a room, corresponding filters are estimated for each individual speaker for echo cancellation. In the case of multiple speakers, the acoustic signal mixture includes speech signals from multiple speakers. In such a case, the first DNN generates corresponding first estimates of the target direct-path signal for each of the multiple speakers. To generate a reverberation-reduced mixture for each of the multiple speakers, the steps of determining a first estimate for each speaker, determining a filter for each speaker, and inputting one or a combination of the first estimate and the reverberation-reduced mixture for each speaker are combined and input to the second DNN to generate a second estimate of the target direct-path signal for each of the multiple speakers.
[0023] In some cases, the acoustic signal mixture may be received from a single channel such as a single microphone, or may be received from multiple channels such as an array of microphones. Each different channel measures a different version of the acoustic signal mixture. The DNN can be trained to estimate the reference channel or the target direct-path signal in each channel. The training can be based on composite spectral mapping in one or more channels. The DNN is trained to output an estimate of the target direct-path signal in the time-frequency domain in one or more channels such that the distance between the estimate and the reference in the time-frequency domain of the target direct-path signal in one or more channels is minimized. In the case of an array of microphones, a beamforming output can be obtained. The beamforming output can be obtained based on statistics calculated from a first estimate of the target direct-path signal at each microphone of the microphone array and one or a combination of the mixtures with reduced reverberation of the target direct-path signal. The beamforming output can be input into a second DNN to generate a second estimate of the target direct-path signal for each of a plurality of speakers. Additionally or alternatively, the beamforming output and the reverberation removal result may be used as additional features for the second DNN to perform better separation and reverberation removal tasks.
[0024] In some embodiments, the first DNN can be pre-trained to obtain a first estimate of the target direct-path signal from the observed acoustic signal mixture. The pre-training of the first DNN can be performed using a training dataset of acoustic signal mixtures and the corresponding reference target direct-path signals within this training dataset. In particular, the pre-training of the first DNN can be performed by minimizing a loss function. The loss function can include one or a combination of distance functions defined based on the real and imaginary (RI) components of the first estimate of the target direct-path signal in the complex time-frequency domain and the RI components of the corresponding reference target direct-path signal. Also, the distance function can be defined based on the magnitude obtained from the RI components of the first estimate of the target direct-path signal in the complex time-frequency domain and the corresponding magnitude of the reference target direct-path signal.
[0025] Additionally or alternatively, the distance function may be defined based on the reconstructed waveform obtained from the RI components of the first estimate of the target direct-path signal by reconstruction in the time domain and the corresponding waveform of the reference target direct-path signal.
[0026] In some alternative embodiments, the distance function may be defined based on the RI components of the first estimate in a second complex time-frequency domain obtained by further transforming the reconstructed waveform in a second time-frequency domain and the corresponding RI components of the reference target direct-path signal in the second time-frequency domain.
[0027] In some alternative embodiments, the distance function may be defined based on the magnitude obtained from the RI components of the first estimate in a second complex time-frequency domain obtained by further transforming the reconstructed waveform in a second time-frequency domain and the corresponding magnitude of the reference target direct-path signal in the second time-frequency domain.
[0028] In some exemplary embodiments, a first estimate of a target direct path signal can be replaced with a second estimate of the target direct path signal to obtain an updated first estimate of the target direct signal. The steps of obtaining the first estimate, obtaining a filter, and inputting the first estimate and a mixture with reduced reverberation can be repeated for the updated first estimate of the target direct signal to obtain an updated second estimate of the target direct signal.
[0029] In some examples, in a multi-speaker scenario, the above steps are repeated for each of the multiple speakers to generate corresponding filters for each of the multiple speakers. Further, by removing the reverberant speech of other speakers among the multiple speakers from the acoustic signal mixture, a portion of the received acoustic signal mixture corresponding to a certain speaker among the multiple speakers can be extracted. An estimate of the reverberant speech of another speaker among the multiple speakers is obtained by adding the first estimate of the target direct path signal for the other speaker to the result of applying the corresponding filter for the other speaker to the first estimate of the target direct path signal for the other speaker. After extraction, a filter for estimating the mixture with reduced reverberation for each speaker of the multiple speakers can be estimated based on the said portion of the received mixture.
[0030] Some embodiments provide evaluation results regarding speech reverberation removal and speaker separation, showing the effectiveness of reverberation removal of speech signals based on a convolutional prediction approach.
[0031] Accordingly, one embodiment of the present disclosure discloses a method executed by a computer to estimate an indoor impulse response (RIR). The method is executed by a processor coupled with stored instructions for implementing the method. The instructions, when executed by the processor, perform the steps of the method. The method includes receiving an acoustic signal mixture including a target direct path signal propagating in the indoor environment and reverberations of the target direct path signal in the indoor environment. The acoustic signal mixture is received via a wired communication channel or a wireless communication channel. The method further includes inputting the received acoustic signal mixture into a first deep neural network (DNN) to generate an estimated value of the direct path signal. The method further includes estimating a filter that models an indoor impulse response (RIR) representing the relationship between the direct path signal and the reverberations of the direct path signal, where the filter, when applied to the estimated value of the direct path signal, generates a result that is closest to one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimated value of the target direct path signal according to a distance function. Thereafter, the filter is transmitted. The filter is used in further other applications such as performing audio processing, generating extended training data, separating multiple speakers, etc.
[0032] In some embodiments, the first DNN is trained to estimate the amplitude and phase of the direct path signal, and the application of the filter to the estimated value of the direct path signal includes operations related to both the estimated amplitude and phase. Further, the distance function between the result of applying the filter to the estimated value of the direct path signal and the acoustic signal mixture is measured in both the amplitude domain and the phase domain.
[0033] In some embodiments, the coefficients of the filter are estimated in the time-frequency domain.
[0034] In various embodiments, the direct path signal is the signal transmitted from the source of the original signal to the sensor. To that end, the direct path signal represents the original signal at the source as would be measured by the sensor in the absence of reverberation, and the reverberation of the direct path signal represents one or more transmissions of the original signal from the source to the sensor over one or more paths longer than the shortest path such that the RIR is modeled with respect to the direct path signal.
[0035] In some embodiments, the filter that models the RIR is a linear filter estimated using a linear convolution prediction module. Linear convolution prediction enhances the linear structure of the filter to approximate the linear reflections of the direct path signal from the surfaces in the room.
[0036] Some embodiments provide applying a filter that models the RIR to an external signal measured outside the room to synthesize a new acoustic mixture, which results in the acoustic effects of the propagation of the external signal in the room. The new acoustic mixture can be used to update the first DNN.
[0037] Some embodiments provide updating an estimate of the direct path signal based on the application of a filter that models the RIR. The training data is further updated based on the acoustic signal mixture and the updated estimate of the direct path signal as a pseudo-label. The updated training data is used to retrain the first DNN based on the updated training data.
[0038] Various embodiments provide for creating extended data based on a filter that models an RIR and creating an extended training data set, the extended training data set obtained by applying a filter that models an RIR to data that includes one or more of external data measured outside a room and data obtained as one or more of a first measurement of a target direct-path signal of another acoustic signal mixture in training data and a second measurement of the target direct-path signal. Further, one or more of the first DNN and the second DNN can be trained to output a first estimate or a second estimate of the direct-path signal based on the extended training data set. Also, in some embodiments, the extended training data set can be used to train a system for automatic speech recognition or a system for audio event detection.
[0039] In some embodiments, a filter that models an RIR is used to perform acoustic analysis for at least one or a combination of indoor acoustic parameter analysis, indoor geometry reconstruction, voice enhancement, and reverberation removal of a speech signal. Indoor acoustic parameters include at least one of direct-to-reverberation ratio, reverberation time (RT60), early decay time (EDT), center time, clarity C80, and resolution D50.
[0040] Accordingly, another embodiment of the present disclosure discloses a system for estimating the RIR. The system comprises an input interface configured to receive an acoustic signal mixture including a direct path signal propagating indoors and reverberation of the direct path signal indoors via a wired communication channel or a wireless communication channel. The system further comprises a memory storing at least a first DNN. The system further comprises a processor configured to input the received acoustic signal mixture into the first DNN to generate an estimated value of the direct path signal. The processor is further configured to estimate a filter that models the RIR representing the relationship between the direct path signal and the reverberation of the direct path signal, and the filter, when applied to the estimated value of the direct path signal, generates a result closest to one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimated value of the target direct path signal according to a distance function. The system further comprises an output interface configured to output a filter that models the RIR via a communication channel.
[0041] Further features and advantages will become more readily apparent from the following detailed description when considered in conjunction with the accompanying drawings.
[0042] The present disclosure will be further described in the following detailed description with reference to a plurality of drawings as non-limiting examples of exemplary embodiments of the present disclosure. In the drawings, like reference numerals represent like parts throughout several views of the drawings. The drawings shown are not necessarily to scale and, overall, emphasis is placed on explaining the principles of the embodiments disclosed herein.
Brief Description of the Drawings
[0043]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 2C
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 8D
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Figure 15
Figure 16
Figure 17
Figure 18
Modes for Carrying Out the Invention
[0044] [Description of Embodiments] The above drawings illustrate the embodiments disclosed in this specification. However, as described in the description, other embodiments are also conceivable. The present disclosure is shown as an illustration rather than a limitation of exemplary embodiments. Many other modifications and embodiments that fall within the scope and spirit of the principles of the embodiments disclosed in this specification may be devised by those skilled in the art.
[0045] In the following description, for the sake of convenience of explanation, a number of specific details are set forth in order to enable a complete understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. In other instances, devices and methods are shown only in block diagram form to avoid obscuring the present disclosure. Various changes can be made in the functions and arrangements of the elements without departing from the spirit and scope of the subject matter disclosed as set forth in the appended claims.
[0046] As used in this specification and the claims, the terms "for example", "such as", and "such", as well as the verbs "comprising", "having", "including", and their other verb forms, when used with a list of one or more components or other items, are each to be construed as open-ended, meaning that the list should not be considered to exclude other additional components or items. The term "based on" means at least in part based on. Further, it should be understood that the phrases and terms used herein are for purposes of explanation and should not be considered limiting. Any headings used herein are for convenience only and have no legal or limiting effect.
[0047] Specific details are provided in the following description to provide a complete understanding of the embodiments. However, one skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form in order not to obscure the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail in order not to obscure the embodiments. Further, like reference numerals and names in the various drawings indicate like elements.
[0048] Although speech is used as the target sound source for most of the description, the same method can be applied to other types of audio signals.
[0049] [Overview of the System] FIG. 1A shows a representation of an environment 100A for reverberation modeling of a speech signal according to an embodiment of the present disclosure. The environment 100A may correspond to a closed environment having speakers 102 such as speaker 102A and speaker 102B. In FIG. 1A, a device 104 including at least a microphone or an array of microphones is also shown. In some exemplary embodiments, the device 104 may correspond to an automatic speech recognition (ASR) system, a voice signal processing system, or any voice processing system.
[0050] In the illustrated exemplary scenario, when the speaker 102 outputs voice, the corresponding acoustic speech signal may travel towards the device 104 via different paths. The acoustic speech signal may be linearly distorted by object reflections such as wall reflections and ceiling reflections as shown in FIG. 1A. In particular, the acoustic speech signal of the speaker 102 is distorted in a multipath direction before reaching the device 104, resulting in reverberation of the acoustic speech signal.
[0051] Therefore, the device 104 receives such an acoustic speech signal of the speaker 102A as an acoustic signal mixture. The acoustic signal mixture includes an anechoic speech signal and a reverberant speech signal. The anechoic speech signal is the target direct path signal 106A. Hereinafter, the reverberant speech signal, collectively referred to as reverberation 108A, includes non-direct path signals or multipath signals. There may be a plurality of speakers in the environment 100, such as when the speaker 102A is present together with another speaker 102B. In such a case, the acoustic signal mixture includes the target direct path signal 106B and a reverberant speech signal, hereinafter collectively referred to as reverberation 108B, corresponding to the speaker 102B. The acoustic signal mixture may include a reverberation noise signal 110A of a non-target source such as an air conditioner 110 in the environment 100A.
[0052] The speech signal of the speaker 102A and / or the speaker 102B may be blocked before reaching the device 104, which is shown in FIG. 1B.
[0053] Figure 1B shows an exemplary representation for reverberation modeling of a speech signal, according to another embodiment of the present disclosure. As shown in the environment 100B of Figure 1B, the speech signal of speaker 102A or speaker 102B is blocked by block 114 before reaching device 104. Block 114 may reverberate the speech signal of the corresponding speaker (such as speaker 102A or speaker 102B) in different directions. Such reverberation may increase the attenuated copy and the delayed copy (not shown in Figure 1B) of the speech signal of speaker 102A or speaker 102B. When device 104 is blocked by block 114, the speech signal of speaker 102A or speaker 102B may not have a corresponding target direct path signal. Instead, the speech signal may include a shortest path such as the shortest path 106C of the speech signal corresponding to speaker 102A and / or the shortest path 106D of the speech signal corresponding to speaker 102B. In such a situation, for the purpose of explanation in this specification, the shortest path signal is regarded as the target direct path signal, and the signal corresponding to a path longer than the shortest path is regarded as reverberation.
[0054] Device 104 may use a system 112 that may be integrated or incorporated within device 104 to reduce such reverberation, such as reverberations 108A and 108B. System 112 will be further described with reference to Figures 2A and 2B.
[0055] Figure 2A is a schematic block diagram showing a system 200a for reverberation removal of a speech signal, according to an embodiment of the present disclosure. System 200a corresponds to system 112 of Figures 1A and 1B.
[0056] In some exemplary embodiments, system 200a includes an input interface 202, a memory 204 storing a first deep neural network (DNN1) (e.g., DNN1206A), a processor 208, and an output interface 210. The input interface 202 and the output interface 210 are further coupled to a wired communication channel or a wireless communication channel for communication of system 200a with components external to system 200a.
[0057] The input interface 202 is configured to receive an acoustic signal mixture including a target direct path signal (e.g., target direct path signal 106A or target direct path signal 106B) and reverberations of the target direct path signal (e.g., reverberations 108A and / or reverberations 108B) over a wired communication channel or a wireless communication channel. In some exemplary embodiments, the input interface 202 may be configured to connect to at least the microphone of device 104, or an array of microphones of device 104.
[0058] The processor 208 inputs an acoustic signal mixture including target direct path signal 106A and reverberation 108A into DNN1206A. DNN1206A generates a first estimate of target direct path signal 106A. In a plurality of speaker scenarios including speakers 102A and 102B that generate sound signals in environment 100A or environment 100B, the target direct path signals corresponding to each of speakers 102A and 102B are estimated by DNN1206A. DNN1206A may generate corresponding estimates of the target direct path signal, one by one or simultaneously, for each of speakers 102A and 102B. For example, DNN1206A simultaneously determines a first estimate of target direct path signal 106A of speaker 102A and a first estimate of target direct path signal 106B of speaker 102B.
[0059] The first estimated value of the target direct path signal 106A is used with the received acoustic signal mixture to estimate a filter that models the room impulse response (RIR) for the first estimated value of the generated target direct path signal 106A. The RIR is the impulse response of the room, such as environment 100A or environment 100B, between the sound sources (e.g., speaker 102A and speaker 102B) and the microphone within device 104. The filter that models the RIR, when applied to the first estimated value of the target direct path signal, produces a result that is closest to at least one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimated value of the target direct path signal according to a distance function. This estimated filter that models the RIR is transmitted over the communication channel. To that end, the filter that models the RIR is output via output interface 210 and can then be transmitted over a wired or wireless communication channel connected to system 200a.
[0060] The filter that models the RIR is a linear filter estimated using a linear convolution prediction module. The linearity of this filter is modeled for the purpose of reducing computational complexity in applications related to reverberation removal of audio signals. Further, the linear filter provides a good approximation of the physical relationship between audio signals and their reflections. However, in reality, the reflection of sound is a non-linear phenomenon. Nevertheless, one of ordinary skill in the art can consider the choice of making the filter linear or non-linear to be a design choice and well within the scope of the present disclosure.
[0061] The linear filter is selected for applications where simplicity and accuracy are required. Thereby, linear convolution prediction enhances the linear structure of the filter to approximate the linear reflections of the direct path signal from the indoor surfaces.
[0062] In some embodiments, the filter is applied to a first estimate of the target direct path signal in the time-frequency domain. This is functionally and computationally superior to applying the filter to the first estimate of the target direct path signal in the time domain because the time-frequency domain enables an accurate modeling of the filter. Further, the estimation of the filter coefficients in the time-frequency domain provides a low-cost closed-form solution to the estimation problem. This may also be useful because other types of processing are performed in the time-frequency domain and there is no need to go back and forth between the time domain and the time-frequency domain.
[0063] When the estimated filter that models the RIR is applied to the first estimate of the target direct path signal 106A, a corresponding result is obtained. This result is closest to one of the acoustic signal mixture and the residual between the acoustic signal mixture and the first estimate of the target direct path signal according to a distance function. The distance function provides the distance between the residual between the acoustic signal mixture and the first estimate of the target direct path signal and the result of applying the filter to the first estimate of the target direct path signal. In some embodiments, the distance function may correspond to a weighted distance having weights at each time-frequency point in the time-frequency domain. The weights may be determined by one or a combination of the received acoustic signal mixture and the estimate of the target direct path signal. In an exemplary embodiment, the distance function may be based on the least squares distance.
[0064] In some embodiments, the first DNN, i.e., DNN1206A, is trained to estimate both the amplitude and the phase of the target direct path signal. Thus, the application of the filter estimate to the first estimate of the target direct path signal (which is also referred to, interchangeably and equivalently, as the estimated filter that models the RIR) involves operations on both the estimated amplitude and the phase, and the distance function between the result of applying the filter to the estimate of the direct path signal and the residual between the acoustic signal mixture and the direct path signal is measured in both the amplitude domain and the phase domain. By operating on both the amplitude and the phase, it becomes possible to use the information regarding the phase of the estimated direct path signal in the estimation of the filter, leading to an improvement in the accuracy of estimating the filter.
[0065] In some embodiments, the first DNN, i.e., DNN1206A, is pre-trained to obtain a first estimate of the target direct path signal from the observed acoustic signal mixture. The pre-training of DNN1206A is performed using a training data set of acoustic signal mixtures and the corresponding reference direct path signals within the training data set by minimizing a loss function.
[0066] When the result of applying the filter estimate to the first estimate of the target direct path signal 106A is removed from the acoustic signal mixture, a mixture with reduced reverberation of the target direct path signal 106A is obtained in order to achieve the purpose of reverberation modeling of the speech signal, like the acoustic signal mixture.
[0067] FIG. 2B is a schematic block diagram showing a system 200b for reverberation modeling of a speech signal according to an embodiment of the present disclosure. System 200b corresponds to system 112 of FIGS. 1A and 1B. System 200b includes all the components of system 200a and an additional second DNN 2206B.
[0068] To that end, system 200b includes an input interface 202, a memory 204 that stores a first DNN1 (e.g., DNN1 206A) and a second DNN2 (e.g., DNN2 206B), a processor 208, and an output interface 210.
[0069] When the result of applying the filter estimate to the first estimate of the target direct path signal 106A is removed from the acoustic signal mixture, a mixture with reduced reverberation of the target direct path signal 106A is obtained. The mixture with reduced reverberation of the target direct path signal 106A is provided as an input to DNN2 206B. DNN2 206B generates a second estimate of the target direct path signal 106A. The second estimate of the target direct path signal is used with the received acoustic signal mixture to estimate an improved filter (or improved filter estimate) that models the RIR of the target direct path signal with higher accuracy than the filter obtained using the first estimate. The improved filter estimate is output via the output interface 210.
[0070] Similarly, for speaker 102B, the first estimate of the target direct path signal 106B is used together with the received acoustic signal mixture to estimate a filter that models the RIR of the target direct path signal 106B. The filter estimate is applied to the first estimate of the target direct path signal 106B to obtain the corresponding result. This result is removed from the acoustic signal mixture to obtain a mixture with reduced reverberation of the target direct path signal 106B. The mixture with reduced reverberation of the target direct path signal 106B is input to a DNN 2206B that generates a second estimate of the target direct path signal 106B. The second estimate of the target direct path signal 106B of speaker 102B is used together with the received acoustic signal mixture to estimate an improved filter that models the RIR of the target direct path signal of speaker 102B with higher accuracy than the filter obtained using the first estimate. The improved filter that models the RIR is output via an output interface 210.
[0071] In some embodiments, a downstream DNN is used to apply one or more of the filter estimates output by the first DNN, i.e., DNN1206A, or the improved filter estimates output by the second DNN, i.e., DNN2206B, to perform different types of speech processing.
[0072] FIG. 2C shows a downstream DNN, i.e., DNN3206C, that is used to perform a plurality of tasks that are not necessarily related to speech transcription and / or separation. For example, the downstream DNN, i.e., DNN3206C, is used for tasks such as speaker identification, emotion recognition, event detection, localization, etc. in some embodiments. Therefore, for all of these tasks, having information regarding the reverberation of the target direct path signal, modeled as the RIR of the target direct path signal, is beneficial for training or inference purposes. In various embodiments, the downstream DNN utilizes the RIR of the filter determined by the first DNN or, if necessary, by the second DNN.
[0073] For example, in an application, the RIR is modeled as a filter estimate value at a location with settings different from the user's current location. For example, in the case of a movie, the RIR is determined for a church scene and then needs to be applied to a recording obtained in a small conference room. In this way, different settings of RIR can be combined with the RIR of the current location / settings to obtain a new acoustic effect. In this way, the reverberation in the church setting is synthesized as if it were applied in the same indoor setting. In this way, any number and type of sounds that provide various uses of the filter can be synthesized.
[0074] Also, as described above, the second estimate of the target direct path signal (such as the second estimate of the target direct path signal 106A or the second estimate of the target direct path signal 106B) is obtained as the reverberation-removed speech signal of the corresponding speaker (such as speaker 102A or speaker 102B). The reverberation modeling of the speech signal by any of the systems 200a, 200b, or 200c will be described in more detail with reference to FIGS. 3A, 3B, and 3C.
[0075] FIG. 3A is a schematic block diagram showing a process 300 for reverberation modeling of a speech signal according to an embodiment of the present disclosure. The process 300 is executed by the system 200. In an exemplary embodiment, the acoustic signal mixture 302 (Y) is received via the input interface 202 of the system 200. The acoustic signal mixture includes a target direct path signal such as the target direct path signal 106A of speaker 102A and the reverberation 108A of the target direct path signal 106A, or the target direct path signal 106B of speaker 102B and the reverberation 108B of the target direct path signal 106B, along with the reverberation of other sources such as the noise signal 110A of the device 110. The received acoustic signal mixture 302 is input into the DNN 1206A.
[0076] DNN 1206A determines a first estimate 304 of a target direct path signal, such as target direct path signal 106A or target direct path signal 106B. Further, a filter estimate 306 (hereinafter referred to interchangeably as filter 306) is determined to model an indoor impulse response (RIR) 308 for the first estimate 304 of target direct path signal 106A. The RIR model 308, hereinafter referred to as RIR 308, may correspond to the impulse response of an environment, such as environment 100A or environment 100B, between a source, such as speaker 102A and / or speaker 102B, and a receiver, such as device 104. For that purpose, the absolute delay and attenuation due to propagation from the source to the microphone are not modeled, and only the relative delay and attenuation with respect to the direct path signal are modeled. The impulse response is considered relatively with respect to the direct path signal received in the mixture, rather than with respect to the actual dry source signal at the source position. For ease of explanation, the application of filter estimate 306 to the direct path signal includes only the early reflections and late reverberations of the direct path signal and is such that the direct path signal is not included. The associated full filter estimate is equivalently obtained by modifying filter estimate 306 to further include the direct path signal. These two filter estimates are equivalent and one can be directly obtained from the other.
[0077] In some exemplary embodiments, the acoustic signal mixture 302 may correspond to a monaural signal recorded in an environment with reverberant noise, such as environment 100A or environment 100B. Such a monaural signal can be formulated in a physical model in the time domain. This physical model represents the relationship between the acoustic signal mixture 302 (y), the reverberant target speech signal (x) (including both a target direct path signal, such as target direct path signal 106A, and reverberation, such as reverberation 108A), and a non-target source (v) (such as device 110) including a reverberant noise signal (e.g., reverberant noise signal 110A) and a reverberant interfering speaker (e.g., speaker 102B).
[0078]
Number
[0079] The term "r d ", "r e " and "r l " represent the direct part, the initial part, and the late part of the RIR308 of the environment 100, respectively. The term "s" represents a target direct path signal (such as the target direct path signal 106A), and the target direct path signal is defined as s = a * r d . The term "h" represents an indirect path signal (e.g., reverberation 108A), and the indirect path signal is the sum of the early reflection a * r e and the late reverberation a * r l , that is, h = a * r e + a * r l = a * r e+l . The part r d+e of the RIR308 corresponding to both the direct path and the early reflection can be defined as a set of impulses up to 50 milliseconds after the direct path peak of r, and the early reflection component re of the RIR can be defined as re = r e = r d+e - r d . The filter for modeling the RIR in this application is considered in relation to r d . That is, the starting point of the time of the filter is implicitly considered to be the time of the impulse of r d , and the scaling of the elements of the filter is considered based on the height of the impulse of r d .
[0080]
Number
[0081] In Equation (2), the target direct path signal 106A represented as S(t,f) is estimated from the STFT coefficients (Y(t,f)) of the acoustic signal mixture 302 using a DNN. As the first estimate value 304 of the target direct path signal 106A, the restored target direct path signal 106A (S(t,f)) can be used.
[0082]
Number
[0083]
Number
[0084]
Number
[0085]
Number
[0086]
Number
[0087]
Number
[0088]
Number
[0089]
Number
[0090] The estimation of filter 306 can be improved by calculating a full filter estimation using Equation (7) when an estimation of the reverberant target speech X can be obtained. In some embodiments, the estimated reverberant speech for each speaker is repeatedly removed from the acoustic signal mixture 302 to refine the reverberant direct signal used in the estimation of filter 306.
[0091] In this embodiment, the formula (4) of the FCP can remove the reverberation related to the target speaker 102A. The ability to obtain the reverberation of the target speaker 102A can be particularly useful in a multi-speaker separation task because each target speaker can be convolved with a different RIR. For this purpose, in some embodiments, different filters may be calculated to remove the reverberation of each speaker (explained in FIG. 6). The estimated filter, for example, filter 306, may focus on reducing the reverberation of the target speaker 102A rather than the reverberation of a combination of another speaker (e.g., speaker 102B) and a non-target source (e.g., air conditioner 110). Even when there is a non-target source, the output of the DNN1206A such as the first estimated value 304 of the target direct path signal 106A and the mixture with reduced reverberation obtained using the filter 306 can be used for removing the reverberation of the speech signal. For this purpose, the first estimated value 304 and the mixture with reduced reverberation obtained using the filter 306 are input to the DNN2206B, and the second estimated value 314 of the target direct path signal 106A (or the target direct path signal 106B) can be output. The output generated by the DNN2206B such as the second estimated value 314 is superior to the output of the DNN1206A because the input to the DNN2206B (i.e., the first estimated value 304 and the mixture with reduced reverberation obtained using the filter 306) is more refined than the input to the DNN1206A. For example, the first estimated value 304 and the mixture with reduced reverberation obtained using the filter 306 output by the DNN1206A may have less interference. When the mixture with reduced reverberation obtained using these first estimated value 304 and the filter 306 with less interference is processed by the DNN2206B, the corresponding output (i.e., the second estimated value 314) can be superior to the output of the DNN1206A (i.e., the first estimated value 304).Accordingly, another iteration of the convolutional prediction can be performed using the second estimated value generated by the DNN2206B to obtain a second filter and a second mixture with reduced reverberation, and the second mixture with reduced reverberation can be input to the DNN2206B together with the second estimated value to generate a refined output.
[0092] In some exemplary embodiments, the corresponding RIR for each speaker, such as the RIR308, can be estimated by solving a linear regression problem for each frequency in the time-frequency domain or the time domain. The filter 306 that models the RIR308 can be used to identify the delayed and attenuated copies of the target direct path signals of the speaker 102A and / or the speaker 102B. The delayed and attenuated copies, which are repetitive patterns due to reverberation, can be removed from the received acoustic signal mixture 302. For this purpose, the filter 306 is applied to the first estimated value 304 to output a result 310. The result 310 can be closest to the residual between the acoustic signal mixture 302 and the first estimated value 304 of the target direct path signal based on a distance function such as a weighted least squares distance function. When the result 310 is removed from the acoustic signal mixture 302, a mixture 312 with reduced reverberation is obtained.
[0093]
Number
[0094] The removal of the result 310 reduces the delayed and attenuated copies from the mixture 312 with reduced reverberation. The delayed and attenuated copies can correspond to the late reverberation and the early reflections of the target direct path signal. These early reflections and late reverberation can be identified from the RIR308 modeled by the filter estimate 306. The RIR308 with early reflections and late reverberation is shown in FIG. 3B.
[0095] FIG. 3B shows an expression 316 of an indoor impulse response (RIR) model 316A for the source of the original signal from a speaker such as speaker 102A, showing the impulse corresponding to the target direct path signal 320A, the impulse corresponding to the early reflection 320B, and the impulse corresponding to the late reverberation 320C. In the present application, instead of the source of the original signal from the speaker, the target direct path signal is regarded as a reference. In other words, the application of the RIR to the target direct path signal results in the reverberation signal of the speaker, which is the sum of the target direct path signal and the early reflection and late reverberation of the target direct path signal.
[0096] FIG. 3C shows an application 326 of a filter 316B that models the RIR316A at a frequency bin f according to an embodiment of the present disclosure. The RIR model 316A corresponds to the RIR model 308, and the filter estimate 316B corresponds to the full filter estimate related to the filter estimate 306.
[0097] The RIR model 316A has a structure that can be represented as a sequence of impulses in the time domain. For example, the RIR model 316A is represented as a graph plot having an amplitude 318A and a number on the tap number axis 318B representing the time delay. The structure of the RIR model 316A is due to the reverberation in an environment such as environment 100, including the impulse corresponding to the target direct path signal 320A (r d ), and a plurality of impulses corresponding to the discrete early reflections 320B (r d ) of the target direct path signal 320A followed by the late reverberation 320C (r1) of the target direct path signal 320A (r e ). The target direct path signal 320A may correspond to the target direct path signal 106A or the target direct path signal 106B.
[0098] In some exemplary embodiments, the early reflections 320B and the late reverberations 320C are identified from the RIR model 316A. Assuming that a filter is modeled using K coefficients at each frequency f, the coefficients of the filter estimate 307 at frequency f are obtained by summing the results of multiplying the k-th coefficient by the time-frequency bin of the first estimate at time t-k+1 (for all k = 1, ..., K) at the same frequency f, such that the application 326 of the filter to the first estimate 304 of the target direct-path signal best approximates the reverberant mixture 322 at the same frequency f at the current time t.
[0099] As shown in the graph plot 316B, it represents approximating the acoustic signal mixture 322 (Y) by the application 326 of the K-tap filter 324 to the first estimate 304 of the target direct-path signal. The filter 324 is estimated by optimizing the forward filtering of the first estimate 304 of the target direct-path signal 302A. The filter 324 is an example of the filter 307. For example, the number of taps K of the filter 324 may be set to 40, which corresponds to a filter length of ((40-1)×8 + 32) ms in the time domain.
[0100] There can be different scenarios for reverberation modeling including reverberation removal of speech signals by the DNNs 1206A and 2206B. For example, the acoustic signal mixture 302 can be received by a single microphone or an array of microphones of the device 104 from a single speaker (e.g., speaker 102A) or from multiple speakers (e.g., speakers 102A and 102B). In the case of multiple speakers, the first DNN 1206A estimates different first estimates of the target direct-path signal for each of the multiple speakers. The reverberation removal of speech signals for different scenarios will be further described with reference to FIGS. 4, 5, and 6.
[0101] FIG. 4 is a schematic diagram showing an architectural representation 400 for reverberation modeling of a speech signal according to an embodiment of the present disclosure. As shown in FIG. 4, the architectural representation 400 includes a DNN1 402, a DNN2 406, and a convolutional prediction module 404 between the DNN1 402 and the DNN2 406. The DNN1 402 corresponds to the DNN1 206A, and the DNN2 406 corresponds to the DNN2 206B.
[0102]
Number
[0103]
Number
[0104]
Number
[0105] Some embodiments are based on the recognition that the second estimate 410 is superior to the first estimate 408 because the DNN 2406 processes a refined mixture of acoustic signals that is the reverberation-reduced mixture 412. The second estimate 410 can be further improved to function better than the first estimate 408. To that end, one or a combination of the acoustic signal mixture 302 and the first estimate 408 may be input to the DNN 2406 to generate the second estimate 410. In some cases, the acoustic signal mixture 302, the first estimate 408, and the reverberation-reduced mixture 412 may be input to the DNN 2406 to generate the second estimate 410. In other cases, the first estimate 408 and the reverberation-reduced mixture 412 may be input to the DNN 2406 to generate the second estimate 410. Further, the filter estimation, the acquisition of the reverberation-reduced mixture 412, and the input of the reverberation-reduced mixture 412 may be repeated to gradually refine the second estimate 410 of the target direct path signal 106A and improve the reverberation removal of the speech signal for the speaker 102A. This repetition may end when an end condition is met. The end condition may correspond to a condition defined by the user. Thus, the second estimate 410 will be superior to the first estimate 408 because it is refined with the reverberation 412-reduced mixture. In some embodiments, the DNN 2406 may be trained using the acoustic signal mixture 302, the reverberation-reduced mixture 412, and the first estimate 408 to output a second estimate 410 that improves the reverberation removal of the speech signal.
[0106] In the case of multiple speakers, the received acoustic signal mixture 302 may include speech signals from multiple speakers such as speaker 102A and speaker 102B. In such a case, the DNN 1402 can generate multiple outputs such as different first estimates of the target direct path signal, from which different filters can be obtained that model the corresponding RIRs for the multiple speakers, which will be further described with reference to FIG. 5.
[0107] FIG. 5 is a schematic diagram showing an architectural representation 500 for reverberation modeling of speech signals for a plurality of speakers (e.g., speakers 102A and 102B) according to some embodiments of the present disclosure. As shown in FIG. 5, the architectural representation 500 corresponds to a plurality of speaker scenarios and includes a plurality of instances of convolutional prediction modules such as DNN1 502, DNN2 506, and convolutional prediction modules 504A and 504B between DNN1 502 and DNN2 506. DNN1 502 corresponds to DNN1 206A, and DNN2 506 corresponds to DNN2 206B.
[0108]
Number
[0109]
Number
[0110] The reverberation-reduced mixtures 510A and 510B are concatenated and provided as an input to DNN2 506 to generate corresponding second estimates 512A and 512B for speaker 102A and speaker 102B. In some exemplary embodiments, DNN2 506 may receive the first estimate 508A together with the reverberation-reduced mixture 510A, the first estimate 508B together with the reverberation-reduced mixture 510B, and the acoustic signal mixture 302 as inputs to output the second estimates 512A and 512B.
[0111] In some exemplary embodiments, for each of the plurality of speakers 102A and 102B, to generate corresponding filters and a mixture with reduced reverberation, the first estimates 508A and 508B are replaced with the second estimates 512A and 512B, so that the filters for the first estimates 508A and 508B and the mixtures 510A with reduced reverberation and the mixtures 510B with reduced reverberation can be repeated. This repetition ends when a user-defined end condition is met. This end condition may include a user-defined end condition such as ending after three repetitions.
[0112] In some exemplary embodiments, the mixtures 510A with reduced reverberation and the mixtures 510B with reduced reverberation can be combined into a tensor. The tensor is a dimensional data structure representing all the mixtures with reduced reverberation of the plurality of speakers 102A and 102B. The tensor is input into the DNN 2506, and the second estimates 512A and the second estimates 512B corresponding to each of the plurality of speakers 102A and 102B are output.
[0113] The corresponding second estimates for each of the plurality of speakers 102A and 102B may be estimated one by one, which will be described in FIG. 6 below.
[0114] FIG. 6 is a schematic diagram showing an architectural representation 600 for reverberation modeling of speech signals of a plurality of speakers 102A and 102B according to some other embodiments of the present disclosure. As shown in FIG. 6, the architectural representation 600 corresponds to a multi-speaker scenario and includes a DNN 1602, a plurality of instances of a second DNN such as DNNs 2606A and 2606B, and a plurality of instances of convolution prediction modules such as convolution prediction modules 604A and 604B between the DNN 1602 and the plurality of instances of the second DNN such as DNNs 2606A and 2606B. The DNN 1602 corresponds to the DNN 1206A, and each of the DNNs 2606A and 2606B corresponds to the DNN 2206B.
[0115] DNN 1602 receives an acoustic signal mixture 302 and estimates a corresponding target direct-path signal for each of a plurality of speakers 102A and 102B. For example, DNN 1602 estimates a first estimated value 608A of the target direct-path signal 106A of speaker 102A. DNN 1602 estimates a first estimated value 608B of the target direct-path signal 106B of speaker 102B. The first estimated value 608A is input to the convolutional prediction module 604A, and the first estimated value 608B is input to the convolutional prediction module 604B.
[0116] The convolutional prediction module 604A estimates a filter that models the RIR of the first estimated value 608A. The filter that models the RIR is applied to the first estimated value 608A to obtain a mixture with reduced reverberation 610A of the target direct-path signal 106A. Similarly, the convolutional prediction module 604B estimates a filter that models the RIR of the first estimated value 608B. The estimated filter that models the RIR, output by the convolutional prediction module 604B, is applied to the first estimated value 608B to obtain a mixture with reduced reverberation 610B of the target direct-path signal 106B.
[0117]
Number
[0118] Furthermore, each of the reverberation-reduced mixtures is input into an instance of DNN2. Each of the reverberation-reduced mixture 610A and the reverberation-reduced mixture 610B is input into corresponding instances DNN2 606A and DNN2 606B (basically the same DNN2 but applied to different inputs), respectively. DNN2 606A generates a second estimate 612A of the target direct-path signal 106A of speaker 102B. DNN2 606B generates a second estimate 612B of the target direct-path signal 106B of speaker 102B. A plurality of instances of the second DNN, such as DNN2 606A and DNN2 606B, which generate respective second estimates 612A and 612B of corresponding speakers 102A and 102B, can be used to obtain clear speech of individual speakers from a plurality of speakers.
[0119] To improve the second estimate 612A and the second estimate 612B, DNN2 606A and DNN2 606B can be input with the acoustic signal mixture, the first estimates 608A and 608B, and one or a combination of the reverberation-reduced mixtures 610A and 610B.
[0120] In some exemplary embodiments, the first estimate 608A can be replaced with the second estimate 612A to generate an updated first estimate 608A of the target direct-path signal 106A. Similarly, the first estimate 608B can be replaced with the second estimate 612B to generate an updated first estimate 608B of the target direct-path signal 106B. Further, the estimation of the filter by DNN1 602, the estimation of the reverberation-reduced mixtures 610A and 610B, and the input of the reverberation-reduced mixtures 610A and 610B can be repeated to output updated second estimates of the target direct-path signal for each of the plurality of speakers 102A and 102B.
[0121] In some other exemplary embodiments, a portion of the acoustic signal mixture corresponding to a speaker (e.g., speaker 102A) can be extracted. This portion is extracted by removing the reverberant speech of other speakers, such as speaker 102B, from the acoustic signal mixture. An estimate of the reverberant speech of other speakers among a plurality of speakers is obtained by adding a first estimate of the target direct-path signal of the other speaker to the result of applying the corresponding filter of the other speaker to the first estimate of the target direct-path signal of the other speaker. After extracting the portion of the acoustic signal corresponding to speaker 102A, a filter for the first estimate of the extracted portion is estimated. This filter is used to estimate a mixture with reduced reverberation of speaker 102A based on the portion. The processing of the portion can improve the quality of the estimated filter for the speaker and the quality of the corresponding second estimate.
[0122] In some exemplary embodiments, the acoustic signal mixture of a single speaker 102A and / or multiple speakers 102A and 102B may be received from a single microphone or may be received from an array of microphones. Therefore, DNNs such as DNN1602 and DNN2606A and DNN2606B can be trained based on spectral mappings corresponding to a single microphone and an array of microphones. The spectral mapping trains DNN1602 to predict the real and imaginary (RI) components (i.e., frequencies) of an estimate, such as the first estimate 608A of the target direct-path signal 106A, from the RI components of the acoustic signal mixture 704. The RI components of the acoustic signal mixture 704 and the RI components of the first estimate 608A are input into DNN2606A, and a second estimate of the target direct-path signal 106A can be predicted. DNN1602 can be pre-trained using a training data set of the acoustic signal mixture and a training data set of the corresponding reference target direct-path signal within the training data set.
[0123] In some embodiments, the pre-training of the DNN 1602 can be performed by minimizing a loss function. This loss function can include one or a combination of distance functions defined based on the RI component of the target direct path signal 106A in the first time-frequency domain and the RI component of the reference target direct path signal in the first time-frequency domain. The reference target direct path signal can be obtained from the training data set of the utterance, and the corresponding reverberant mixture can be obtained by convolving the reference target direct path signal with a recorded RIR or a synthetic RIR and summing it with other interfering signals. The distance function can be defined based on the magnitude obtained from the RI component of the estimated target direct path signal in the first time-frequency domain and the corresponding magnitude of the reference target direct path signal.
[0124] In an alternative embodiment, the distance function may be defined based on the reconstructed waveform obtained from the RI component of the estimated target direct path signal in the first time-frequency domain by reconstruction in the time domain and the waveform of the reference target direct path signal. Also, the distance function may be defined based on the RI component in the complex time-frequency domain obtained by further converting the reconstructed waveform in the second time-frequency domain and the RI component of the reference target direct path signal in the second time-frequency domain. Also, the distance function may be defined based on the magnitude obtained from the RI component in the second time-frequency domain obtained by converting the reconstructed waveform in the second time-frequency domain and the corresponding magnitude of the reference target direct path signal in the second time-frequency domain.
[0125]
Number
[0126]
Number
[0127]
Number
[0128]
Number
[0129]
Number
[0130] In some exemplary embodiments, the acoustic signal mixture 302 may correspond to a multi-channel signal received from an array of microphones. Beamforming is performed on such multi-channel signals, which is further described with reference to FIG. 7.
[0131] FIG. 7 is a schematic diagram showing an architectural representation 700 for improving reverberation removal of a speech signal according to some embodiments of the present disclosure. The architectural representation 700 is similar to the architectural representation of FIG. 5, but further includes a plurality of instances of a Minimum Variance Distortionless Response (MVDR) beamforming module 704. In some exemplary embodiments, each instance of the MVDR module may output a beamforming output for a multi-channel signal. The beamforming filter can be obtained based on statistics calculated from one or a combination of a first estimate such as the first estimate 508A (and / or the first estimate 508B) output by a first DNN such as DNN1502, the reverberation-reduced mixture 510A (and / or the reverberation-reduced mixture 510B), and a second estimate such as the second estimate 512A (and / or the second estimate 512B) output by a second DNN such as DNN2506. The second estimate would have been obtained using a previous iteration of the architectural representation of FIG. 5 that includes only a convolutional prediction module between the two DNNs, or the architectural representation of FIG. 7 that includes MVDR beamforming. The beamforming output for the speaker can be obtained by applying the beamforming filter to the reverberation-reduced mixture 510A or the mixture 502. The MVDR beamforming module can be used between two DNNs such as DNN1502 and DNN2506. The output of the MVDR beamforming module such as the beamforming output 514A (and / or the beamforming output 514B) can be used as an input to a second DNN such as DNN2506. In some exemplary embodiments, the output of the MVDR beamforming module such as the beamforming output 514A can be combined with one or a combination of a first estimate such as the first estimate 508A, a reverberation-reduced mixture such as the reverberation-reduced mixture 510A, and the mixture 502.In some exemplary embodiments, the beamforming output for all speakers is used as an input to DNN2506, combined with a reverberation-reduced mixture for all speakers and a first estimate for all speakers. In some exemplary embodiments, the MVDR beamforming module may output beamforming using MVDR technology so as to be able to combine signals from multiple channels to derive an excellent estimate of the target direct path signal.
[0132] To that end, MVDR beamforming can be applied to the reverberation-reduced mixture to further improve the reverberation removal task and the separation task.
[0133]
Number
[0134]
Number
[0135] Furthermore, DNNs, such as DNN1602 and DNN2606A, can be easily replaced with magnitude or time domain models and more advanced DNN architectures. One such model will be further described with reference to FIGS. 8A, 8B, 8C, and 8D.
[0136] FIGS. 8A, 8B, 8C, and 8D are schematic diagrams showing a network architecture 800 for reverberation removal of speech signals according to some other embodiments of the present disclosure. The network architecture 800 corresponds to DNNs such as DNN1206A and DNN2206B.
[0137] The network architecture 800 is a Temporal Convolutional Network (TCN) 806. The TCN 806 includes four layers, and each of these layers has six dilated convolutional blocks such as dilated convolutional block 802A, dilated convolutional block 802B, dilated convolutional block 802C, dilated convolutional block 802D, dilated convolutional block 802E, and dilated convolutional block 802F (hereinafter referred to as dilated convolutional blocks 802A to 802F). In each of the dilated convolutional blocks 802A to 802F, one one-dimensional (1D) depthwise separable convolution 804 is used to reduce the number of parameters. For example, each of the dilated convolutional blocks 802A to 802F may include approximately 6.9 million parameters for reverberation removal of the speech signal. These large numbers of parameters can be reduced by the 1D depthwise separable convolution 804.
[0138] Furthermore, TCN806 is sandwiched by a U-Net including an encoder 808 and a decoder 810. In each of the encoder 808 and the decoder 810, DenseNet blocks are inserted at multiple frequency scales. The DenseNet block is an architecture that trains DNNs such as DNN1602 and DNN2606A using shorter connections between layers of the DNN. For example, the encoder 808 includes DenseNet blocks 808A, 808B, 808C, 808D, and 808E (hereinafter simply referred to as DenseNet blocks 808A to 808E) at multiple frequency scales. Similarly, the decoder 810 of the U-Net includes DenseNet blocks 810A, 810B, 810C, 810D, and 810E (hereinafter simply referred to as DenseNet blocks 810A to 810E) at multiple frequency scales. The U-Net can maintain a fine local structure through skip connections and model context information along the frequency through downsampling and upsampling. TCN806 utilizes the long-term information of the received acoustic signal mixture by using dilated convolution along the time domain. The DenseNet blocks 808A to 808E enable feature reuse and improve the discriminability of the speech signals of multiple speakers 102A and 102B in the speaker separation task.
[0139] The encoder 808 includes one two-dimensional (2D) convolution 812 and seven convolution blocks such as convolution block 814A, convolution block 814B, convolution block 814C, convolution block 814D, convolution block 814E, convolution block 814F, and convolution block 814G (hereinafter referred to as convolution blocks 814A - 814G). Each of the convolution blocks 814A - 814G includes a 2D convolution, an Exponential Linear Unit (ELU) non-linearity, and Instance Normalization (IN) to perform downsampling, that is, to reduce the sampling rate or sample size (bits per sample) of an input signal such as the audio signal mixture 704. The 2D convolution forms an essential component of feature extraction corresponding to the estimated value of the target direct path signal. The ELU is an activation function for DNNs (e.g., DNN 1602 and DNN 2606A), and the IN is a normalization layer to stabilize the hidden state dynamics in DNN 1602 and DNN 2606A.
[0140] The decoder 810 includes seven blocks of 2D transposed convolutions such as transposed convolution 816A, transposed convolution 816B, transposed convolution 816C, transposed convolution 816D, transposed convolution 816E, transposed convolution 816F, and transposed convolution 816G (hereinafter referred to as transposed convolutions 816A - 816G) together with the ELU, the IN, and one 2D transposed convolution 820 for upsampling by adding zero-value samples between the original samples to increase the sampling rate.
[0141] As described above, the mixtures with reduced reverberation of the plurality of speakers 102A and 102B (such as the reverberation-reduced mixture 510A and the reverberation-reduced mixture 510B) are represented by tensors. The tensor is in the form of featureMaps timeSteps frequencyChannels. Each of the convolutional blocks 814A to 814G (i.e., Conv2D + ELU + IN) and the transposed convolutional blocks 816A to 816G (i.e., Deconv2D + ELU + IN) is specified in the form of kernelSizeTime kernelSizeFreq, (stridesTime, stridesFreq), (paddingsTime, paddingsFreq) and featureMaps.
[0142] Each of the DenseNet blocks 808A to 808E, such as DenseBlock(g1,g2), includes five Conv2D + ELU + IN blocks having the growth rate g1 of the first four layers and the growth rate g2 of the last layer of the DenseNet blocks 808A to 808E. The tensor shape after each TCN block 806 is in the form of featureMaps timeSteps. Each IN + ELU + Conv1D block is specified in the form of kernelSizeTime, stridesTime, paddingsTime, dilationTime, featureMaps.
[0143] FIG. 9A is a flowchart showing a method 900a for estimating an RIR for reverberation modeling of a speech signal according to an embodiment of the present disclosure. The method 900a is executed by the system 200. The method 900a includes, in operation 902, receiving, via a communication channel, an acoustic signal mixture (such as the acoustic signal mixture 302) including a target direct path signal (e.g., the direct path signal 106A) and the reverberation of the target direct path signal. The acoustic signal mixture may include at least one of a single-channel signal or a multi-channel signal received from a single microphone or an array of microphones connected to an input interface connected to the communication channel.
[0144] In operation 904, the received acoustic signal mixture is input into a first DNN, such as DNN 1206, to generate a first estimate of the target direct path signal 106A (e.g., first estimate 408). In a multi-speaker scenario, the first DNN determines a corresponding first estimate for each of the plurality of speakers. The corresponding first estimates may be determined one by one for each of the plurality of speakers or may be determined simultaneously for the plurality of speakers. In some embodiments, based on the training data set of the acoustic signal mixture and the corresponding reference target direct path signal in the training data set, the first DNN may be pre-trained to generate a first estimate from the observed acoustic signal mixture. The pre-training of the first DNN may be performed by minimizing a loss function.
[0145] In operation 906, a filter (e.g., filter 306) that models an indoor impulse response (RIR) (e.g., RIR model 308) is estimated for a first estimate 408 of the target direct path signal 106A, and the filter is estimated such that the result of applying the filter to the first estimate of the target direct path signal is closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal according to a distance function (e.g., least squares distance function). In some embodiments, the filter corresponds to a linear filter structure estimated based on convolutional prediction. The first estimate is filtered forward frequency by frequency in the time-frequency domain using a linear filter of convolutional prediction (described in FIGS. 3A, 3B, 4, 5, and 6). In some exemplary embodiments, the received acoustic signal mixture includes speech signals from multiple speakers. The first DNN generates a plurality of outputs, each output including a first estimate of the target direct path signal for a certain speaker from the multiple speakers. In some embodiments, the early reflections (e.g., early reflection 320B) and late reverberations (e.g., late reverberation 320C) of the first estimate can be identified based on the RIR modeled by the filter. The identified early reflections and late reverberations can be removed from the first estimate to estimate a mixture with reduced reverberation.
[0146] In operation 908, the estimated filter that models the RIR is transmitted over a communication channel. The communication channel is any combination of wired communication channels or wireless communication channels required for each application of the filter after transmission. To that end, method 900a may be extended with additional steps of processing, as described in FIG. 9B below.
[0147] FIG. 9B is a flowchart showing another method 900b for reverberation removal of speech signals according to an embodiment of the present disclosure. Method 900b is executed by system 200.
[0148] Method 900b includes, in operation 910, receiving, via an input interface, an acoustic signal mixture (e.g., acoustic signal mixture 302) that includes a target direct path signal (e.g., target direct path signal 106A) and reverberation of the target direct path signal. The acoustic signal mixture can include at least one of a single-channel signal or a multi-channel signal that can be received from a single microphone or an array of microphones connected to the input interface.
[0149] In operation 912, the received acoustic signal mixture is input into a first DNN, such as DNN 1206, to generate a first estimate (e.g., first estimate 408) of the target direct path signal 106A. In a scenario of multiple speakers, the first DNN determines a corresponding first estimate for each of the multiple speakers. The corresponding first estimates may be determined one by one for each of the multiple speakers, or may be determined simultaneously for the multiple speakers. In some embodiments, the first DNN may be pre-trained to generate the first estimate based on at least one of the observed acoustic signal mixture, or a training data set of the acoustic signal mixture, and the corresponding reference target direct path signal within the training data set. The pre-training of the first DNN may be performed by minimizing a loss function.
[0150] In some embodiments, the first estimate of the target direct path signal is updated based on the application of a filter that models the RIR. Next, the updated first estimate and the acoustic signal mixture are used to update the training data of the training data set. The updated first estimate of the target direct path signal is used as a pseudo-label to identify the target direct path signal in the updated training data set. Further, the first DNN is then re-trained using this updated training data set.
[0151] In operation 914, a filter (e.g., filter 306) that models an indoor impulse response (RIR) (e.g., RIR model 308) is estimated for a first estimate 408 of the target direct path signal 106A such that the result of applying the filter that models the RIR to the first estimate of the target direct path signal is closest to the residual between the acoustic signal mixture and the first estimate of the target direct path signal according to a distance function (e.g., least squares distance function). In some embodiments, the filter corresponds to a linear filter structure estimated based on convolutional prediction. The first estimate is filtered forward frequency-by-frequency in the time-frequency domain using the linear filter of the convolutional prediction (described in FIGS. 3A, 3B, 4, 5, and 6). In some exemplary embodiments, the received acoustic signal mixture includes speech signals from a plurality of speakers. The first DNN generates a plurality of outputs, each output including a first estimate of the target direct path signal for one of the plurality of speakers. In some embodiments, the early reflections (e.g., early reflection 320B) and late reverberations (e.g., late reverberation 320C) of the first estimate can be identified based on the RIR modeled by the filter. The identified early reflections and late reverberations can be removed from the first estimate to estimate the acoustic signal mixture.
[0152] In operation 916, the mixture with reduced reverberation of the target direct path signal 106A is obtained by removing from the received mixture the result of applying a filter to a first estimate 408 of the target direct path signal 106A. In some embodiments, the second DNN can be trained based on a training data set created from extended data obtained using a set of estimated filters and a set of estimated target direct path signals to create the reverberant mixture. For example, the extended data includes external data measured outside the room or one or more of the other acoustic signal mixtures within the training data set. Next, a reverberation removal process based on the techniques described above is performed on this extended data. Further, the results of applying the filter and the reverberation removal process include extended data that is then used to further train the second DNN.
[0153] In operation 918, the mixture with reduced reverberation is input into a second DNN (e.g., DNN 2206B) to generate a second estimate of the target direct path signal. In some exemplary embodiments, one or a combination of the received acoustic signal mixture and the first estimate of the target direct path signal is input into the second DNN to generate a second estimate of the target direct path signal. In some other exemplary embodiments, the received acoustic signal mixture, the first estimate of the target direct path signal, and the mixture with reduced reverberation are input into the second DNN to generate a second estimate of the target direct path signal. In some yet other exemplary embodiments, the first estimate of the target direct path signal and the mixture with reduced reverberation are input into the second DNN to generate a second estimate of the target direct path signal. In some embodiments, the second DNN is trained based on a training data set created from extended data obtained using a set of estimated filters and a set of estimated target direct path signals such that a reverberant mixture can be created.
[0154] In operation 920, a second estimate of the target direct path signal is output via an output interface such as output interface 210. To further improve the reverberation removal of the speech signal, the steps of estimating a filter, obtaining a mixture with reduced reverberation, and inputting the mixture with reduced reverberation can be repeated for each of the plurality of outputs of the first DNN. Also, the output interface can be configured to output the RIR modeled by the filter. The output RIR can be used for performing speech analysis for one or a combination of indoor acoustic parameter analysis, room geometry reconstruction, speech enhancement, and reverberation removal of the speech signal.
[0155] In some exemplary embodiments, the reverberation removal of the speech signal using the estimates, i.e., the first and second estimates of the target direct path signal, and the filter of the target direct path signal, is evaluated for three tasks: 1) speech reverberation removal using weak stationary noise, 2) two-speaker separation in reverberant situations using white noise, and 3) two-speaker separation in reverberant situations using highly difficult non-stationary noise. The evaluation results are shown in FIGS. 10, 11, and 12.
[0156] FIG. 10 is a diagram showing a tabular representation 1000 corresponding to a simulated test set for reverberation removal of a speech signal according to an embodiment of the present disclosure. The tabular representation 1000 shows the data set used for reverberation removal, the reverberant speaker separation and speech enhancement tasks, the hyperparameter settings, and the baseline system for reverberation removal of the speech signal. The tabular representation 1000 also shows the results regarding the ASR task of the REVERB corpus.
[0157] For the reverberation removal of speech signals, DNNs such as DNN1206A and DNN2206B can be trained using a simulated reverberation dataset in a state where the air-conditioning noise is weak. In addition to evaluating the trained DNN on a simulated test set, the DNN is directly applied to the Reverberant Voice Enhancement and Recognition Benchmark (REVERB) corpus to show its effectiveness for processing actually recorded noisy reverberant utterances. The REVERB corpus is a benchmark for the evaluation of automatic speech recognition technology. The dataset also includes clean signals for simulation obtained from the WSJCAM0 corpus. The WSJCAM0 corpus contains 7,861 utterances, 742 utterances, and 1,088 utterances in its training set, validation set, and test set, respectively. Using these utterances in the WSJCAM0 corpus, reverberant mixtures containing 39,305 (7,861×5) noises, reverberant mixtures containing 2,968 (742×4) noises, and reverberant mixtures containing 3,264 (1,088×3) noises are simulated as the training set, validation set, and test set, respectively. Subsequently, a data spatialization process is executed, and for each utterance, at random room characteristics as well as speaker and microphone positions, the room is randomly sampled using the RIR estimated for the reverberation removal of the speech signal. The distance between the speaker and the microphone is sampled from the range [0.75, 2.5] m. The reverberation time (T60) is derived from the range [0.2, 1.3] seconds. For each utterance, diffuse air-conditioning noise is sampled from the REVERB corpus and added to the reverberant speech of the speaker. The signal-to-noise ratio between the anechoic speech and the noise is sampled from the range [5, 25] dB. The sampling rate is 16 kHz.
[0158] The trained model is applied to actual reverberation recordings without retraining and is applied to the ASR task of REVERB. The test mixtures are obtained from actual recordings made in a room (e.g., Environment 100) with a reverberation time T60 of approximately 0.7 seconds, a distance between the speaker and the microphone of approximately 1 m in the near field and 2.5 m in the far field. The recorded noise is diffuse air-conditioning noise and is weak.
[0159] To construct a backend for ASR trained using noisy-reverberant speech and the clean source signals of REVERB, the official REVERB corpus is used in software such as Kaldi. In an exemplary embodiment, subsequently, a plug-and-play approach for ASR is implemented and the enhanced time-domain signal is directly input to the backend for decoding.
[0160] For the reverberant speaker separation task, the 6-channel Spatialized Multi-Speaker Wall Street Journal (SMS-WSJ) dataset is used. The SMS-WSJ dataset contains simulated two-speaker mixtures in reverberant scenarios. The clean speech is sampled from the WSJ0 and WSJ1 datasets. The corpus contains 33,561 two-speaker mixtures, 982 two-speaker mixtures, and 1,332 two-speaker mixtures for training, validation, and testing, respectively. The distance between the speaker and the array is sampled from the range [1.0, 2.0] m, and T60 is derived from the range [0.2, 0.5] seconds. Weak white noise is added to simulate microphone noise. The energy level between the sum of the reverberant target speech signals and the noise is sampled from the range [20, 30] dB. The sampling rate is 8 kHz. The first channel of the 6-channel SMS-WSJ dataset is used for training and evaluation. Additionally, the direct sound is used as the training target, and both the reverberation removal task and the separation task are performed.
[0161] For ASR, the default Kaldi-based backend acoustic model defined in the SMS-WSJ dataset is used. This model is trained using single-speaker noisy reverberant speech as input and the state alignment of its corresponding direct-path signal as labels. Signals in the first, third, and fifth channels (i.e., more signals than microphones) are used for training the acoustic model. The task standard trigram language model is used for decoding.
[0162] The noisy reverberant speaker separation task is evaluated using the noisy reverberant WHAMR! (WSJ0 Hipster Ambient Mixtures) dataset. WHAMR! pairs the two-speaker mixtures in the wsj0-2mix dataset with the noise backgrounds scenes used for noisy reverberant binaural two-speaker separation. In this evaluation, the clean two-speaker mixtures are reused from the WSJ0-2mix dataset, and each clean signal is reverberated and the non-stationary environmental noise recorded in WHAM! is added. The reverberation time T60 is randomly sampled from the range [0.2, 1.0] seconds. The signal-to-noise ratio between the louder speaker and the noise is derived from the range [-6, 3] dB. The energy level between the two speakers in each mixture is sampled from the range [-5, 5] dB. The distance between the speaker and the array is sampled from the range [0.66, 2.0] m. The training set, validation set, and test set each have 20,000 binaural mixtures, 5,000 binaural mixtures, and 3,000 binaural mixtures, respectively. The corpus used is the 1-minute and 8 kHz version.
[0163] In the case of STFT, the window length is 32 milliseconds, the hop size is 8 milliseconds, and the analysis window is the square root of the Hann window. When the sampling rate is 16 kHz, a 512-point FFT is applied to extract 257-dimensional STFT features, and when the sampling rate is 8 kHz, a 256-point FFT is used to extract 129-dimensional features. Sentence-level or global-level mean-variance normalization is not performed on the input features. For each mixture, its sample variance is normalized to 1 before any processing. During training, the target signal needs to be scaled by the same coefficient as that used for scaling the mixture.
[0164] In the case of WPE and DNN-WPE, the number of filter taps K is set to 37, and the filter delay Δ is set to 3. The number of iterations in WPE is set to 3. PSD context is not used. Based on the validation set, K and Δ are adjusted to 40 and 0, 39 and 1, 38 and 2, 37 and 3, and 36 and 4, among which setting the filter taps and filter delay to 37 and 3 worked best across the entire dataset. For convolutional prediction, K is set to 40, resulting in the same amount of context as in WPE. This means that the filter length in the time domain is 344 (= (40 - 1) × 8 + 32) milliseconds. The filter taps K are increased up to 125, which corresponds to an RIR length up to 1.0 second. This incurs an increase in the amount of computation spent on the linear regression step, but there is no significant difference in terms of the evaluation score. The RIR mostly has energy within the range of 0.35 seconds after the peak impulse. The floor value ε used for calculating the reverberation removal result is set to 1.0, indicating that no weights are used, or set to 0.001. The PSD in each T-F unit will be 30 dB lower than the T-F unit with the highest energy.
[0165] For all tasks, the main evaluation metric is the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). The SI-SDR measures the quality of time-domain sample-level prediction. The extended Short-Time Objective Intelligibility (eSTOI) and the Perceptual Evaluation of Speech Quality (PESQ) scores are measured. For PESQ, the narrowband MOS-LQO score is reported based on the ITU P.862.1 standard using the Pearson-pesq toolkit. The reference for metric calculation is used from the target direct-path signal obtained by setting the reverberation time T60 parameter to zero in the RIR. The Word Error Rate (WER) of the ASR is also shown in Table Format 1000.
[0166] In Table Format 1000, the target direct-path signal is represented by "d", the target direct-path signal with early reflections is represented by "d+e", and the target direct-path signal with early reflections and noise is represented by "d+e+v".
[0167] As shown in Table Format 1000, when the first estimate of the first DNN (DNN1) is considered the final prediction, the training target of DNN1 shows better performance than the other two (i.e., "d+e" and "d+e+v"). There is no significant difference in DNN1-WPE, which applies the DNN1 output to improve WPE, compared to using various things as the training target of DNN1. However, it can be seen that training DNN1 using the target direct-path signal shows a performance improvement in DNN1+DNN2, which stacks two DNNs using the acoustic signal mixture and the output of DNN1, i.e., DNN1+DNN2 that trains the second DNN2 using the first estimate of the target direct-path signal.
[0168] In addition, the tabular representation 1000 includes a comparison of using the Inverse Convolutive Prediction (ICP) method, the Forward Convolutive Prediction (FCP) method, or the Weighted Prediction Error (WPE) method between two DNNs, namely DNN1 and DNN2. DNN1 + FCP + DNN2 with the floor value ε set to 0.001 shows better performance than DNN1 + WPE + DNN2 and DNN1 + ICP + DNN2. As shown in the tabular representation 1000, by performing linear or convolutional prediction and DNN2 one or more times repeatedly at runtime, DNN1+(WPE + DNN2)×2 and DNN1+(ICP + DNN2)×2 show a slight improvement in terms of SI-SDR and PESQ, and a slight decrease in terms of the Word Error Rate (WER). On the other hand, DNN1+(FCP + DNN2)×2 shows an improvement in all metrics. These results indicate that the DNN1 + FCP + DNN2 approach is more effective than WPE and DNN1 + WPE + DNN2.
[0169] In DNN1+ICP+DNN2, the SI-SDR score and PESQ score were improved by setting the floor value ε to 1.0. When the floor value is set to 0.001, the SI-SDR score and PESQ score in DNN1+FCP+DNN2 are further improved. For example, when the floor value is 1.0, the SI-SDR score is 11.9 and the PESQ score is 3.15. When the floor value is 0.001, the SI-SDR score is 12.3 and the PESQ score is 3.18. The floor values of 1.0 and 0.001 are also used to evaluate the trained DNN1 using ICP and FCP. As shown in tabular representation 1000, in DNN1+ICP with a floor value of 1.0, the SI-SDR score is 3.2 and the PESQ score is 1.78, in DNN1+ICP with a floor value of 0.001, the SI-SDR score is 0.7 and the PESQ score is 1.77, in DNN1+FCP with a floor value of 1.0, the SI-SDR score is 3.6 and the PESQ score is 1.82, and in DNN1+FCP with a floor value of 0.001, the SI-SDR score is 3.0 and the PESQ score is 1.82. Therefore, DNN1+FCP+DNN2 shows better scores than training DNN1 using the ICP method and the FCP method.
[0170] Overall, for voice reverberation removal, the mixture SI-SDR and PESQ are improved from -3.6 dB and 1.64 to 8.2 dB and 2.65 by using one DNN (i.e., DNN1), to 9.1 dB and 2.82 by using two DNNs (i.e., DNN1+DNN2), to 12.3 dB and 3.18 by adding an FCP module between the two DNNs (DNN1+FCP+DNN2), and to 12.8 dB and 3.24 by using one additional repetition for FCP and DNN2 (DNN1+(FCP+DNN2)×2).
[0171] Finally, during the training of the second DNN2, a loss in the magnitude domain is added. While improvements in word error rate (WER) and PESQ are obtained, SI-SDR decreases by approximately 0.5 dB.
[0172] FIG. 11 shows a tabular representation 1100 showing evaluation results for reverberation removal of speech signals using a test data set according to an embodiment of the present disclosure. The evaluation results show the performance regarding the SMS-WSJ data set and the oracle results obtained by using oracle masks such as a target direct-path signal with early reflections or a target direct-path signal without early reflections and a spectral intensity mask (|S| / |Y|) and a phase-sensitive mask (|S| / |Y|cos(∠S - ∠Y)). As shown in the tabular representation 1100, by using an oracle target direct-path signal for ASR, a better WER is obtained than using a target direct-path signal with early reflections (6.4% vs. 7.04%), which indicates the potential benefit of removing early reflections.
[0173]
Number
[0174] For DNN-WPE, two variants for the multi-speaker scenario are used. The first variant calculates a different WPE for each speaker using the PSD of each estimated target speaker generated by DNN1. In the tabular representation 1100, DNN-WPE in the multi-speaker scenario is represented as DNN1 + mfWPE + DNN2, where "mf" indicates a multi-filter. The multi-filter sums up all the estimated target speakers provided by DNN1 and calculates a single WPE filter using the PSD of the summed signal to remove reverberation from the mixture. The second variant is represented as DNN1 + sfWPE + DNN2, where "sf" indicates a single filter.
[0175] As shown in the tabular representation 1100, in DNN1+sfWPE+DNN2, slightly better performance was obtained than in DNN1+mfWPE+DNN2, suggesting that calculating separate filters for each target speaker is not effective for WPE.
[0176] The scenario where all speakers provide speech signals is represented by "allSpks" in the tabular representation 1100, and DNN2 is trained to emphasize all target speakers simultaneously. As shown in the tabular representation 1100, compared to DNN1+sfWPE+DNN2 and DNN1+ICP+DNN2, DNN1+FCP+DNN2 shows excellent performance in all metrics. This proves that forward filtering of convolutional prediction (explained in Figures 5 and 6) is more effective than WPE in reverberation removal in the presence of competing speakers.
[0177] When DNN2 is trained to emphasize target speakers one by one as explained in Figure 6 (represented by "perSpk" in the tabular representation 1100), further improvement is achieved. This suggests that reverberation removal for each speaker individually can improve the voice emphasis of the speaker. As shown in the tabular representation 1100, by repeating convolutional prediction and DNN2 one or more times, gradual improvement can be realized. Also, DNN2 trained by including magnitude-level loss improves PESQ, eSTOI, and WER, but decreases SI-SDR.
[0178] In the tabular representation 1100, it is further shown that for DNN1+(FCP+DNN2)×2 trained using the magnitude-level loss function, the scores of SI-SDR, PESQ, eSTOI, and WER are 12.2, 3.24, 89.0, and 12.77 respectively. DNN1+(FCP+DNN2)×2 trained using the magnitude-level loss function can function better in terms of performance than DNN1+(FCP+DNN2)×2 trained using another complex spectral mapping for a single microphone such as a single-input single-output microphone (SISO1) (SI-SDR is 12.5 dB vs. 5.1 dB). DNN1+(FCP+DNN2)×2 trained using the magnitude-level loss function can function better than DNN1+(FCP+DNN2)×2 trained using DPRNN-TasNet (SI-SDR is 12.5 dB vs. 6.5 dB).
[0179] Also, the tabular representation 1100 shows the performance of DNN1 and DNN2 trained based on spectral mapping corresponding to an array of microphones such as a 6-microphone SISO (SISO1-BF-SISO2) having beamforming of the microphone array, where the SISO combines a monaural complex spectral mapping, beamforming, and post-filtering. These results suggest that combining an end-to-end DNN and convolutional prediction can be effective in reducing reverberation in an acoustic signal mixture containing speech signals of speakers (e.g., speakers 102A and 102B).
[0180] FIG. 12 is a diagram showing a tabular representation 1200 of evaluation results for reverberation removal of speech signals using a test data set according to some other embodiments of the present disclosure. The tabular representation 1200 shows SI-SDR for the WHAMR! data set. As shown in the tabular representation 1200, DNN1+FCP+DNN2 produces better results than DNN1+mfWPE+DNN2 (SI-SDR is 7.4 dB vs. 6.8 dB). This indicates that DNN-FCP can be more robust than DNN-WPE in reverberation removal in the presence of noise and competing speakers.
[0181] Also, the tabular representation 1200 shows a comparison with end-to-end speech separation systems such as Wavesplit. For DNN1+(FCP+DNN2)×2, the SI-SDR score is 7.5 dB, which is higher than the SI-SDR score of Wavesplit, i.e., 5.9 dB. Wavesplit may use speaker identity as secondary information during training for target speaker extraction. DNN1+(FCP+DNN2)×2 does not rely on the availability of speaker identity information. Also, dynamic mixing may be applied for data augmentation, which results in a better SI-SDR (7.1 dB). DNN1+(FCP+DNN2)×2 may be trained without data augmentation, and this works better than Wavesplit with dynamic mixing.
[0182] FIG. 13 is a block diagram showing an audio signal processing system 1300 according to an embodiment of the present disclosure. The audio signal processing system 1300 uses the system 200. In some exemplary embodiments, the system 200 having a DNN for reverberation removal of speech signals, such as DNNs 1206A and 2206B, may be implemented on a remote server or within a cloud network. In some embodiments, the audio signal processing system 1300 (hereinafter referred to as the system 1300) may receive an RIR model, such as the RIR model 316A, to the audio signal processing system 1300. The system 1300 may process this RIR model to perform audio analysis for at least one or a combination of room geometry reconstruction, voice enhancement, and reverberation removal of speech signals.
[0183] In some exemplary embodiments, the system 1300 includes one or more sensors 1302, such as acoustic sensors, that collect data including the acoustic signal 1204 from the environment 1306. The environment 1306 corresponds to the environment 100.
[0184] The acoustic signal 1304 may include one or more target direct path signals and their reverberations. For example, the acoustic signal 1304 may include multiple speakers with overlapping speech and their reverberations. Further, the sensor 1302 may convert an acoustic input into the acoustic signal 1304.
[0185] The audio signal processing system 1300 includes a hardware processor 1308 that communicates with a computer storage memory such as the memory 1310. The memory 1310 includes stored data that includes algorithms, instructions, and other data that can be executed by the hardware processor 1308. Depending on the requirements of a particular application, it is conceivable that the hardware processor 1308 may include two or more hardware processors. These two or more hardware processors can be either internal or external. The audio signal processing system 1300 may be incorporated into other components, particularly an output interface and a transceiver, among other devices.
[0186] In some alternative embodiments, the hardware processor 1308 can be connected to a network 1312, and the network 1312 communicates with one or more data sources 1314, computer devices 1316, mobile phone devices 1318, and storage devices 1320. The network 1312 can include, by way of non-limiting example, one or more local area networks (LANs) and / or wide area networks (WANs). Also, the network 1312 can include enterprise-scale computer networks, intranets, and the Internet. The audio signal processing system 1300 can include one or more client devices, storage components, and data sources. Each of the one or more client devices, storage components, and data sources can include a single device or can include multiple devices that cooperate in a distributed environment of the network 1312.
[0187] In some other alternative embodiments, the hardware processor 1308 may be connected to a network-enabled server 1322 connected to the client device 1324. The hardware processor 1308 may be connected to an external memory device 1326 and a transmitter 1328. Further, outputs may be output for each target speaker according to the intended use 1330 of a particular user. For example, the intended use 1330 of a particular user may correspond to displaying speech as text (such as a speech command) on one or more display devices such as a monitor or screen, or inputting the text for each target speaker for further analysis into a computer-related device.
[0188] The data source 1314 may include data resources for training DNNs such as DNN1206A and DNN2206B for the voice separation task. For example, in one embodiment, the training data may include acoustic signals of a plurality of speakers such as speakers 102A and 102B speaking simultaneously. Also, the training data may include acoustic signals of a single speaker speaking alone, acoustic signals of one or more speakers speaking in a noisy environment, and acoustic signals of a noisy environment (e.g., environment 100 having a reverberant noise signal 110A).
[0189] Also, the data source 1314 may include data resources for training DNN1206A and DNN2206B for the speech recognition task. The data provided by the data source 1314 may include labeled data and unlabeled data such as transcribed data and non-transcribed data. For example, in one embodiment, the data may include one or more sounds and may also include corresponding transcription information or labels that can be used for the initialization of the speech recognition task.
[0190] Furthermore, the unlabeled data in the data source 1314 may be provided by one or more feedback loops. For example, usage data from verbal search queries executed on a search engine may be provided as non-transcribed data. Other examples of data sources, by way of example and not limitation, include streaming sound or video, web queries, mobile device camera or voice information, webcam feeds, smart glass and smart watch feeds, customer care systems, security camera feeds, web documents, catalogs, user feeds, SMS logs, instant messaging logs, transcripts of spoken words, voice commands or captured images (e.g., depth camera images) such as game system user interactions, tweets, chat or video call recordings, or various spoken language voice or image sources including social networking media. The particular data source 1314 used may be determined based on an application including whether the data is data related to a particular class of data (e.g., data related only to a particular type of sound including machine systems, entertainment systems) or is in fact general (not specific to a class) data.
[0191] In addition, the voice signal processing system 1300 may include a third-party device that can be configured with any type of computing device, such as an automatic speech recognition (ASR) system on a computing device. For example, the third-party device may include a computer device or a mobile device 1318. The mobile device 1318 may include a personal data assistant (PDA), a smartphone, a smartwatch, smart glasses (or other wearable smart devices), an augmented reality headset, a virtual reality headset, a laptop, a tablet, a remote control device, an entertainment system, a vehicle computer system, an embedded system controller, an appliance, a home computer system, a security system, a consumer electronics device, or other similar electronic devices. Also, the mobile device 1318 may include a microphone or a line input terminal for receiving voice information, a camera for receiving video information or image information, or a communication component (e.g., Wi-Fi function) for receiving such information from another source such as the Internet or a data source 1314. In one exemplary embodiment, the mobile device 1318 may be capable of receiving input data such as voice information and image information. For example, the input data may include a speaker's query to the microphone of the mobile device 1318 while multiple speakers in a room are speaking. To determine the content of the query, the input data may be processed by the ASR within the mobile device 1318 using the system 200. The system 200 enhances the input data by reducing noise in the speaker's environment, separating the speaker from other speakers, or emphasizing the voice signal of the query, so that the ASR can output an accurate response to the query.
[0192] In some exemplary embodiments, storage 1320 may store information including data, computer instructions (e.g., software program instructions, routines, or services), and / or data related to DNNs such as DNN 1206A and DNN 2206B of system 200. For example, storage 1320 may store data from one or more data sources 1314, one or more deep neural network models, information for generating and training deep neural network models, and computer-usable information output by one or more deep neural network models.
[0193] FIG. 14A is a block diagram showing a system 1400A for reverberation removal of a speech signal according to some exemplary embodiments of the present disclosure. System 1400A can be used to estimate a target audio signal from an input audio signal 1402 obtained from a sensor 1404 that monitors environment 1406.
[0194] The input audio signal 1402 includes an acoustic signal mixture that includes a target direct path signal (e.g., target direct path signal 106A) and a corresponding reverberation (e.g., reverberation 108A). The system 1400 processes the audio signal 1402 via the processor 1408 using the feature extraction module 1410. The feature extraction module 1410 calculates an audio feature sequence from the input audio signal 1402. The first target direct path signal estimation module 1412 processes the audio feature sequence and outputs a first estimated value (e.g., the first estimated value 408 of the target direct path signal 106A). The first estimated value of the target direct path signal is processed by the filter estimation module 1414 to output a filter that models the room impulse response that affects the target direct path signal. For example, the target direct path signal can be affected to change to a target reverberation signal. The filter is applied to the first estimated value to output a mixture with reduced reverberation. The filter and the first estimated value are further processed by the target direct path reverberation reduction mixture estimation module 1416 that estimates a mixture with reduced target direct path reverberation. The mixture with reduced target direct path reverberation, the first estimated value, and the features are further processed by the second target direct path estimation module 1418 to calculate a signal estimated value 1424 (e.g., the second estimated value 410) of the target direct path signal. The signal estimated value 1424 is output via the output interface 1422. In some embodiments, the room impulse response modeled by the filter can be output via the output interface 1422. The output room impulse response can be used in an audio analysis application to perform one or a combination of room geometry reconstruction, audio enhancement, and reverberation removal of the speech signal.
[0195] In some exemplary embodiments, the network parameter 1420 may be input to the first target direct path signal estimation module 1412, the filter estimation module 1414, the target direct path reverberation reduction mixture estimation module 1416, and the second target direct path estimation module 1418. The network parameter 1420 may include labeled data and unlabeled data such as transcription data and non-transcription data for various sounds or utterances that may be used for initialization of the speech recognition task.
[0196] FIG. 14B is a block diagram showing a system 1400B for reverberation removal of a speech signal according to some other exemplary embodiments of the present disclosure.
[0197] System 1400B includes a processor 1426 configured to execute stored instructions and a memory 1428 storing instructions regarding a neural network 1430 that includes a voice separation network 1432 that enables voice separation and reverberation reduction. The processor 1426 may be a single-core processor, a multi-core processor, a graphics processing unit (GPU), a computing cluster, or any number of other configurations. The memory / storage 1428 may include random access memory (RAM), read only memory (ROM), flash memory, or other suitable memory systems. Also, the memory 1328 may include a hard drive, an optical drive, a thumb drive, an array of drives, or any combination thereof. The processor 1426 is connected via a bus 1434 to one or more input and output interfaces / devices. Further, the system 1400B may include one or more microphones 1438 connected via the bus 1434. The system 1400B is configured to receive / acquire a speech signal 1456 via one or more microphones 1438 or via a network interface 1452 and a network 1454 connected to a data source of the speech signal 1456.
[0198] The memory 1428 stores a neural network 1430 trained to convert an acoustic signal mixture including a speech signal mixture and corresponding reverberation into a separated speech signal with reduced reverberation. The processor 1426, which executes the stored instructions, performs voice separation using the neural network 1430 retrieved from the memory 1428. The neural network 1430 is trained to convert an acoustic signal including a speech signal mixture into a separated speech signal. The neural network 1430 may include a voice separation network 1432 trained to estimate the separated signal from the acoustic characteristics of the acoustic signal.
[0199] FIG. 15 shows a use case 1500 for reverberation removal of a speech signal according to some exemplary embodiments of the present disclosure. The use case 1500 corresponds to a teleconference room including a group of speakers such as speaker 1502A, speaker 1502B, speaker 1502C, speaker 1502D, speaker 1502E, and speaker 1502F (the group of speakers 1502A-1502F). The speech signals of one or more speakers among the group of speakers 1502A-1502F are received by the audio receiver 1506 of the device 1504. The audio receiver 1506 includes the system 200 and receives an acoustic speech signal of a certain speaker or one or more speakers from the group of speakers 1502A-1502F.
[0200] The voice receiver 1506 may include a single microphone and / or an array of microphones for receiving an acoustic signal mixture and a noise signal from the group of speakers 1502A-1502F in the teleconference room. These acoustic signal mixtures from the group of speakers 1502A-1502F can be processed using the system 200. For example, the system 200 may analyze the RIR model of the teleconference room. This RIR model can be used to generate the room geometry structure of the teleconference room. The room geometry structure can be used for the arrangement of the reflective boundaries in the teleconference room. For example, the corresponding room geometry structure can be used to determine the installation location of the speakers, the seating arrangement of the group of speakers 1502A-1502F, etc. to cancel out noise and other disturbances in the teleconference room. Further, the RIR model can be used to remove reflections and reverberations of the speech signals of one or more speakers among the group of speakers 1502A-1502F.
[0201] In the illustrated exemplary scenario, multiple speakers among the group of speakers 1502A - 1502F may output speech signals simultaneously. In such a scenario, system 200 reduces reverberation within the teleconference room and separates each speech signal of speakers 1502A - 1502F. Also, system 200 may perform beamforming on the acoustic signal mixture from the microphone array to enhance the corresponding speaker's speech signal among the group of speakers 1502A - 1502F. The enhanced speech signal can be used for the transcription of the speaker's utterance. For example, device 1504 may include an ASR module. The ASR module may receive the enhanced speech signal and output a transcription. The transcription may be displayed by the display screen of device 1504.
[0202] FIG. 16 is a diagram showing a use case 1600 for reverberation removal of speech signals according to some other exemplary embodiments of the present disclosure. Use case 1600 corresponds to a factory floor including one or more speakers such as speaker 1602A and speaker 1602B. This factory floor may have high reverberation signals and noise due to the operation of various industrial machines. Also, this factory floor may be equipped with an audio device 1604 to facilitate communication between a control operator of the factory floor (not shown) and one or more speakers 1602A and 1602B within the factory floor. Audio device 1604 may include system 200.
[0203] In the illustrated exemplary scenario, audio device 1604 may be transmitting a voice command addressed to person 1602A who manages the factory floor. This voice command may include "Please report the status of machine 1". Speaker 1602A may utter "Machine 1 is in operation". However, the speech signal of speaker 1602A's utterance may be mixed with noise from the machine, noise from the background, and other utterances from speaker 1602B within the background.
[0204] Such noise and reverberation signals can be reduced by system 200. System 200 outputs clean speech of speaker 1602A. This clean speech is input to audio device 1604. Audio device 1604 receives this clean speech and captures a response to an audio command from the clean speech corresponding to the utterance of speaker 1602A. System 200 enables the audio device to improve communication with an intended speaker such as speaker 1602A.
[0205] FIG. 17 is a diagram showing a use case 1700 for reverberation removal of a speech signal according to some further other exemplary embodiments of the present disclosure. Use case 1700 corresponds to a driver assistance system 1702. Driver assistance system 1702 is implemented in a vehicle such as a manually operated vehicle, an automated vehicle, or a semi-automated vehicle. The vehicle is occupied by one or more persons such as person 1704A and person 1704B. Driver assistance system 1702 includes system 200. For example, driver assistance system 1702 may be remotely connected to system 1702 via a network such as network 1754. In some alternative exemplary embodiments, system 200 may be incorporated within driver assistance system 1702.
[0206] Additionally, the driver assistance system 1702 may include one or more microphones for receiving an acoustic signal mixture. This acoustic signal mixture may include speech signals from persons 1704A and 1704B and external noise signals such as the horn sound of other vehicles. In some cases, when person 1704A is transmitting a speech command to the driver assistance system 1702, another person 1704B may speak louder than person 1704A. The speech from person 1704B may interfere with the speech command of person 1704A. For example, the speech command of person 1704A may be "Find the nearest parking lot," and the speech of person 1704B may be "Find a shopping mall for parking." In such cases, the system 200 processes the speech of each of persons 1704A and 1704B simultaneously or separately. The system 200 separates the speech of person 1704A from the speech of person 1704B. The separated speech is used by the driver assistance system 1702. The driver assistance system 1702 may process and execute the speech command of person 1704A and the speech of person 1704B and, in response, output a response to each speech.
[0207] FIG. 18 shows a use case 1800 for reverberation removal of speech signals according to some further exemplary embodiments of the present disclosure. In some exemplary embodiments, systems 200a, 200b, 200c (illustrated in FIGS. 2A, 2B, 2C) may process pre-recorded data or live recordings of sound to determine an estimated value of a target direct path signal. The pre-recorded data of sound may be accessed from a database via a network 1808. The network 1808 is an example of the network 1312. Similarly, live recordings of the source may be streamed from a corresponding source at a remote location via the network 1808.
[0208] The estimated value of the target direct path signal can be filtered by the system 200 to determine the RIR model. This RIR model can be analyzed by an audio signal processing system such as the audio signal processing system 1300 connected to the system 200. The audio signal processing system 1300 can process the RIR model for room acoustic simulation 1802 of an environment such as a music concert hall 1806. The RIR model can be convolved using a recorded sound track source to encode the acoustics of the music concert hall 1806 based on the room acoustic simulation 1802. Using the room acoustic simulation 1802, a simulated environment or virtual reality environment of the actual situation of the music concert hall 1806 can be created. The simulated environment of the music concert hall 1806 can enable a rehearsal to be performed before a musician actually performs in the music concert hall 1806.
[0209] In some cases, the room acoustic simulation 1802 can be used to model the room acoustic behavior for room geometry reconstruction 1804. The room geometry reconstruction 1804 can provide an architectural aspect to the design and structure for maximizing the listening experience of the audience in a music concert hall such as the music concert hall 1806.
[0210] In some embodiments, the room acoustic simulation 1802 can be used to model room acoustic parameters such as at least one of the direct-to-reverberant ratio, reverberation time (RT60), early decay time (EDT), center time, clarity C80, and resolution D50. These parameters are well known in the art.
[0211] By incorporating operations 902 - 920 in the manner described above, methods 900a and 900b, which are executed using processor 208 disposed within system 200, can enable the estimation of a filter that includes both the magnitude and phase of reverberation, thus improving the removal of reverberation from speech signals. Since the filter is estimated based on a convolutional prediction approach, it enables the filter to reduce the initial reflections of the target direct-path signal. Further, since the filter models the signal propagation within the room, i.e., the RIR, it can improve the accuracy of the reverberation estimation. Also, the use of two DNNs in system 200 can improve the performance of reverberation removal from speech signals, as well as tasks such as voice enhancement and speaker separation. More specifically, the first DNN estimates a first estimate of the target direct-path signal from an acoustic signal mixture that includes reverberation. The second DNN uses the first estimate along with other data such as the filter and the reduction of reverberation estimated by the filter to estimate a refined estimate of the target direct-path signal. Thus, the two DNNs enable the identification and discrimination of the target direct-path signal from high reverberation and noise in an efficient and feasible manner.
[0212] Also, individual embodiments may be described as a process depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart can describe operations as a sequential process, but many of the operations can be executed in parallel or concurrently. Further, the order of the operations can be rearranged. A process can end when its operations are completed, but may have additional steps not discussed or included in the figure. Further, not all operations in a particular process occur in all embodiments. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the end of the function can correspond to the return of the function to the calling function or the main function.
[0213] Furthermore, embodiments of the disclosed subject matter can be implemented, at least in part, either manually or automatically. Examples of manual or automatic implementation can be executed or at least assisted by the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing the required tasks can be stored on a machine-readable medium. A processor may perform the required tasks.
[0214] The above-described embodiments of the present disclosure may be implemented in any of a number of ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or collection of processors, whether provided on a single computer or distributed among a plurality of computers. Such processors may be implemented as an integrated circuit and may include one or more processors within the integrated circuit component. However, the processor may be implemented using any suitable form of circuitry.
[0215] The various methods or processes outlined herein can be coded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Additionally, such software can be written using any of a number of suitable programming languages and / or programming tools or scripting tools and can be compiled as executable machine language code or intermediate code to be executed on a framework or virtual machine. The functionality of program modules can typically be combined or distributed as desired in various embodiments.
[0216] Embodiments of the present disclosure may be embodied as a method, and examples thereof are provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, even though shown as consecutive acts in the illustrative embodiments, embodiments may be constructed in which the acts are performed in a different order than that shown, including performing some acts simultaneously. Accordingly, it is an object of the appended claims to cover all such variations and modifications as fall within the true spirit and scope of the present disclosure.
[0217] The present disclosure has been described with reference to particular preferred embodiments, but it should be understood that various other adaptations and modifications may be made within the spirit and scope of the present disclosure. Accordingly, it is an aspect of the appended claims to cover all such variations and modifications as fall within the true spirit and scope of the present disclosure.
Claims
1. A method for estimating an indoor impulse response (room impulse response: RIR), the method using a processor coupled with stored instructions for executing the method, the instructions, when executed by the processor, performing the steps of the method, the method comprising: receiving, via a wired communication channel or a wireless communication channel, an acoustic signal mixture comprising a target direct-path signal propagating in a room and reverberation of the target direct-path in the room; inputting the received acoustic signal mixture into a first deep neural network (deep neural network: DNN) to generate a first estimate of the target direct-path signal; estimating a filter that models an indoor impulse response (RIR) representing a relationship between the target direct-path signal and the reverberation of the target direct-path signal; wherein the filter, when applied to the first estimate of the target direct-path signal, generates a result that is closest to at least one of the acoustic signal mixture and a residual between the acoustic signal mixture and the first estimate of the target direct-path signal according to a distance function, and the method further comprises: transmitting the filter that models the RIR. A method comprising the steps of:
2. The first DNN is trained to estimate the amplitude and phase of the target direct-path signal, applying the filter that models the RIR to the first estimate of the target direct-path signal includes operations related to both the estimated amplitude and the phase, The distance function between the result of applying the filter that models the RIR to the first estimate of the target direct-path signal and the acoustic signal mixture is measured in both the amplitude domain and the phase domain. The method according to claim 1.
3. The method according to claim 1, wherein the first DNN is pre-trained to obtain the first estimate of the target direct-path signal from the observed acoustic signal mixture.
4. Pre-training the first DNN comprises: A distance function defined based on the real and imaginary (RI) components of the first estimated value of the target direct-path signal in the first time-frequency domain and the RI components of the corresponding reference direct-path signal in the first time-frequency domain, A distance function defined based on the magnitude and phase obtained from the RI components of the first estimated value of the target direct-path signal in the first time-frequency domain and the corresponding magnitude and phase of the reference direct-path signal in the first time-frequency domain, A distance function defined based on the reconstructed waveform obtained from the RI components of the first estimated value of the target direct-path signal in the first time-frequency domain by reconstruction in the time domain and waveform of the reference direct-path signal, A distance function defined based on the RI components of the first estimated value in the second time-frequency domain obtained by further converting the waveform reconstructed in the second time-frequency domain and the RI components of the reference direct-path signal in the second time-frequency domain, A distance function defined based on the magnitude obtained from the RI components of the first estimated value of the target direct-path signal in the second time-frequency domain obtained by further converting the waveform reconstructed in the second time-frequency domain and the corresponding magnitude of the reference direct-path signal in the second time-frequency domain, wherein the method according to claim 3 is performed by minimizing a loss function including one or a combination of these. **Claim 5** The method according to claim 1, wherein the coefficients of the filter for modeling the RIR are estimated in the time-frequency domain. **Claim 6** The target direct-path signal is a signal transmitted from the source of the original signal to the sensor, The target direct-path signal represents the original signal at the source as would be measured by the sensor in the absence of reverberation. The reverberation of the target direct path signal represents one or more transmissions of the original signal from the source to the sensor via one or more paths longer than the shortest path such that the RIR is modeled with respect to the target direct path signal, the method according to claim 1.
7. The filter for modeling the RIR is a linear filter estimated using a linear convolution prediction module, the method according to claim 1.
8. The linear convolution prediction enhances the linear structure of the filter for modeling the RIR to approximate the linear reflection of the target direct path signal from the surfaces in the room, the method according to claim 7.
9. The filter for modeling the RIR is applied to the first estimated value of the target direct path signal in the time-frequency domain, The distance function is a weighted distance having weights at each time-frequency point in the time-frequency domain determined by one or a combination of the received acoustic signal mixture and the first estimated value of the target direct path signal, The distance function is based on the least squares distance, the method according to claim 1.
10. The method further comprises applying the filter for modeling the RIR to an external signal measured outside the room to synthesize a new acoustic mixture, The new acoustic mixture provides the acoustic effect of the propagation of the external signal in the room, the method according to claim 1.
11. The method according to claim 10 further comprises updating the first DNN using the new acoustic mixture.
12. Obtaining a mixture with reduced reverberation of the target direct path signal by removing the result of applying the filter for modeling the RIR to the first estimated value of the target direct path signal from the received acoustic signal mixture; Feeding the mixture with reduced reverberation into a second DNN to generate a second estimated value of the target direct path signal; The method according to claim 1 further comprises outputting the second estimated value of the target direct path signal via an output interface.
13. Updating the first estimate of the target direct path signal based on the application of the filter that models the RIR; Updating training data based on the received acoustic signal mixture and the updated first estimate of the target direct path signal as a pseudo label; The method according to claim 1, further comprising retraining the first DNN based on the updated training data.
14. Further comprising generating extended data based on the filter that models the RIR, wherein the extended data is obtained by applying the filter that models the RIR to data including one or more of external data measured outside the room and data obtained as one or more of the first measurement value and the second measurement value of the target direct path signal of another acoustic signal mixture in the training data, and the method further comprises creating an extended training data set based on the extended data. The method according to claim 12.
15. The method according to claim 14, further comprising training a system for automatic speech recognition using the extended training data set.
16. The method according to claim 14, further comprising training a system for audio event detection using the extended training data set.
17. further comprising processing the filter that models the RIR to perform audio analysis for at least one or a combination of analysis of room acoustic parameters, reconstruction of room geometry, voice enhancement, and reverberation removal of speech signals, The method according to claim 1, wherein the room acoustic parameters include at least one of a direct-to-reverberation ratio, a reverberation time (RT60), an early decay time (EDT), a center time, a clarity C80, and a resolution D50.
18. The received acoustic signal mixture includes speech signals from a plurality of speakers, The first DNN generates a plurality of outputs, Each output includes the first estimate of the target direct path signal for one of the plurality of speakers. The method according to claim 1.
19. The method according to claim 18, further comprising, for each speaker among the plurality of speakers, estimating a corresponding filter for modeling the corresponding RIR.
20. The method further comprises processing the filter for modeling the RIR using a downstream DNN trained to perform a task based on the RIR, wherein the task includes one or a combination of speaker identification, emotion recognition, event detection, and localization, according to claim 1.
21. A system for estimating a room impulse response (RIR), comprising an input interface configured to receive, via a wired communication channel or a wireless communication channel, an acoustic signal mixture including a target direct-path signal propagating in a room and reverberation of the target direct path in the room, a memory storing a first Deep Neural Network (DNN), a processor, wherein the processor inputs the received acoustic signal mixture into the first DNN to generate a first estimate of the target direct-path signal, is configured to estimate a filter for modeling the RIR representing a relationship between the target direct-path signal and the reverberation of the target direct-path signal, wherein when applied to the first estimate of the target direct-path signal, the filter generates a result that is closest to a residual between the acoustic signal mixture and the first estimate of the target direct-path signal according to a distance function, and the system further comprises an output interface configured to output the filter for modeling the RIR via the wired communication channel or the wireless communication channel.